VVC intra-frame prediction mode rapid selection method
By using the Sobel operator in VVC intra prediction to extract CU texture information, establish a relationship with the prediction mode, and quickly filter out the optimal mode, the problem of complex and low accuracy of RDCost calculations in the prior art is solved, and more efficient mode selection and better video quality are achieved.
Patent Information
- Application Number
- CN202311803205.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, when selecting the VVC intra prediction mode, RDCost calculation is complex and has low accuracy, resulting in the final selected mode being not the optimal mode, large memory usage, and poor video quality.
By performing data processing on the original image, the texture information of pixels in the CU is extracted, the relationship between the texture information and the intra prediction mode is established using the Sobel operator, and the optimal mode is quickly filtered out, reducing the time-consuming of SATD and RDO calculations.
It realizes the optimal mode of intra prediction quickly and accurately, reduces memory usage, improves video quality, and greatly accelerates the intra prediction mode selection process.
Smart Images

Figure CN120223880A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video compression encoding and decoding, and particularly relates to a fast intra prediction mode selection method for VVC. Background Art
[0002] While retaining the original HEVC framework, VVC also upgrades and improves in some aspects. Intra prediction in VVC includes 67 prediction modes, an increase of 32 modes compared to HEVC. Among them, the Planar mode and the DC mode are non-angle modes, and the remaining 65 modes are angle modes. Selecting a correct mode can result in better image quality and less storage space. The calculation of RDCost in intra prediction of VVC is the most time-consuming. Most of the existing technologies either sample the modes step by step to screen the optimal mode or use machine learning methods to screen the optimal mode, with a high error rate and being time-consuming. As Figure 1 shown, it is a schematic diagram of 67 prediction modes.
[0003] In the prior art, for example, as described in Chinese Patent CN201910496326.0, "A Fast Intra Prediction Angle Mode Selection Method for VVC": Calculate the SATD values of modes 2, 18, 34, 50, and 66, and compare the SATD values of the five modes to obtain the mode i0 corresponding to the minimum SATD value; calculate the SATD values of modes i0 - 8 and i0 + 8, and compare the SATD values of modes i0, i0 - 8, and i0 + 8 to obtain the mode i1 corresponding to the minimum SATD value; where if mode i0 - 8 or i0 + 8 does not exist, the calculation is not performed; if the mode with the minimum SATD value is i1 = 2 or i1 = 66, go to S4; otherwise, calculate the SATD values of modes i1 - 4 and i1 + 4, and compare the SATD values of modes i1, i1 - 4, and i1 + 4 to obtain the mode i2 corresponding to the minimum SATD value; determine the RDO mode set, and use the RDO technology of VVC to calculate the RDO values of each angle mode in the RDO mode set, and select the angle mode with the minimum RDO value as the optimal angle mode.
[0004] Therefore, the deficiencies of the prior art are:
[0005] The RDCost decision uses distortion and bits as metrics. When selecting the best mode, predictions are made on the modes. Since RDCost is very complex in calculation, to reduce the RDCost calculation, a certain number of modes are initially screened for RDCost calculation. Most of the existing optimal mode screening methods are relatively complex and inaccurate, resulting in the finally selected mode not being the optimal mode, leading to large memory occupancy and poor video quality.
[0006] In addition, the commonly used terms in the prior art include:
[0007] VVC: It is a video compression coding standard and a lossy compression scheme standard, which is an extension based on the HEVC / H.265 coding standard.
[0008] CU: The coding unit in VVC, with sizes of 128*128, 64*64, 32*32, 16*16, 8*8, and 4*4. The CU contains image data of the corresponding size.
[0009] Distortion: After the current CU is encoded, there is a certain loss of image quality compared with the CU before encoding, and this loss is the distortion.
[0010] Bitstream size: The size of the memory space used after the current CU is encoded.
[0011] RDCost: Generally, the larger the video compression bit rate, the smaller the distortion; and the smaller the bitstream, the larger the distortion. RDCost is an expression for calculating the bit rate and distortion, RDCost = D + λR. Where D is the distortion, R is the bitstream size, and λ is an external input. Through the comparison of RDCost, the mode is judged to maintain the balance between the bitstream size and the distortion.
[0012] Intra prediction: Without referring to other images, only the spatial redundancy information within the current encoded image is eliminated. Mode selection: There are 67 intra prediction modes in VVC, including 65 angular modes, PLANAR mode (mode 0), and DC mode (mode 1). The encoder will select one of these methods for encoding, and in the official model, the mode with the smallest RDCost is selected.
[0013] SATD: Use the Hadamard transform to roughly estimate the RDCost when the current CU uses a certain mode. It is much faster than RDCost, but the accuracy is worse than RDCost, and it is often used for fast mode selection. MPM mode list: The mode list constructed according to the optimal modes of the left and upper CUs of the current CU. Each CU has an MPM mode list.
[0014] Sobel operator: An operator used to detect image texture. Summary of the Invention
[0015] In order to solve the above problems, the purpose of this application is to: through data processing of the original image, extract effective features, quickly and accurately find the optimal intra prediction mode, so that the encoded video occupies less memory and has good image quality. Use the Sobel operator to extract the texture information of the pixels in the CU, establish the relationship between the texture information and the intra prediction mode, which can effectively reduce the time consumption of calculations such as SATD and RDO, and greatly accelerate the mode selection process of intra prediction.
[0016] Specifically, the present invention provides a fast intra prediction mode selection method for VVC, and the method includes the following steps:
[0017] S1. Feature extraction: Select the gradient direction and gradient magnitude of the feature selection image to detect the image texture direction and texture intensity in the CU block; further including:
[0018] S1.1. Expand the CU pixel values;
[0019] S1.2. Use the Sobel operator to calculate the gradient direction and corresponding gradient magnitude of each pixel in the CU;
[0020] S2. Establish the correspondence between the gradient direction and the prediction mode, that is, map the gradient direction to the prediction mode: Map the gradient direction calculated in step S1 to 65 angular modes of intra prediction;
[0021] S3. Use the ratio of the maximum gradient magnitude in the combination to the total gradient magnitude of all combinations and the MPM mode to comprehensively screen the optimal mode.
[0022] The step S1.1 of expanding the CU pixel values further includes:
[0023] To calculate the gradient direction and gradient magnitude in the CU block, it is necessary to expand the CU pixels. The expansion of CU pixel values is divided into two cases:
[0024] (1). The CU is not an edge block of the image, so the pixels around the CU are available, and directly fill with the available pixels around the CU;
[0025] (2). The CU is an edge block of the image. The CU is located at the edge of the image. If there are available pixels around, directly use the available pixels for filling. If there are no available pixels, use the nearest pixels for filling.
[0026] The step S1.1 includes CUs of size 4×4, and the expansion methods for CUs of other sizes are the same as those of size 4×4:
[0027] (1). The CU is not an edge block of the image. As shown in the following table, the part P inside the solid line is the pixels in the CU, and the part S outside the solid line is the available pixels around the CU;
[0028]
[0029] (2). The CU is an edge block of the image. As shown in the following table, the CU is in the upper right corner of the table. There are available pixels on the left and below the CU, and directly use the available pixels for filling. There are no available pixels above and to the right of the CU, and use the nearest pixels for filling;
[0030]
[0031] In step S1.2, the Sobel operator is used for feature extraction. The horizontal direction template and vertical direction template of the Sobel operator are shown in Table 1 and Table 2 below:
[0032]
[0033]
[0034] The gradient is obtained by multiplying and adding the template and the covered pixels one by one;
[0035] Assume the calculation of pixels P1 shown in Table 3 below and P6 shown in Table 4. The calculation process is as follows: The thin dashed line part is the position where the image is covered by the template. The middle position of the template is the pixel for which the feature needs to be calculated. Gx and Gy are the horizontal gradient and vertical gradient of the pixel respectively. The calculation of other pixels is the same as this, and only the pixel gradient and gradient magnitude of CU need to be calculated. The filled pixels do not need to be calculated;
[0036] Table 3:
[0037]
[0038] Horizontal direction gradient of P1: Gx P1 =(S19 - S1)+2*(P2 - S2)+(P6 - S3);
[0039] Vertical direction gradient of P1: Gy P1 =(S1 - S3)+2*(S20 - P5)+(S19 - P6);
[0040] Table 4:
[0041]
[0042] Horizontal direction gradient of P6: Gx P6 =(P3 - P1)+2*(P7 - P5)+(P11 - P9);
[0043] Vertical direction gradient of P6: Gy P6 =(P1 - P9)+2*(P2 - P10)+(S3 - P11);
[0044] The gradient magnitude can be obtained from the gradients in two directions. The gradient magnitude adopts an approximate calculation, which is simple in operation and faster in speed. The formula is as follows:
[0045] G = |G x |+|G y |
[0046] The corresponding gradient direction can be obtained through the gradients in two directions. The formula for calculating the gradient direction is as follows:
[0047]
[0048] According to the above two steps, the gradient direction and gradient amplitude of each pixel in the CU can be obtained.
[0049] In step S2, it further includes: dividing the angle mode 2 - 66 into 32 combinations according to the angle, as shown in the following table:
[0050]
[0051] That is, mapping the gradient direction calculated in step S1 to the 65 angle modes of intra prediction.
[0052] Step S3 further includes:
[0053] S3.1, according to the corresponding relationship between the gradient direction and the prediction mode, classifying the gradient - corresponding amplitudes into the corresponding combinations, adding the mode gradient amplitudes in the same combination, finding the combination Q with the maximum gradient amplitude, and the corresponding amplitude sum is T. The expression is as follows:
[0054] T = max(X i (i = 1, 2, 3...32)), where X is the amplitude sum of each combination;
[0055] S3.2, calculating the proportion P of the maximum amplitude T to the sum SUM of the gradient amplitudes of all combinations. The expression is as follows: P = T / SUM;
[0056] S3.3, setting the threshold Th to 0.5. If P is greater than the threshold Th, then perform step S3.4, indicating that the final angle mode is probably in this maximum interval; otherwise, perform step S3.5;
[0057] S3.4, performing SATD calculation on the modes in combination Q, selecting the mode with the minimum SATD as the optimal mode, and going to step S3.8;
[0058] S3.5, judging whether the mode in combination Q is in the MPM mode list. If it is in the MPM list, then enter step S3.6; otherwise, go to step S3.7;
[0059] S3.6, adding the modes in MPM, the DC mode, and the PLANAR mode to combination Q, removing duplicate modes, calculating the SATD within the combination, and selecting the mode with the minimum SATD as the optimal mode, and going to step S3.8;
[0060] S3.7. Add the modes in MPM, the DC mode, and the PLANAR mode to combination Q, calculate the RDCost within the combination, select the mode with the minimum RDCost as the optimal mode, and go to step S3.8; S3.8. Select the optimal mode.
[0061] Therefore, the advantages of this application are as follows:
[0062] 1. The method flow steps are simple and have high accuracy. The optimal mode can be screened out with only a few steps.
[0063] 2. The method calculation is simple. There is no very complex operation, and the result can be obtained with only a few simple calculations.
[0064] 3. It is fast. Due to the simplicity of the method flow and the method formula calculation, the running speed of the encoder is very fast. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0066] Figure 1 It is a schematic diagram of 67 prediction modes.
[0067] Figure 2 It is a schematic flowchart of the method of this application.
[0068] Figure 3 It is a schematic diagram of the specific flow steps of the method of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings.
[0070] The present invention belongs to an optimized implementation of a model quantization method in the field of model compression, and is a fast method for intra prediction mode selection in the VVC standard. This method uses the sobel operator to calculate the gradient direction and gradient magnitude, and maps the gradient direction to the prediction mode; uses the ratio of the maximum gradient magnitude in the combination to the gradient magnitudes of all combinations and the MPM mode to comprehensively screen the optimal mode. Specifically, as Figure 2 、 3 shown, the method includes:
[0071] S1. Feature extraction: The features select the gradient direction and gradient magnitude of the image to detect the image texture direction and texture intensity in the CU block;
[0072] The calculation steps of the gradient direction and gradient magnitude are as follows:
[0073] S1.1. To facilitate the calculation of the gradient direction and gradient magnitude in the CU block, it is necessary to expand the CU pixels. Taking a CU of size 4*4 as an example, the expansion method for other-sized CUs is the same. The expansion of CU pixel values is divided into two cases.
[0074] (1) The CU is not an edge block of the image, so the pixels around the CU are available. Directly fill with the available pixels around the CU. The schematic diagram is as follows. The part P inside the solid line is the pixels inside the CU, and the part S outside the solid line is the available pixels around the CU.
[0075]
[0076] (2) The CU is an edge block of the image. Taking the following figure as an example, the CU is in the upper right corner of the image. There are available pixels on the left and below the CU. Directly fill with the available pixels. There are no available pixels above and to the right of the CU, so fill with the nearest pixels.
[0077]
[0078] S1.2. Use the Sobel operator to extract features. The horizontal direction template and vertical direction template of the Sobel operator are shown in Table 1 and Table 2 below:
[0079]
[0080] The gradient is obtained by multiplying and adding the template with the covered pixels one by one.
[0081] Taking the calculation of pixels P1 and P6 shown in Table 3 as an example, the calculation process and schematic are as follows. The part with the thin dashed line is the position where the image is covered by the template. The middle position of the template is the pixel for which the feature needs to be calculated. Gx and Gy are the horizontal gradient and vertical gradient of the pixel respectively. The calculation of other pixels is the same. Only the pixel gradients and gradient magnitudes of the CU need to be calculated. The filled pixels do not need to be calculated.
[0082] Table 3:
[0083]
[0084] Horizontal direction gradient of P1: Gx P1 =(S19 - S1)+2*(P2 - S2)+(P6 - S3);
[0085] Vertical direction gradient of P1: Gy P1 =(S1 - S3)+2*(S20 - P5)+(S19 - P6);
[0086] Table 4:
[0087]
[0088] Horizontal gradient of P6: Gx P6 = (P3 - P1) + 2 * (P7 - P5) + (P11 - P9);
[0089] Vertical gradient of P6: Gy P6 = (P1 - P9) + 2 * (P2 - P10) + (S3 - P11);
[0090] The gradient magnitude can be obtained from the gradients in two directions. The gradient magnitude uses an approximate calculation, which is simple and faster in operation. The formula is as follows:
[0091] G = |G x | + |G y |
[0092] The corresponding gradient direction can be obtained from the gradients in two directions. The formula for calculating the gradient direction is:
[0093]
[0094] According to the above two steps, the gradient direction and gradient magnitude of each pixel in the CU can be obtained.
[0095] S2. Establish the relationship between the gradient direction and the prediction mode:
[0096] Map the gradient direction calculated in the previous step to the 65 angular modes of intra-frame prediction. The angular modes 2 - 66 are divided into 32 combinations according to the angle, as shown in the following table:
[0097]
[0098] S3. Use the ratio of the maximum gradient magnitude in the combination to the gradient magnitudes of all combinations and the MPM mode to comprehensively screen the optimal mode;
[0099] S3.1. According to the correspondence between the gradient direction and the prediction mode, classify the magnitudes corresponding to the gradients into the corresponding combinations, add up the mode gradient magnitudes in the same combination, and find the combination Q with the maximum gradient magnitude. The corresponding magnitude sum is T. The expression is as follows,
[0100] T = max(X i (i = 1, 2, 3...32)), where X is the magnitude sum of each combination;
[0101] Assume that the modes in combination Q are 2 and 3, and the magnitude sum T is 6000;
[0102] S3.2. Calculate the ratio P of the maximum magnitude T to the sum SUM of the gradient magnitudes of all combinations. The expression is as follows: P = T / SUM; Assume SUM is 10000, P = 6000 / 1000 = 0.6;
[0103] In S3.3, the threshold Th is set to 0.5. If P is greater than the threshold Th, then proceed to step S3.4, indicating that the final angle mode is likely to be within this maximum interval; otherwise, proceed to step S3.5;
[0104] In S3.4, perform SATD calculation on the modes within the combination Q, select the mode with the smallest SATD as the optimal mode, and go to step S3.8;
[0105] The SATD is a common term in the field of coding and decoding. It uses the Hadamard transform to roughly estimate the RDCost when the current CU uses a certain mode. It is much faster than RDCost, but less accurate than RDCost, and is often used for fast mode selection;
[0106] The SATD calculation is as follows:
[0107] SATD = ∑∑|HXH|
[0108] where HXH is as follows, and r in the formula is the pixel value of the CU
[0109]
[0110] In S3.5, determine whether the modes in the combination Q are in the MPM mode list. If they are in the MPM list, then enter step S3.6; otherwise, go to step S3.7;
[0111] Suppose the modes in the MPM mode list are DC, 5, and 6, and the modes in the Q combination are 2 and 3. Then the modes in the combination Q are not in the MPM mode list;
[0112] In S3.6, add the modes in the MPM, the DC mode, and the PLANAR mode to the combination Q, remove duplicate modes, calculate the SATD of each mode within the combination, select the mode with the smallest SATD as the optimal mode, and go to step S3.8;
[0113] Suppose the modes in the MPM mode list are DC, 5, and 6, and the modes in the Q combination are 2 and 3. Add the MPM modes to the Q combination, and the mode list of the Q combination becomes DC, 5, 6, 2, and 3. Then add the DC and PLANAR modes. Since there is already a DC mode in the list, only the PLANAR mode needs to be added;
[0114] In S3.7, add the modes in the MPM, the DC mode, and the PLANAR mode to the combination Q, remove duplicate modes, calculate the RDCost within the combination, select the mode with the smallest RDCost as the optimal mode, and go to step S3.8;
[0115] Assume that the patterns in the MPM pattern list are DC, 5, and 6, and the patterns in the Q combination are 2 and 3. Add the MPM patterns to the Q combination. The pattern list of the Q combination becomes DC, 5, 6, 2, and 3. Then add the DC and PLANAR patterns. Since the DC pattern already exists in the list, only the PLANAR pattern needs to be added;
[0116] S3.8, select the optimal pattern.
[0117] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A fast intra prediction mode selection method for VVC, characterized in that The method includes the following steps: S1. Feature extraction: The gradient direction and gradient magnitude of the feature selection image are used to detect the image texture direction and texture intensity in the CU block; further including: S1.
1. Expand the CU pixel values; S1.
2. Use the Sobel operator to calculate the gradient direction and corresponding gradient magnitude of each pixel in the CU; S2. Establish the correspondence between the gradient direction and the prediction mode, that is, map the gradient direction to the prediction mode: Map the gradient direction calculated in step S1 to the 65 angular modes of intra-frame prediction; S3. Use the ratio of the maximum gradient magnitude in the combination to the total gradient magnitude of all combinations and the MPM mode to comprehensively screen the optimal mode.
2. The fast intra prediction mode selection method for VVC according to claim 1, wherein The step S1.1 of expanding the CU pixel values further includes: To calculate the gradient direction and gradient magnitude in the CU block, the CU pixels need to be expanded. The expansion of CU pixel values is divided into two cases: (1). The CU is not an edge block of the image, so the pixels around the CU are available, and the available pixels around the CU are directly used for filling; (2). The CU is an edge block of the image. The CU is located at the edge of the image. If there are available pixels around, the available pixels are directly used for filling. If there are no available pixels, the nearest pixels are used for filling.
3. A fast intra prediction mode selection method for VVC according to claim 2, characterized in that The step S1.1 includes CUs of size 4*4. The expansion method for CUs of other sizes is the same as that of 4*4 size: (1). The CU is not an edge block of the image, as shown in the following table. The part P inside the solid line is the pixel inside the CU, and the part S outside the solid line is the available pixel around the CU: (2). The CU is an edge block of the image, as shown in the following table. The CU is in the upper right corner of the table. If there are available pixels on the left and below the CU, the available pixels are directly used for filling. If there are no available pixels above and to the right of the CU, the nearest pixels are used for filling; 4. A fast intra prediction mode selection method for VVC according to claim 2, characterized in that, In the step S1.2, the Sobel operator is used for feature extraction. The horizontal direction template and vertical direction template of the Sobel operator are shown in Table 1 and Table 2 below: The gradient is obtained by multiplying and adding the template with the covered pixels one by one; Assume the calculation of pixels P1 shown in Table 3 below and P6 shown in Table 4 below. The calculation process is as follows: The part of the image covered by the template is shown by the thin dotted line. The middle position of the template is the pixel for which the feature needs to be calculated. Gx and Gy are the horizontal gradient and vertical gradient of the pixel respectively. The calculation of other pixels is the same as this, and only the pixel gradients and gradient magnitudes of the CU need to be calculated. The filled pixels do not need to be calculated; Table 3: Horizontal gradient of P1: Gx P1 =(S19 - S1)+2*(P2 - S2)+(P6 - S3); Vertical direction gradient of P1: Gy P1 =(S1 - S3) + 2 * (S20 - P5) + (S19 - P6); Table 4: Horizontal gradient of P6: Gx P6 = (P3 - P1) + 2 * (P7 - P5) + (P11 - P9); Vertical gradient of P6: Gy P6 = (P1 - P9) + 2 * (P2 - P10) + (S3 - P11); The gradient magnitude can be obtained from the gradients in two directions. The gradient magnitude adopts an approximate calculation, which is simple in operation and faster in speed. The formula is as follows: G = |G x | + |G y | The corresponding gradient direction can be obtained from the gradients in two directions. The gradient direction calculation formula: According to the above two steps, the gradient direction and gradient magnitude of each pixel in the CU can be obtained.
5. A fast intra prediction mode selection method for VVC according to claim 1, characterized in that, In the step S2, it further includes: Divide the angular modes 2 - 66 into 32 combinations according to the angle, as shown in the following table: That is, map the gradient direction calculated in step S1 to the 65 angular modes of intra-frame prediction.
6. A fast intra prediction mode selection method for VVC according to claim 1, characterized in that, The step S3 further includes: S3.
1. Classify the magnitudes corresponding to the gradients into the corresponding combinations according to the correspondence between the gradient direction and the prediction mode, add the mode gradient magnitudes in the same combination, find the combination Q with the maximum gradient magnitude, and the corresponding magnitude sum is T. The expression is as follows: T = max(X i (i = 1, 2, 3...32)), where X is the sum of the amplitudes of each combination; S3.
2. Calculate the ratio P of the maximum magnitude T to the sum SUM of the gradient magnitudes of all combinations. The expression is as follows: P = T / SUM; S3.
3. Set the threshold Th to 0.
5. If P is greater than the threshold Th, then perform step S3.4, indicating that the final angle mode is probably within this maximum interval; otherwise, perform step S3.5; S3.
4. Calculate the SATD for the modes in combination Q, select the mode with the minimum SATD as the optimal mode, and go to step S3.8; S3.
5. Determine whether the modes in combination Q are in the MPM mode list. If they are in the MPM list, then enter step S3.6; otherwise, go to step S3.7; S3.
6. Add the modes in MPM, the DC mode, and the PLANAR mode to combination Q, remove duplicate modes, calculate the SATD within the combination, select the mode with the minimum SATD as the optimal mode, and go to step S3.8; S3.
7. Add the modes in MPM, the DC mode, and the PLANAR mode to combination Q, calculate the RDCost within the combination, select the mode with the minimum RDCost as the optimal mode, and go to step S3.8; S3.
8. Select the optimal mode.
Citation Information
Patent Citations
VVC intra-frame prediction angle mode rapid selection method
CN110213586A