Method and apparatus for intra prediction
By determining chroma reference samples first and using them to select luma reference samples for linear model coefficient calculation, the method addresses coding errors in existing video coding methods, enhancing coding efficiency and prediction accuracy.
Patent Information
- Application Number
- JP2023173303
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-10-06
- Filing Date
- 2023-10-05
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-09-30
AI Technical Summary
Existing video coding methods face challenges in efficiently compressing video data, leading to coding errors due to the lack of corresponding chroma reference samples for available luma reference samples during intra prediction.
The proposed method determines the availability of chroma reference samples first, then uses the available chroma reference samples to determine the luma reference samples for calculating linear model coefficients, thereby enhancing the efficiency of cross-component linear model prediction (CCLM).
This approach improves the coding efficiency of video signals by reducing coding errors and enhancing the prediction accuracy of chroma blocks, leading to better compression ratios without sacrificing video quality.
Smart Images

Figure 0007693767000010 
Figure 0007693767000011 
Figure 0007693767000012
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This patent application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 742,266, filed on October 5, 2018; U.S. Provisional Patent Application No. 62 / 742,355, filed on October 6, 2018; U.S. Provisional Patent Application No. 62 / 742,275, filed on October 6, 2018; and U.S. Provisional Patent Application No. 62 / 742,356, filed on October 6, 2018. The foregoing patent applications are hereby incorporated by reference in their entirety.
[0002] Embodiments of the present disclosure generally relate to the field of video coding, and more specifically, to the field of intra - prediction using a cross - component linear model prediction (CCLM).
Background Art
[0003] The amount of video data required to depict even relatively short videos can be quite large, which can cause difficulties when streaming or otherwise communicating data over a communication network with limited bandwidth capacity. Thus, video data is typically compressed before being communicated over modern telecommunications networks. When storing video on a storage device, the size of the video can also be a problem since memory resources may be limited. Video compression devices typically use software and / or hardware at the source to encode the video data prior to transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that increase the compression ratio without sacrificing much or any of the video quality are desired since network resources are limited and the demand for higher video quality is increasing. High Efficiency Video Coding (HEVC) is the latest video compression issued by the ISO / IEC Moving Picture Experts Group and the ITU-T Video Coding Experts Group as ISO / IEC 23008-2 MPEG-H Part 2, i.e., ITU-T H.265, which doubles the data compression ratio at the same level of video quality or significantly improves the video quality at the same bitrate.
Summary of the Invention
[0004] Examples of the present disclosure provide an intra prediction apparatus and method for encoding and decoding an image, the intra prediction apparatus and method being capable of enhancing the efficiency of cross-component linear model prediction (CCLM), thereby enhancing the coding efficiency of a video signal. This disclosure is described in detail in the examples and claims included in this document.
[0005] The foregoing and other objects are achieved by the subject matter of the independent claims. Further embodiments are apparent from the dependent claims, the description of the specification and the figures.
[0006] Certain embodiments are outlined in the appended independent claims and other embodiments are outlined in the dependent claims.
[0007] According to a first aspect, the present disclosure relates to a method for performing intra prediction using a linear model. The method includes determining a luma block corresponding to a current chroma block; obtaining luma reference samples of the luma block based on determining L available chroma reference samples of the current chroma block, wherein the obtained luma reference samples of the luma block are downsampled luma reference samples; calculating linear model coefficients based on the luma reference samples and chroma reference samples corresponding to the luma reference samples; and obtaining a prediction of the current chroma block based on the linear model coefficients and values of the downsampled luma block of the luma block. The chroma reference samples of the current chroma block include reconstructed adjacent samples of the current chroma block. The L available chroma reference samples are determined from the reconstructed adjacent samples. Similarly, the adjacent samples of the luma block are also reconstructed adjacent samples of the luma block (i.e., reconstructed adjacent luma samples). In one example, the obtained luma reference samples of the luma block are obtained by downsampling reconstructed adjacent luma samples selected based on the available chroma reference samples.
[0008] In existing methods, luma reference samples are used to determine the availability of reference samples for determining linear model coefficients. However, in some scenarios, this can lead to coding errors because there are no corresponding chroma reference samples for the available luma reference samples. The techniques presented herein address this problem by determining the availability of reference samples after examining the availability of chroma reference samples. In some examples, the chroma reference samples are available when the chroma reference samples are not outside the current image, slice, or tile and the reference samples are reconstructed. In some examples, the chroma reference samples are available when the chroma reference samples are not outside the current image, slice, or tile, the reference samples are reconstructed, and the reference samples are not omitted based on coding decisions. The available reference samples for the current chroma block can be the reconstructed adjacent samples available for the chroma block. The luma reference samples corresponding to the available chroma reference samples are used to determine the linear model coefficients.
[0009] In a possible implementation of the method according to the first aspect itself, the step of determining L available chroma reference samples includes determining that the L upper adjacent chroma samples of the current chroma block are available, where 1 <= L <= W2, W2 represents the upper reference sample range, L and W2 are positive integers, and the L upper adjacent chroma samples are used as the available chroma reference samples.
[0010] In a possible implementation of the method according to the first aspect itself or any preceding implementation of the first aspect, W2 is equal to either 2*W or W + H, where W represents the width of the current chroma block and H represents the height of the current chroma block.
[0011] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the step of determining L available chroma reference samples includes determining that the L left adjacent chroma samples of the current chroma block are available, where 1 <= L <= H2, H2 represents the left reference sample range, L and H2 are positive integers, and the L left adjacent chroma samples are used as the available chroma reference samples.
[0012] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, H2 is equal to either 2*H or W+H, where W represents the width of the current chroma block and H represents the height of the current chroma block.
[0013] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the step of determining L available chroma reference samples includes determining that L1 upper adjacent chroma samples and L2 left adjacent chroma samples of the current chroma block are available, where 1 <= L1 <= W2, 1 <= L2 <= H2, W2 represents the upper reference sample range, H2 represents the left reference sample range, L1, L2, W2, and H2 are positive integers, L1 + L2 = L, and the L1 upper adjacent chroma samples and L2 left adjacent chroma samples are used as the available chroma reference samples.
[0014] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the luma reference sample is obtained by downsampling only the adjacent samples that are above the luma block and are selected based on L available chroma reference samples, or by downsampling only the adjacent samples that are to the left of the luma block and are selected based on L available chroma reference samples. For example, when L is 4, the luma reference sample is obtained by downsampling 24 adjacent samples that are above the luma block and are selected based on 4 available chroma reference samples, or by downsampling 24 adjacent samples that are to the left of the luma block and are selected based on 4 available chroma reference samples, and a 6-tap filter is used in the downsampling process.
[0015] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the downsampled luma block of the luma block is obtained by downsampling the reconstructed luma block of the luma block corresponding to the current chroma block.
[0016] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, when the luma reference sample is obtained based only on the adjacent samples above the luma block and the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), only one row of the reconstructed adjacent luma samples of the reconstructed version of the luma block is used to obtain the luma reference sample.
[0017] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the step of calculating the linear model coefficients based on the luma reference samples and the chroma reference samples corresponding to the luma reference samples includes: determining a maximum luma value and a minimum luma value based on the luma reference samples; obtaining a first chroma value based at least in part on the position of the luma reference sample associated with the maximum luma value; obtaining a second chroma value based at least in part on the position of the luma reference sample associated with the minimum luma value; and calculating the linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value.
[0018] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the step of obtaining a first chroma value based at least in part on the position of the luma reference sample associated with the maximum luma value includes obtaining the first chroma value based at least in part on one or more positions of one or more luma reference samples associated with the maximum luma value, and the step of obtaining a second chroma value based at least in part on the position of the luma reference sample associated with the minimum luma value includes obtaining the second chroma value based at least in part on one or more positions of one or more luma reference samples associated with the minimum luma value.
[0019] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the linear model coefficients α and β are calculated based on the following: α=(y B -y A ) / (x B -x A )、 β=y A -αx A where x B represents the maximum luma value, y B represents the first chroma value, x A represents the minimum luma value, and y A represents the second chroma value.
[0020] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the prediction of the current chroma block is obtained based on the following: pred C (i,j)=α·rec’ L (i,j)+β, where pred C (i,j) represents the predicted value of the chroma samples of the current chroma block, and rec’ L (i,j) represents the sample value of the corresponding luma sample of the downsampled luma block of the reconstructed luma block of the luma block.
[0021] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the luma reference samples are obtained based only on the adjacent samples on the left side of the luma block, and when the current chroma block is at the left boundary of the current coding tree unit (CTU), only one column of the reconstructed adjacent luma samples of the reconstructed luma block is used to obtain the luma reference samples.
[0022] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the linear model includes a multi-directional linear model (MDLM).
[0023] According to a second aspect, the present disclosure relates to a method for performing intra prediction using a linear model. The method includes determining a luma block corresponding to a current chroma block; obtaining luma reference samples of the luma block based on determining L available chroma reference samples of the current chroma block, wherein the obtained luma reference samples of the luma block are downsampled luma reference samples obtained by downsampling adjacent samples (i.e., reconstructed adjacent luma samples) of the luma block corresponding to the L available chroma reference samples; calculating linear model coefficients based on the luma reference samples and chroma reference samples corresponding to the luma reference samples; and obtaining a prediction of the current chroma block based on the linear model coefficients and the values of the downsampled luma block of the luma block. The chroma reference samples of the current chroma block include reconstructed adjacent samples of the current chroma block. The L available chroma reference samples are determined from the reconstructed adjacent samples. Similarly, the adjacent samples of the luma block are also reconstructed adjacent samples (i.e., reconstructed adjacent luma samples) of the luma block. The obtained luma reference samples of the luma block are obtained by downsampling the reconstructed adjacent luma samples corresponding to the available chroma reference samples.
[0024] Regarding the "reconstructed adjacent luma samples corresponding to the available chroma reference samples", it can also be understood that the correspondence between the reconstructed adjacent luma samples and the available chroma reference samples is not limited to a "one-to-one correspondence", and the correspondence between the reconstructed adjacent luma samples and the available chroma reference samples can be an "M-to-N correspondence". For example, when a 6-tap filter is used for downsampling, M = 24 and N = 4.
[0025] According to a third aspect, the present invention relates to an apparatus for encoding video data. The apparatus includes a video data memory and a video encoder. The video encoder is configured to determine a luma block corresponding to a current chroma block and, based on (or by) determining L available chroma reference samples of the current chroma block, obtain luma reference samples of the luma block, where the obtained luma reference samples of the luma block are downsampled luma reference samples obtained by downsampling 24 adjacent samples (i.e., reconstructed adjacent luma samples) of the luma block corresponding to the L available chroma reference samples; calculate linear model coefficients of a linear model based on the luma reference samples and chroma reference samples corresponding to the luma reference samples; and obtain a prediction of the current chroma block based on the linear model coefficients and values of a downsampled version of the luma block. For example, when L is 4, the obtained luma reference samples of the luma block are 4 downsampled luma reference samples obtained by downsampling 24 adjacent samples (i.e., reconstructed adjacent luma samples) of the luma block corresponding to the 4 available chroma reference samples, and a 6 - tap filter is used in the downsampling process.
[0026] According to a fourth aspect, the present invention relates to an apparatus for decoding video data. The apparatus has a video data memory and a video decoder. The video decoder is configured to determine a luma block corresponding to a current chroma block; obtain luma reference samples of the luma block based on determining L available chroma reference samples of the current chroma block, wherein the obtained luma reference samples of the luma block are downsampled luma reference samples; calculate linear model coefficients based on the luma reference samples and chroma reference samples corresponding to the luma reference samples; and obtain a prediction of the current chroma block based on the linear model coefficients and the values of the downsampled luma block of the luma block.
[0027] In existing approaches, luma reference samples are used to determine the availability of reference samples in order to determine linear model coefficients. However, in some scenarios, this can lead to coding errors because there are no corresponding chroma reference samples for the available luma reference samples. The technique presented herein addresses this problem by examining the availability of chroma reference samples and then determining the availability of reference samples. In some examples, the chroma reference samples are available when they are not outside the current image, slice, or tile and the reference samples have been reconstructed. In some examples, the chroma reference samples are available when they are not outside the current image, slice, or tile, the reference samples have been reconstructed, and the reference samples have not been omitted based on coding decisions. The available chroma reference samples of the current chroma block can be the available reconstructed adjacent samples of the chroma block (i.e., the available reconstructed adjacent chroma samples). The luma reference samples corresponding to the available chroma reference samples are used to determine the linear model coefficients.
[0028] In a possible embodiment of the apparatus according to the third or fourth aspect itself, determining the L available chroma reference samples includes determining that the L upper adjacent chroma samples of the current chroma block are available, where 1 <= L <= W2, W2 represents the upper reference sample range, L and W2 are positive integers, and the L upper adjacent chroma samples are used as the available chroma reference samples.
[0029] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, W2 is equal to either 2*W or W + H, where W represents the width of the current chroma block and H represents the height of the current chroma block.
[0030] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, determining the L available chroma reference samples includes determining that the L left adjacent chroma samples of the current chroma block are available, where 1 <= L <= H2, H2 represents the left reference sample range, L and H2 are positive integers, and the L left adjacent chroma samples are used as the available chroma reference samples.
[0031] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, H2 is equal to either 2*H or W + H, where W represents the width of the current chroma block and H represents the height of the current chroma block.
[0032] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, determining the L available chroma reference samples includes determining that the L1 upper adjacent chroma samples and the L2 left adjacent chroma samples of the current chroma block are available, where 1 <= L1 <= W2, 1 <= L2 <= H2, W2 represents the upper reference sample range, H2 represents the left reference sample range, L1, L2, W2, and H2 are positive integers, L1 + L2 = L, and the L1 upper adjacent chroma samples and the L2 left adjacent chroma samples are used as the available chroma reference samples.
[0033] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, the luma reference samples are obtained by downsampling only the adjacent samples that are above the luma block and are selected based on the L available chroma reference samples, or by downsampling only the adjacent samples that are to the left of the luma block and are selected based on the L available chroma reference samples.
[0034] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, the downsampled luma block of the luma block is obtained by downsampling the reconstructed luma block of the luma block corresponding to the current chroma block.
[0035] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, when the luma reference samples are obtained based only on the adjacent samples above the luma block and the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), only one row of the reconstructed adjacent luma samples of the reconstructed version of the luma block is used to obtain the luma reference samples.
[0036] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, the linear model coefficients α and β are calculated based on the following, α=(y B -y A ) / (x B -x A ), β=y A -αx A where x B represents the maximum luma value, y B represents the first chroma value, x A represents the minimum luma value, and y A represents the second chroma value.
[0037] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, the prediction of the current chroma block is obtained based on the following, pred C (i,j)=α·rec’ L (i,j)+β, where pred C (i,j) represents the predicted value of the chroma sample of the current chroma block, and rec’ L (i,j) represents the sample value of the corresponding luma sample of the downsampled luma block of the reconstructed luma block of the luma block.
[0038] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, when the luma reference sample is obtained based only on the adjacent samples on the left side of the luma block and the current chroma block is at the left boundary of the current coding tree unit (CTU), only one column of the reconstructed adjacent luma samples of the reconstructed luma block is used to obtain the luma reference sample.
[0039] In a possible embodiment of the apparatus according to the third or fourth aspect itself or any preceding embodiment of the third or fourth aspect, the linear model includes a multi-directional linear model (MDLM).
[0040] According to a fifth aspect, the present invention relates to a method for coding an intra chroma prediction mode in a bitstream of a video signal. The method includes performing an intra prediction of a chroma block of the video signal based on the intra chroma prediction mode, wherein the intra chroma prediction mode is selected from a first mode set, a second mode set including at least one of a CCLM_L mode or a CCLM_T mode, or a third mode set; and generating a bitstream of the video signal by including a syntax element indicating the intra chroma prediction mode, wherein the number of bits of the syntax element when the intra chroma prediction mode is selected from the first mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set, and the number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the third mode set.
[0041] The proposed method for encoding the intra chroma prediction mode enables representing CCLM_L and CCLM_T using a binary string and including it in the bitstream of the video signal.
[0042] In a possible embodiment of the method according to the fifth aspect itself, the first mode set includes at least one of a derived mode (DM) or a cross-component linear model (CCLM) prediction mode, and the third mode set includes at least one of a vertical mode, a horizontal mode, a DC mode, or a plane mode.
[0043] In a possible embodiment of the method according to the fifth aspect itself or any preceding embodiment of the fifth aspect, the syntax element for the DM mode is 0; the syntax element for the CCLM mode is 10; the syntax element for the CCLM_L mode is 1110; the syntax element for the CCLM_T mode is 1111; the syntax element for the planar mode is 11000; the syntax element for the vertical mode is 11001; the syntax element for the horizontal mode is 11010; the syntax element for the DC mode is 11011.
[0044] In a possible embodiment of the method according to the fifth aspect itself or any preceding embodiment of the fifth aspect, the syntax element for the DM mode is 00; the syntax element for the CCLM mode is 10; the syntax element for the CCLM_L mode is 110; the syntax element for the CCLM_T mode is 111; the syntax element for the planar mode is 0100; the syntax element for the vertical mode is 0101; the syntax element for the horizontal mode is 0110; the syntax element for the DC mode is 0111.
[0045] According to a sixth aspect, the present invention relates to a method for decoding an intra chroma prediction mode in a bitstream of a video signal. The method includes analyzing a plurality of syntax elements from the bitstream of the video signal; determining an intra chroma prediction mode based on the syntax elements from the plurality of syntax elements, wherein the intra chroma prediction mode is determined from one of a first mode set, a second mode set including at least one of a CCLM_L mode or a CCLM_T mode, or a third mode set, and the number of bits of the syntax elements when the intra chroma prediction mode is selected from the first mode set is smaller than the number of bits of the syntax elements when the intra chroma prediction mode is selected from the second mode set, and the number of bits of the syntax elements when the intra chroma prediction mode is selected from the second mode set is smaller than the number of bits of the syntax elements when the intra chroma prediction mode is selected from the third mode set; and performing an intra prediction of a current chroma block of the video signal based on the intra chroma prediction mode.
[0046] According to a seventh aspect, the present invention relates to an apparatus for encoding video data. The apparatus has a video data memory and a video encoder. The video encoder performs intra prediction of a chroma block of a video signal based on an intra chroma prediction mode, where the intra chroma prediction mode is selected from a first mode set, a second mode set including at least one of a CCLM_L mode or a CCLM_T mode, or a third mode set, and performs intra prediction; and generates a bitstream of the video signal by including a syntax element indicating the intra chroma prediction mode, where the number of bits of the syntax element when the intra chroma prediction mode is selected from the first mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set, and the number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the third mode set.
[0047] According to an eighth aspect, the present invention relates to an apparatus for decoding video data. The apparatus has a video data memory and a video decoder. The video decoder analyzes a plurality of syntax elements from a bitstream of a video signal, determines an intra chroma prediction mode based on the syntax elements from the plurality of syntax elements, where the intra chroma prediction mode is determined from one of a first mode set, a second mode set including at least one of a CCLM_L mode or a CCLM_T mode, or a third mode set, and the number of bits of the syntax elements when the intra chroma prediction mode is selected from the first mode set is smaller than the number of bits of the syntax elements when the intra chroma prediction mode is selected from the second mode set, and the number of bits of the syntax elements when the intra chroma prediction mode is selected from the second mode set is smaller than the number of bits of the syntax elements when the intra chroma prediction mode is selected from the third mode set, and performs an intra prediction of a current chroma block of the video signal based on the intra chroma prediction mode.
[0048] The proposed apparatus for encoding video data and the apparatus for decoding video data represent CCLM_L and CCLM_T using a binary string, include them in the bitstream of the video signal in the encoding apparatus, and enable decoding in the decoding apparatus. In possible embodiments of the apparatus according to the seventh and eighth aspects themselves, the first mode set includes at least one of a derived mode (DM) or a cross-component linear model (CCLM) prediction mode, and the third mode set includes at least one of a vertical mode, a horizontal mode, a DC mode, or a plane mode.
[0049] In a possible embodiment of the apparatus according to the seventh and eighth aspects themselves or any preceding embodiment of the seventh and eighth aspects, the syntax element of the DM mode is 0; the syntax element of the CCLM mode is 10; the syntax element of the CCLM_L mode is 1110; the syntax element of the CCLM_T mode is 1111; the syntax element of the planar mode is 11000; the syntax element of the vertical mode is 11001; the syntax element of the horizontal mode is 11010; the syntax element of the DC mode is 11011.
[0050] In a possible embodiment of the method according to the seventh and eighth aspects themselves or any preceding embodiment of the seventh and eighth aspects, the syntax element of the DM mode is 00; the syntax element of the CCLM mode is 10; the syntax element of the CCLM_L mode is 110; the syntax element of the CCLM_T mode is 111; the syntax element of the planar mode is 0100; the syntax element of the vertical mode is 0101; the syntax element of the horizontal mode is 0110; the syntax element of the DC mode is 0111.
[0051] According to a ninth aspect, the present invention relates to a method for performing intra prediction using a cross-component linear model (CCLM) prediction mode. The method includes determining a luma block corresponding to a current chroma block; obtaining luma reference samples of the luma block by downsampling adjacent samples of the luma block, wherein the luma reference samples include only luma reference samples obtained based on adjacent samples above the luma block or only luma reference samples obtained based on adjacent samples to the left of the luma block; determining a maximum luma value and a minimum luma value based on the luma reference samples; obtaining a first chroma value based at least in part on one or more positions of one or more luma reference samples associated with the maximum luma value; obtaining a second chroma value based at least in part on one or more positions of one or more luma reference samples associated with the minimum luma value; calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; and generating a prediction of the current chroma block based on the linear model coefficients and values of a downsampled version of the luma block.
[0052] In a possible embodiment of the method according to the ninth aspect itself, the number of luma reference samples is greater than or equal to the width of the current chroma block or greater than or equal to the height of the current chroma block.
[0053] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, the luma reference samples used to determine the maximum luma value and the minimum luma value are the available luma reference samples of the luma block.
[0054] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, the available luma reference samples of the luma block are determined based on the available chroma reference samples of the current chroma block.
[0055] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, up to 2*W luma reference samples are used to derive linear model coefficients when the luma reference samples are obtained based only on adjacent samples above the luma block, where W represents the width of the current chroma block.
[0056] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, up to 2*H luma reference samples are used to derive linear model coefficients when the luma reference samples are obtained based only on adjacent samples to the left of the luma block, where H represents the height of the current chroma block.
[0057] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, up to N available luma reference samples are used to derive linear model coefficients, where N is the sum of W and H, W represents the width of the current chroma block, and H represents the height of the current chroma block.
[0058] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, the linear model coefficients α and β are calculated based on the following: α=(y B -y A ) / (x B -x A )、 β=y A -αx A where x B represents the maximum luma value, y B represents the first chroma value, x A represents the minimum luma value, and y A represents the second chroma value.
[0059] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, the prediction of the current chroma block is obtained based on the following: pred C (i,j)=α·rec’L (i, j) + β, Here, pred C (i, j) represents the predicted value of the chroma sample of the current chroma block, and rec’ L (i, j) represents the sample value of the corresponding luma sample of the downsampled version of the reconstructed version of the luma block.
[0060] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, the downsampled version of the luma block is obtained by downsampling the reconstructed version of the luma block corresponding to the current chroma block.
[0061] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, when the luma reference sample is obtained based only on the adjacent samples above the luma block, and the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), or the current chroma block is at the upper boundary of the current coding tree unit (CTU), only one row of the reconstructed adjacent luma samples of the reconstructed version of the luma block is used to obtain the luma reference sample.
[0062] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, when the luma reference sample is obtained based only on the adjacent samples on the left side of the luma block and the current chroma block is at the left boundary of the current coding tree unit (CTU), only one column of the reconstructed adjacent luma samples of the reconstructed luma block is used to obtain the luma reference sample.
[0063] In a possible embodiment of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, the CCLM includes a multi-directional linear model (MDLM).
[0064] According to a tenth aspect, the present invention relates to an encoder configured to execute the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect.
[0065] According to an eleventh aspect, the present invention relates to a decoder configured to execute the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect.
[0066] According to an eleventh aspect, the present invention relates to an intra prediction method by using a cross-component linear prediction mode (CCLM). This method includes steps of: obtaining reference samples of a current luma block, where the reference samples only belong to an upper template of the current luma block; obtaining a maximum luma value and a minimum luma value based on the reference samples; obtaining a first chroma value based on a sample position of the maximum luma value; obtaining a second chroma value based on a sample position of the minimum luma value; calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; and obtaining a predictor of a current chroma block based on the linear model coefficients, where the current chroma block corresponds to the current luma block.
[0067] In a possible embodiment of the method according to the eleventh aspect itself, the number of reference samples is greater than or equal to the width of the current chroma block.
[0068] In a possible embodiment of the method according to the eleventh aspect itself or any preceding embodiment of the eleventh aspect, the reference samples are available.
[0069] In a possible embodiment of the method according to the eleventh aspect itself or any preceding embodiment of the eleventh aspect, the method further includes a step of checking the availability of reference samples within a range, where the length of the range is 2*W or the length of the range is the sum of W and H. Here, W represents the width of the current chroma block and H represents the height of the current chroma block.
[0070] In a possible embodiment of the method according to the 11th aspect itself or any preceding embodiment of the 11th aspect, a maximum of 2*W available reference samples are used to derive the linear model coefficients, where W represents the width of the current chroma block.
[0071] In a possible embodiment of the method according to the 11th aspect itself or any preceding embodiment of the 11th aspect, a maximum of N available reference samples are used to derive the linear model coefficients, where N is the sum of W and H, W represents the width of the current chroma block, and H represents the height of the current chroma block.
[0072] In a possible embodiment of the method according to the 11th aspect itself or any preceding embodiment of the 11th aspect, the linear model coefficients α and β are calculated based on the following: α=(y B -y A ) / (x B -x A ) β=y A -αx A where x B represents the maximum luma value, y B represents the first chroma value, x A represents the minimum luma value, and y A represents the second chroma value.
[0073] In a possible embodiment of the method according to the 11th aspect itself or any preceding embodiment of the 11th aspect, the predictor of the current chroma block is obtained based on the following: pred C (i,j)=α·rec L ’(i,j)+β where pred C (i,j) represents a chroma sample, and rec L (i,j) represents the corresponding reconstructed luma sample.
[0074] In a possible embodiment of the method according to the 11th aspect itself or any preceding embodiment of the 11th aspect, the number of reference samples is greater than or equal to the size of the current chroma block.
[0075] In a possible embodiment of the method according to the 11th aspect itself or any preceding embodiment of the 11th aspect, the reference samples are downsampled luma samples.
[0076] In a possible embodiment of the method according to the 11th aspect itself or any preceding embodiment of the 11th aspect, when the current chroma block is at the upper boundary, only one row of the reconstructed adjacent luma samples is used to obtain the reference samples.
[0077] In a possible embodiment of the method according to the 11th aspect itself or any preceding embodiment of the 11th aspect, the CCLM is a multi-directional linear model (MDLM), and the linear model coefficients are used to obtain the MDLM.
[0078] In a possible embodiment of the method according to the 11th aspect itself or any preceding embodiment of the 11th aspect, this method is called CCIP_T.
[0079] According to the 12th aspect, the present invention relates to an intra prediction method by using a cross-component linear prediction mode (CCLM). This method includes obtaining reference samples of the current luma block, wherein the reference samples belong only to the left template of the current luma block; obtaining a maximum luma value and a minimum luma value based on the reference samples; obtaining a first chroma value based on the sample position of the maximum luma value; obtaining a second chroma value based on the sample position of the minimum luma value; calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; Obtaining a predictor for a current chroma block based on linear model coefficients; The current chroma block corresponds to the current luma block.
[0080] In a possible embodiment of the method according to the twelfth aspect itself, the number of reference samples is greater than or equal to the height of the current chroma block.
[0081] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, the reference samples are available.
[0082] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, the method further includes checking the availability of reference samples within a range, where the length of the range is 2*H, or the length of the range is the sum of W and H. Here, W represents the width of the current chroma block, and H represents the height of the current chroma block.
[0083] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, a maximum of 2*H available reference samples are used to derive the linear model coefficients, where H represents the height of the current chroma block.
[0084] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, a maximum of N available reference samples are used to derive the linear model coefficients, where N is the sum of W and H, W represents the width of the current chroma block, and H represents the height of the current chroma block.
[0085] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, the linear model coefficients α and β are calculated based on the following: α=(y B -y A ) / (x B -x A ) β=y A -αxA Here, x B represents the maximum luma value, and y B represents the first chroma value, and x A represents the minimum luma value, and y A represents the second chroma value.
[0086] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, the predictor of the current chroma block is obtained based on the following, pred C (i,j)=α·rec L ’(i,j)+β Here, pred C (i,j) represents a chroma sample, and rec L (i,j) represents the corresponding reconstructed luma sample.
[0087] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, the number of reference samples is greater than or equal to the size of the current chroma block.
[0088] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, the reference samples are downsampled luma samples.
[0089] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, when the current block of the current chroma block is at the left boundary, only one column of the reconstructed adjacent luma samples is used to obtain the reference samples.
[0090] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, the CCLM is a multi-directional linear model (MDLM), and the linear model coefficients are used to obtain the MDLM.
[0091] In a possible embodiment of the method according to the twelfth aspect itself or any preceding embodiment of the twelfth aspect, this method is called CCIP_L.
[0092] According to a thirteenth aspect, the present invention relates to an intra prediction method by using a cross-component linear prediction mode (CCLM). This method includes obtaining reference samples of the current luma block, where the reference samples belong only to the upper template of the current luma block or only to the left template of the current luma block; obtaining chroma samples of the current chroma block, where the current chroma block corresponds to the current luma block; calculating linear model coefficients based on the reference samples and the chroma samples; obtaining a predictor of the current chroma block based on the linear model coefficients.
[0093] In a possible embodiment of the method according to the thirteenth aspect itself, a maximum of N reference samples are used to derive the linear model coefficients, where N is the sum of W and H, W represents the width of the current chroma block, and H represents the height of the current chroma block.
[0094] In a possible embodiment of the method according to the thirteenth aspect itself or any preceding embodiment of the thirteenth aspect, when the reference samples belong only to the upper template of the current luma block, the number of reference samples is greater than or equal to the width of the current chroma block.
[0095] In a possible embodiment of the method according to the thirteenth aspect itself or any preceding embodiment of the thirteenth aspect, when the reference samples belong only to the upper template of the current luma block, a maximum of 2*W reference samples are used to derive the linear model coefficients, where W represents the width of the current chroma block.
[0096] In a possible embodiment of the method according to the 13th aspect itself or any preceding embodiment of the 13th aspect, when the reference sample belongs only to the left template of the current luma block, the number of reference samples is greater than or equal to the height of the current luma block.
[0097] In a possible embodiment of the method according to the 13th aspect itself or any preceding embodiment of the 13th aspect, when the reference sample belongs only to the left template of the current luma block, a maximum of 2*H reference samples are used to derive the linear model coefficients, where W represents the height of the current chroma block.
[0098] In a possible embodiment of the method according to the 13th aspect itself or any preceding embodiment of the 13th aspect, the number of reference samples is greater than or equal to the size of the current chroma block.
[0099] In a possible embodiment of the method according to the 13th aspect itself or any preceding embodiment of the 13th aspect, the reference sample is a downsampled luma sample.
[0100] In a possible embodiment of the method according to the 13th aspect itself or any preceding embodiment of the 13th aspect, when the reference sample belongs only to the upper template of the current luma block and the current block of the current chroma block is at the upper boundary, only one row of the reconstructed adjacent luma samples is used to obtain the reference sample.
[0101] In a possible embodiment of the method according to the 13th aspect itself or any preceding embodiment of the 13th aspect, when the reference sample belongs only to the left template of the current luma block and the current block of the current chroma block is at the left boundary, only one column of the reconstructed adjacent luma samples is used to obtain the reference sample.
[0102] In a possible embodiment of the method according to the 13th aspect itself or any preceding embodiment of the 13th aspect, the CCLM is a multi-directional linear model (MDLM), and the linear model coefficients are used to obtain the MDLM.
[0103] In a possible embodiment of the method according to the 13th aspect itself or any preceding embodiment of the 13th aspect, a reference sample is available.
[0104] According to a 14th aspect, the present invention relates to a decoder including a processing circuit for executing a method according to any one of the 11th aspect itself or the preceding embodiments of the 11th aspect.
[0105] According to a 15th aspect, the present invention relates to a decoder including a processing circuit for executing a method according to any one of the 12th aspect itself or the preceding embodiments of the 12th aspect.
[0106] According to a 16th aspect, the present invention relates to a decoder including a processing circuit for executing a method according to any one of the 13th aspect itself or the preceding embodiments of the 13th aspect.
[0107] According to a 17th aspect, the present invention relates to an intra prediction method by using a cross-component linear prediction mode (CCLM). This method obtaining reference samples of a current luma block, wherein the reference samples belong only to an upper template of the current luma block; obtaining a maximum luma value and a minimum luma value based on the reference samples; obtaining a first chroma value and a second chroma value based on the maximum luma value and the minimum luma value; calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; obtaining a predictor of a current block based on the linear model coefficients.
[0108] In a possible embodiment of the method according to aspect 17 itself, the number of reference samples is greater than or equal to the width of the current chroma block.
[0109] In a possible embodiment of the method according to aspect 17 itself or any preceding embodiment of aspect 17, the reference samples are available.
[0110] In a possible embodiment of the method according to aspect 17 itself or any preceding embodiment of aspect 17, a maximum of 2*W reference samples are used to derive the model coefficients.
[0111] In a possible embodiment of the method according to aspect 17 itself or any preceding embodiment of aspect 17, this method is called CCIP_T.
[0112] According to aspect 18, the present invention relates to an intra prediction method by using a cross-component linear prediction mode (CCLM). This method includes obtaining reference samples of the current luma block, where the reference samples belong only to the left template of the current luma block; obtaining a maximum luma value and a minimum luma value based on the reference samples; obtaining a first chroma value and a second chroma value based on the maximum luma value and the minimum luma value; calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; obtaining a predictor of the current block based on the linear model coefficients.
[0113] In a possible embodiment of the method according to aspect 18 itself, the number of reference samples is greater than or equal to the height of the current chroma block.
[0114] In a possible embodiment of the method according to aspect 18 itself or any preceding embodiment of aspect 18, the reference samples are available.
[0115] In a possible embodiment of the method according to the 18th aspect itself or any preceding embodiment of the 18th aspect, a maximum of 2*H reference samples are used to derive the model coefficients.
[0116] In a possible embodiment of the method according to the 18th aspect itself or any preceding embodiment of the 18th aspect, this method is called CCIP_L.
[0117] According to the 19th aspect, the present invention relates to a decoder for executing the method according to the 17th aspect itself or any preceding embodiment of the 17th aspect.
[0118] According to the 20th aspect, the present invention relates to a decoder for executing the method according to the 18th aspect itself or any preceding embodiment of the 18th aspect.
[0119] According to the 21st aspect, the present invention relates to a method for intra prediction using a linear model. This method includes the steps of: obtaining reference samples of the current luma block; obtaining a maximum luma value and a minimum luma value based on the reference samples; obtaining a first chroma value and a second chroma value based on the positions of the luma samples having the maximum luma value and the positions of the luma samples having the minimum luma value; calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; and obtaining a predictor of the current block based on the linear model coefficients. The step of obtaining reference samples of the current luma block includes the step of determining L available chroma template samples of the current luma block, where the reference samples of the current luma block are L luma template samples corresponding to the L available chroma template samples, or the step of determining L available adjacent chroma samples of the current chroma block, where the reference samples of the current luma block are L adjacent luma samples corresponding to the L available adjacent chroma samples. L>=1, and L is a positive integer.
[0120] In a possible embodiment of the method according to the 21st aspect itself, the step of determining the L available chroma template samples of the current chroma block includes the step of checking the availability of the adjacent chroma samples above the current chroma block. When the L upper adjacent chroma samples are available, the reference sample of the current luma block is the L adjacent luma samples corresponding to the L upper adjacent chroma samples, where L >= 1 and L <= W2, W2 indicates the upper template sample range, and L and W2 are positive integers.
[0121] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, the step of determining the L available chroma template samples of the current chroma block includes the step of checking the availability of the adjacent chroma samples to the left of the current chroma block. When the L left adjacent chroma samples are available, the reference sample of the current luma block is the L adjacent luma samples corresponding to the L left adjacent chroma samples, where L >= 1 and L <= H2, H2 indicates the left template sample range, and L and H2 are positive integers.
[0122] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, the step of determining L available chroma template samples of the current chroma block includes checking the availability of adjacent chroma samples above the current chroma block and the availability of adjacent chroma samples to the left of the current chroma block. When L1 adjacent chroma samples above are available and L2 adjacent chroma samples to the left are available, the reference sample of the current luma block includes or consists of L1 adjacent luma samples corresponding to the L1 adjacent chroma samples above and L2 adjacent luma samples corresponding to the L2 adjacent chroma samples to the left. L2 >= 1 and L2 <= H2, where H2 indicates the left template sample range, and L and H2 are positive integers. L1 >= 1 and L1 <= W2, where W2 indicates the upper template sample range, and L1 and W2 are positive integers. L = L1 + L2.
[0123] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, when L adjacent chroma samples are available within the template range, then the L template luma samples and the L template chroma samples are used to obtain model coefficients.
[0124] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, the reference sample is available.
[0125] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, the linear model coefficients α and β are calculated based on the following: α=(y B -y A ) / (x B -x A )、 β=y A -αx A Here, x B represents the maximum luma value, and y Brepresents the first chroma value, x A represents the minimum luma value, y A represents the second chroma value.
[0126] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, the predictor of the current chroma block is obtained based on the following, pred C (i,j)=α·rec L ’(i,j)+β, where pred C (i,j) represents a chroma sample, and rec L (i,j) represents the corresponding reconstructed luma sample.
[0127] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, the number of reference samples is greater than or equal to the size of the current luma block.
[0128] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, the reference samples are downsampled luma samples.
[0129] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, when the current block of the current chroma block is at the upper boundary, only one row of the reconstructed adjacent luma samples is used to obtain the reference samples.
[0130] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, the linear model is a multi-directional linear model (MDLM), and the linear model coefficients are used to obtain the MDLM.
[0131] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, this method is called CCIP_T, or this method is called CCIP_L.
[0132] In a possible embodiment of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, the reference sample belongs only to the upper template of the current luma block, or belongs only to the left template of the current luma block, or the reference sample belongs to the upper template of the current luma block and the left template of the current luma block.
[0133] According to the 22nd aspect, the present invention relates to a decoder including a processing circuit for executing the method according to the 21st aspect itself or a preceding embodiment of the 21st aspect.
[0134] According to the 23rd aspect, the present invention relates to an encoder including a processing circuit for executing the method according to the 21st aspect itself or a preceding embodiment of the 21st aspect.
[0135] According to the 24th aspect, the present invention relates to a method for binarizing a chroma mode. This method includes performing intra prediction using a linear model (multi-directional linear model, MDLM, etc.); and generating a bitstream including a plurality of syntax elements, the plurality of syntax elements indicating or including the CCLM mode, the CCIP_L mode, or the CCIP_T mode.
[0136] In a possible embodiment of the method according to the 24th aspect itself, the first indicator (77) indicates the CCLM mode and the intra_chroma_pred_mode index is 4, the second indicator (78) indicates the CCIP_L mode and the intra_chroma_pred_mode index is 5, the third indicator (79) indicates the CCIP_T mode and the intra_chroma_pred_mode index is 6.
[0137] In a possible embodiment of the method according to the 24th aspect itself or any preceding embodiment of the 24th aspect, when sps_cclm_enabled_flag is equal to 1, IntraPredModeC[xCb][yCb] depends on intra_chroma_pred_mode[xCb][yCb] and IntraPredModeY[xCb][yCb].
[0138] According to the 25th aspect, the present invention relates to a decoding method performed by a decoding apparatus. This method includes: analyzing a plurality of syntax elements from a bitstream, the plurality of syntax elements indicating or including a CCLM mode, a CCIP_L mode, or a CCIP_T mode; performing intra prediction using the indicated linear model.
[0139] In a possible embodiment of the method according to the 25th aspect itself, the first indicator (77) indicates a CCLM mode and the intra_chroma_pred_mode index is 4, the second indicator (78) indicates a CCIP_L mode and the intra_chroma_pred_mode index is 5, the third indicator (79) indicates a CCIP_T mode and the intra_chroma_pred_mode index is 6.
[0140] In a possible embodiment of the method according to the 25th aspect itself or any preceding embodiment of the 25th aspect, when sps_cclm_enabled_flag is equal to 1, IntraPredModeC[xCb][yCb] depends on intra_chroma_pred_mode[xCb][yCb] and IntraPredModeY[xCb][yCb].
[0141] According to a 26th aspect, the present invention relates to a decoder including a processing circuit for performing a method according to the 24th and 25th aspects themselves or any preceding embodiment of the 24th and 25th aspects.
[0142] According to a 27th aspect, the present invention relates to an encoder including a processing circuit for performing a method according to the 24th and 25th aspects themselves or any preceding embodiment of the 24th and 25th aspects.
[0143] According to a 28th aspect, the present invention relates to a computer-readable medium storing instructions which, when executed on a processor, cause the processor to perform a method according to the 24th and 25th aspects themselves or any preceding embodiment of the 24th and 25th aspects.
[0144] According to a 28th aspect, the present invention relates to a decoder. The decoder includes one or more processors and a non-transitory computer-readable storage medium coupled to and storing programming executed by the processor, the programming configuring the decoder to perform a method according to the 24th and 25th aspects themselves or any preceding embodiment of the 24th and 25th aspects when executed by the processor.
[0145] According to a 28th aspect, the present invention relates to an encoder. The encoder includes one or more processors and a non-transitory computer-readable storage medium coupled to and storing programming executed by the processor, the programming configuring the encoder to perform a method according to the 24th and 25th aspects themselves or any preceding embodiment of the 24th and 25th aspects when executed by the processor.
[0146] According to a 29th aspect, the present invention relates to an intra prediction method by using a cross-component linear prediction mode (CCLM). The method Obtaining a reference sample of the current luma block; Obtaining a maximum luma value and a minimum luma value based on the reference sample; Obtaining a first chroma value and a second chroma value based on the maximum luma value and the minimum luma value; Calculating a linear model coefficient based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; Obtaining a predictor of the current block based on the linear model coefficient; including The availability of the template sample is to check adjacent chroma samples.
[0147] According to a 30th aspect, the present invention relates to a decoder for executing the method according to the 28th aspect itself.
[0148] According to a 31st aspect, the present invention relates to a decoder for executing the method according to the 29th aspect itself.
[0149] According to a 32nd aspect, the present invention relates to a decoder for executing the method according to the 28th aspect or the 29th aspect.
[0150] According to a 33rd aspect, there is provided a device including a module / unit / component / circuit for executing at least a part of the steps of the above method according to any preceding aspect itself or any preceding embodiment of any preceding aspect.
[0151] The device according to the 33rd aspect can be extended to an embodiment corresponding to the embodiment of the method according to any preceding aspect. Therefore, the embodiment of the device includes the features of the corresponding embodiment of the method according to any preceding aspect.
[0152] The advantages of the device according to any preceding aspect are the same as the advantages of the corresponding embodiments of the method according to any preceding aspect.
[0153] For clarity, any one of the foregoing examples can be combined with any one or more of the other foregoing examples to create new examples within the scope of the present disclosure.
[0154] These and other features will be more clearly understood from the following detailed description, which is to be construed in conjunction with the accompanying drawings and the claims.
Brief Description of the Drawings
[0155] For a more complete understanding of the present disclosure, reference is made to the following brief description, which is to be construed in conjunction with the accompanying drawings and detailed description, and in which like reference numerals represent like parts.
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Figure 6D
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
DETAILED DESCRIPTION OF THE INVENTION
[0156] Exemplary embodiments of one or more examples are provided below, but it should first be understood that the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or existing. The present disclosure should in no way be limited to the exemplary embodiments, drawings, and techniques shown and described herein, including the exemplary designs and embodiments shown herein, but may be varied within the scope of the appended claims, together with all equivalents of their full scope.
[0157] FIG. 1A is a block diagram showing an exemplary coding system 10 that can utilize bidirectional prediction techniques. As shown in FIG. 1A, the coding system 10 includes a source device 12 that provides encoded video data, which is later decoded by a destination device 14. In particular, the source device 12 can provide the video data to the destination device 14 via a computer-readable medium 16. The source device 12 and the destination device 14 can include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some cases, the source device 12 and the destination device 14 can be equipped for wireless communication.
[0158] The destination device 14 can receive the encoded video data via the computer-readable medium 16 and decode the encoded video data. The computer-readable medium 16 can include any type of medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the computer-readable medium 16 can include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data can be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium can include any wireless or wired communication medium such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, a switch, a base station, or any other facility that can help facilitate communication from the source device 12 to the destination device 14.
[0159] In some examples, the encoded data may be output from the output interface 22 to a storage device. Similarly, the encoded data may be accessed from the storage device by an input interface. The storage device may include any of a variety of distributed or local access type data storage media, such as a hard drive, a Blu-ray disk, a digital video disk (DVD), a compact disk read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or other suitable digital storage media for storing encoded video data. In a further example, the storage device may correspond to a file server or another intermediate storage device that may store the encoded video generated by the source device 12. The destination device 14 may access the video data stored in the storage device via streaming or downloading. The file server may be any type of server that can store the encoded video data and transmit the encoded video data to the destination device 14. Examples of file servers include web servers (e.g., for a website), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. The destination device 14 can access the encoded video data via any standard data connection including an Internet connection. This may include a wireless channel (e.g., Wi-Fi connection) suitable for accessing the encoded video data stored on the file server, a wired connection (digital subscriber line (DSL), cable modem, etc.), or a combination of both. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.
[0160] The technology of the present disclosure is not necessarily limited to wireless applications or settings. This technology can be applied to video coding when supporting various multimedia applications, such as terrestrial television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions such as dynamic adaptive streaming via HTTP (DASH), digital video encoded on data storage media, decoding of digital video stored on data storage media, or other applications. In some examples, the coding system 10 can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or videotelephony.
[0161] In the example of FIG. 1A, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The destination device 14 includes an input interface 28, a video decoder 300, and a display device 32. In accordance with the present disclosure, the video encoder 200 of the source device 12 and / or the video decoder 300 of the destination device 14 can be configured to apply techniques for bidirectional prediction. In other examples, the source device and the destination device can include other components or configurations. For example, the source device 12 can receive video data from an external video source such as an external camera. Similarly, the destination device 14 can interface with an external display device instead of including an integrated display device.
[0162] The coding system 10 shown in FIG. 1A is merely an example. The techniques for bidirectional prediction can be performed by any digital video encoding and / or decoding device. The technology of the present disclosure is generally performed by a video coding device, but the technology can also be typically performed by a video encoder / decoder, which is generally called a "CODEC". Further, the technology of the present disclosure can also be performed by a video preprocessor. The video encoder and / or decoder can be a graphics processing unit (GPU) or a similar device.
[0163] The source device 12 and the destination device 14 are merely examples of such a coding device that generates coded video data for the source device 12 to transmit to the destination device 14. In some examples, the source device 12 and the destination device 14 may operate in a substantially symmetric manner such that each of the source and destination devices 12, 14 includes video encoding and decoding components. Thus, the coding system 10 can support one-way or two-way video transmission between the video devices 12, 14, for example, for video streaming, video playback, video broadcasting, or videotelephony.
[0164] The video source 18 of the source device 12 may include a video capture device such as a video camera, a video archive including previously captured video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source 18 can generate computer graphics-based data as source video or as a combination of live video, archived video, and computer-generated video.
[0165] In some cases, when the video source 18 is a video camera, the source device 12 and the destination device 14 may form a so-called camera-equipped mobile phone or videotelephone. However, as described above, the techniques described in this disclosure may generally be applicable to video coding and may be applied to wireless and / or wired applications. In any case, the captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. Next, the encoded video information may be output onto the computer-readable medium 16 by the output interface 22.
[0166] The computer-readable medium 16 can include a temporary medium such as a wireless broadcast or a wired network transmission, or a storage medium (i.e., a non-temporary storage medium) such as a hard disk, a flash drive, a compact disc, a digital video disc, a Blu-ray disc, or other computer-readable media. In some examples, a network server (not shown) can receive the encoded video data from the source device 12 and provide the encoded video data to the destination device 14 via, for example, a network transmission. Similarly, a computing device of a media production facility such as a disk stamping facility can receive the encoded video data from the source device 12 and manufacture a disk including the encoded video data. Thus, the computer-readable medium 16 can be understood to include one or more computer-readable media in various forms in various examples.
[0167] The input interface 28 of the destination device 14 receives information from the computer-readable medium 16. The information of the computer-readable medium 16 can include syntax information defined by the video encoder 20, which is also used by the video decoder 30. The syntax information includes syntax elements that describe the characteristics and / or processing of blocks and other coded units, such as groups of pictures (GOPs). The display device 32 displays the decoded video data to the user and can include any of a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0168] The video encoder 200 and the video decoder 300 can operate in accordance with video coding standards such as the currently under - development High - Efficiency Video Coding (HEVC) standard, and can comply with the HEVC Test Model (HM). Alternatively, the video encoder 200 and the video decoder 300 can operate in accordance with other corporate or industry standards, such as the International Telecommunication Union Telecommunication Standardization Sector (ITU - T) H.264 standard, or MPEG (Motion Picture Expert Group) - 4, Part 10, also known as Advanced Video Coding (AVC), H.265 / HEVC, or extensions of such standards. However, the technology of the present disclosure is not limited to a specific coding standard. Other examples of video coding standards include MPEG - 2 and ITU - T H.263. Although not shown in FIG. 1A, in some embodiments, the video encoder 200 and the video decoder 300 can be integrated with an audio encoder and decoder, respectively, and include an appropriate multiplexer - demultiplexer (MUX - DEMUX) unit, or other hardware and software, to process the encoding of both audio and video in a common data stream or separate data streams. When applicable, the MUX - DEMUX unit can comply with other protocols such as the ITU H.223 multiplexer protocol, or the User Datagram Protocol (UDP).
[0169] The video encoder 200 and the video decoder 300 can each be implemented as any of various suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented partially in software, the apparatus can store software instructions in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, and any of these can be integrated as part of an integrated encoder / decoder (CODEC) within the respective apparatus. An apparatus including the video encoder 200 and / or the video decoder 300 can include an integrated circuit, a microprocessor, and / or a wireless communication device such as a mobile phone.
[0170] FIG. 1B is an exemplary diagram of an exemplary video coding system 40 including the encoder 200 of FIG. 2 and / or the decoder 300 of FIG. 3 according to an exemplary embodiment. The system 40 can perform the technology of this application, such as merge estimation in inter prediction. In the illustrated embodiment, the video coding system 40 can include an imaging device 41, a video encoder 20, a video decoder 300 (and / or a video codec implemented via the logic circuit 47 of the processing unit 46), an antenna 42, one or more processors 43, one or more memory stores 44, and / or a display device 45.
[0171] As shown, imaging device 41, antenna 42, processing unit 46, logic circuit 47, video encoder 20, video decoder 30, processor 43, memory store 44, and / or display device 45 may communicate with each other. As described, both video encoder 200 and video decoder 30(0) are shown, but video coding system 40 may include only video encoder 200 or only video decoder 300 in various actual scenarios.
[0172] As shown, in some examples, video coding system 40 may include antenna 42. Antenna 42 may be configured, for example, to transmit or receive an encoded bitstream of video data. Further, in some examples, video coding system 40 may include display device 45. Display device 45 may be configured to present video data. As shown, in some examples, logic circuit 47 may be implemented via processing unit 46. Processing unit 46 may include, for example, application specific integrated circuit (ASIC) logic, a graphics processor, a general purpose processor, etc. Video coding system 40 may also include optional processor 43, which may similarly include application specific integrated circuit (ASIC) logic, a graphics processor, a general purpose processor, etc. In some examples, logic circuit 47 may be implemented via hardware, video coding dedicated hardware, etc., and processor 43 may be implemented via general purpose software, an operating system, etc. Further, memory store 44 may be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash, etc.). In a non-limiting example, memory store 44 may be implemented by cache memory. In some examples, logic circuit 47 may be able to access memory store 44 (e.g., for implementation of an image buffer). In other examples, logic circuit 47 and / or processing unit 46 may include a memory store (e.g., a cache, etc.) for implementation of an image buffer, etc.
[0173] In some examples, the video encoder 200 implemented via a logic circuit may include an image buffer (e.g., via either the processing unit 46 or the memory store 44), and a graphics processing unit (e.g., via the processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include a video encoder 200 implemented via the logic circuit 47 to embody various modules as described with respect to FIG. 2 and / or any other encoder system or subsystem described herein. The logic circuit may be configured to perform various operations as discussed herein.
[0174] The video decoder 300 can be implemented in a similar manner as implemented via the logic circuit 47 to embody various modules as described with respect to the decoder 300 of FIG. 3 and / or any other decoder system or subsystem described herein. In some examples, the video decoder 300 may be implemented via a logic circuit and may include an image buffer (e.g., via either the processing unit 46 or the memory store 44), and a graphics processing unit (e.g., via the processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include a video decoder 300 implemented via the logic circuit 47 to embody various modules as described with respect to FIG. 3 and / or any other decoder system or subsystem described herein.
[0175] In some examples, the antenna 42 of the video coding system 40 may be configured to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data associated with coding partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators (as discussed), and / or data defining the coding partitions), data, indicators, index values, mode selection data, etc., related to encoding a video frame as discussed herein. The video coding system 40 may also include a video decoder 300 coupled to the antenna 42 and configured to decode the encoded bitstream. The display device 45 is configured to present the video frame.
[0176] FIG. 2 is a block diagram showing an example of a video encoder 200 that can implement the technology of the present application. The video encoder 200 can perform intra-coding and inter-coding of video blocks within a video slice. Intra-coding reduces or eliminates the spatial redundancy of video within a given video frame or image, depending on spatial prediction. Inter-coding reduces or eliminates the temporal redundancy of video in adjacent frames or images of a video sequence, depending on temporal prediction. The intra mode (I mode) may refer to any of several spatial-based coding modes. Inter modes, such as unidirectional prediction (P mode) or bidirectional prediction (B mode), may refer to any of several time-based coding modes.
[0177] FIG. 2 shows a schematic / conceptual block diagram of an exemplary video encoder 200 configured to implement the techniques of the present disclosure. In the example of FIG. 2, the video encoder 200 includes a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 210, a buffer 216, a loop filter unit 220, a decoded picture buffer (DPB) 230, a prediction processing unit 260, and an entropy encoding unit 270. The prediction processing unit 260 may include an inter estimation 242, an inter prediction unit 244, an intra estimation 252, an intra prediction unit 254, and a mode selection unit 262. The inter prediction unit 244 may further include a motion compensation unit (not shown). The video encoder 200 as shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder by a hybrid video codec.
[0178] For example, the residual calculation unit 204, the transform processing unit 206, the quantization unit 208, the prediction processing unit 260, and the entropy encoding unit 270 form the forward signal path of the encoder 200, while for example, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the prediction processing unit 260 form the reverse signal path of the encoder, where the reverse signal path of the encoder corresponds to the signal path of a decoder (see decoder 300 in FIG. 3).
[0179] The encoder 200 is configured to receive an image of a series of images forming, for example, a video or a video sequence, by, for example, an input 202, an image 201, or a block 203 of the image 201. The image block 203 may also be referred to as a current image block or an image block to be coded, and the image 201 may be referred to as a current image or an image to be coded (especially in video coding to distinguish the current image from other images (e.g., images previously encoded and / or decoded in the same video sequence (i.e., the video sequence including the current image))).
[0180] Partitioning
[0181] An embodiment of the encoder 200 may include a partitioning unit (not shown in FIG. 2) configured to partition an image 201 into a plurality of blocks, such as blocks like block 203, typically a plurality of non-overlapping blocks. The partitioning unit may use the same block size and corresponding grid defining the block size for all images of a video sequence, or may vary the block size between an image or a subset or group of images and be configured to partition each image into corresponding blocks.
[0182] In HEVC and other video coding specifications, a set of coding tree units (CTUs) may be generated to produce an encoded representation of an image. Each CTU may include a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples, and a syntax structure used to code the samples of the coding tree block. In a monochrome image or an image having three separate color planes, a CTU may include a single coding tree block and a syntax structure used to code the samples of the coding tree block. The coding tree block may be an N×N block of samples. A CTU may also be referred to as a "tree block" or "largest coding unit" (LCU). The CTUs of HEVC may be generally similar to the macroblocks of other standards such as H.264 / AVC. However, a CTU is not necessarily limited to a particular size and may include one or more coding units (CUs). A slice may include an integral number of CTUs ordered consecutively in raster scan order.
[0183] In HEVC, a CTU is divided into CUs using a quadtree structure shown as a coding tree to adapt to various local characteristics. The decision on whether to code an image region using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. A CU can include a coding block of luma samples, two corresponding coding blocks of chroma samples of an image having an array of luma samples, an array of Cb samples, and an array of Cr samples, and a syntax structure used to code the samples of the coding block. In a monochrome image or an image having three separate color planes, a CU can include a single coding block and a syntax structure used to code the samples of the coding block. A coding block is an N×N block of samples. In some examples, a CU can be the same size as a CTU. Each CU is coded in one coding mode, which can be, for example, an intra coding mode or an inter coding mode. Other coding modes are also possible. The encoder 200 receives video data. The encoder 200 can encode each CTU within a slice of an image of the video data. As part of the encoding of the CTU, the prediction processing unit 260 of the encoder 200 or another processing unit (including, but not limited to, the units of the encoder 200 shown in FIG. 2) can perform partitioning to divide the CTB of the CTU into successively smaller blocks 203. The small blocks can be the coding blocks of the CUs.
[0184] The syntax data in the bitstream can also define the size of the CTU. A slice contains a number of consecutive CTUs in coding order. A video frame or image or picture can be partitioned into one or more slices. As described above, each tree block can be divided into coding units (CUs) according to a quadtree. Generally, a quadtree data structure contains one node for each CU, and the root node corresponds to a tree block (e.g., a CTU). When a CU is divided into four sub-CUs, the node corresponding to the CU contains four child nodes, and each child node corresponds to one of the sub-CUs. The multiple nodes of the quadtree structure include leaf nodes and non-leaf nodes. A leaf node has no child nodes in the tree structure (i.e., the leaf node is not further divided). Non-leaf nodes include the root node of the tree structure. For each non-root node of the multiple nodes, each non-root node corresponds to a sub-CU of the CU corresponding to the parent node in the tree structure of each non-root node. Each non-leaf node has one or more child nodes in the tree structure.
[0185] Each node of the quadtree data structure can provide syntax data for the corresponding CU. For example, a node in the quadtree can include a split flag indicating whether the CU corresponding to the node is divided into sub-CUs. The syntax elements of a CU can be defined recursively and may depend on whether the CU is divided into sub-CUs. When a CU is not further divided, it is called a leaf CU. When the block of a CU is further divided, it can generally be called a non-leaf CU. Each level of partitioning is a quadtree divided into four sub-CUs. The black CU is an example of a leaf node (i.e., a block that is not further divided).
[0186] The CU has a similar purpose as the macroblock in the H.264 standard, except that the CU has no size distinction. For example, a tree block can be divided into four child nodes (also called sub-CUs), and then each child node can be used as a parent node and divided into another four child nodes. The last undivided child node, called the leaf node of the quadtree, contains a coding node also called a leaf CU. The syntax data associated with the coded bitstream can define the maximum number of times the tree block is divided (referred to as the maximum CU depth) and can also define the minimum size of the coding node. Therefore, the bitstream can also define the smallest coding unit (SCU). The term "block" is used to refer to any of the CU, PU, or TU in the context of HEVC, or a similar data structure in the context of other standards (e.g., the macroblock and its sub-blocks in H.264 / AVC).
[0187] In HEVC, each CU can be further divided into one, two, or four PUs according to the PU partition type. Within one PU, the same prediction process is applied, and the relevant information is sent to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree of the CU. One of the important features of the HEVC structure is the concept of multiple partitions including CUs, PUs, and TUs. The PU can be partitioned so that its shape becomes non-square. The syntax data related to the CU can also describe, for example, partitioning the CU into one or more PUs. The TU can have a square or non-square (e.g., rectangular) shape, and the syntax data related to the CU can describe, for example, partitioning the CU into one or more TUs according to a quadtree. The partition mode may vary depending on whether the CU is coded in skip or direct mode, intra prediction mode, or inter prediction mode.
[0188] VVC (Versatile Video Coding) eliminates the separation of the concepts of PUs and TUs while supporting higher flexibility in CU partition shapes. The size of a CU corresponds to the size of the coding node and can be square or non-square (e.g., rectangular) in shape. The size of a CU can range from 4×4 pixels (or 8×8 pixels) up to the size of a tree block of 128×128 pixels or more (e.g., 256×256 pixels).
[0189] After the encoder 200 generates the prediction blocks of a CU (e.g., luma, Cb, and Cr prediction blocks), the encoder 200 can generate the residual block of the CU. For example, the encoder 100 can generate the luma residual block of the CU. Each sample in the luma residual block of the CU indicates the difference between the luma sample in the predicted luma block of the CU and the corresponding sample in the original luma coding block of the CU. Further, the encoder 200 can generate the Cb residual block of the CU. Each sample in the Cb residual block of the CU can indicate the difference between the Cb sample in the predicted Cb block of the CU and the corresponding sample in the original Cb coding block of the CU. The encoder 100 can also generate the Cr residual block of the CU. Each sample in the Cr residual block of the CU can indicate the difference between the Cr sample in the predicted Cr block of the CU and the corresponding sample in the original Cr coding block of the CU.
[0190] In some examples, the encoder 100 skips the application of the transform to the transform block. In such examples, the encoder 200 can process the residual sample values in the same way as the transform coefficients. Thus, in examples where the encoder 100 skips the application of the transform, the following description of the transform coefficients and the coefficient block can be applicable to the transform block of the residual samples.
[0191] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), encoder 200 can quantize the coefficient block to potentially reduce the amount of data used to represent the coefficient block and provide further compression. Quantization generally refers to the process of compressing a range of values into a single value. After encoder 200 quantizes the coefficient block, encoder 200 can entropy code the syntax elements indicating the quantized transform coefficients. For example, encoder 200 can perform context-adaptive binary arithmetic coding (CABAC) or other entropy coding techniques on the syntax elements indicating the quantized transform coefficients.
[0192] Encoder 200 can output a bitstream of encoded image data 271 that includes a sequence of bits forming a representation of the coded image and associated data. Thus, the bitstream includes an encoded representation of the video data.
[0193] In “Block partitioning structure for next generation video coding” by J. An et al., International Telecommunication Union, COM16-C966, September 2015 (hereinafter referred to as “VCEG Proposal COM16-C966”), a quadtree-binary tree (QTBT) partitioning technique was proposed for future video coding standards beyond HEVC. Simulations have shown that the proposed QTBT structure is more efficient than the quadtree structure of HEVC that was used. In HEVC, for small blocks, inter prediction is restricted to reduce memory access for motion compensation, so bidirectional prediction is not supported for 4×8 and 8×4 blocks, and inter prediction is not supported for 4×4 blocks. In JEM's QTBT, these restrictions are removed.
[0194] In QTBT, a CU can have either a square or rectangular shape. For example, a Coding Tree Unit (CTU) is first partitioned by a quadtree structure. A quadtree leaf node can be further partitioned by a binary tree structure. There are two types of binary tree partitions: horizontal symmetric partitioning and vertical symmetric partitioning. In either case, the node is divided by splitting it horizontally or vertically at the center. A binary tree leaf node is called a Coding Unit (CU), and its segmentation is used for prediction and transformation processing without further partitioning. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. A CU may be composed of coding blocks (CBs) of different color components. For example, in the case of P slices and B slices in 4:2:0 chroma format, one CU contains one luma CB and two chroma CBs. It may also be composed of single-component CBs. For example, in the case of I slices, one CU contains only one luma CB (b) or only two chroma CBs.
[0195] The following parameters are defined for the QTBT partitioning scheme.
[0196] - CTU size: The size of the root node of the quadtree, the same concept as in HEVC.
[0197] - MinQT size: The minimum allowable quadtree leaf node size.
[0198] - MaxBT size: The maximum allowable binary tree root node size.
[0199] - MaxBT depth: The maximum allowable binary tree depth.
[0200] MinBT size: The minimum allowable binary tree leaf node size.
[0201] In an example of the QTBT partition structure, the CTU size is set as a 128×128 luma sample including two corresponding 64×64 blocks of chroma samples, the MInQT size is set as 16×16, the MaxBT size is set as 64×64, the MinBT size (both width and height) is set to 4×4, and the MaxBT depth is set as 4. To initially generate the quadtree leaf nodes, the quadtree partition is applied to the CTU. The quadtree leaf nodes can have sizes from 16×16 (i.e., the MinQT size) to 128×128 (i.e., the CTU size). When a quadtree node has a size equal to the MinQT size, no further quadtree is considered. When the quadtree leaf node is 128×128, it is not further divided by the binary tree because its size exceeds the MaxBT size (i.e., 64×64). Otherwise, the leaf quadtree node can be further partitioned by the binary tree. Thus, the quadtree leaf node is also the root node of the binary tree, which has a binary tree depth of 0. When the binary tree depth reaches the MaxBT depth (i.e., 4), no further division is considered. When a binary tree node has a width equal to the MinBT size (i.e., 4), no further horizontal division is considered. Similarly, when a binary tree node has a height equal to the MinBT size, no further vertical division is considered. The leaf nodes of the binary tree are further processed by prediction and transformation processing without further partitioning. In JEM, the maximum CTU size is 256×256 luma samples. The leaf nodes of the binary tree (CU) can be further processed (e.g., by performing a prediction process and a transformation process) without further partitioning.
[0202] Furthermore, the QTBT scheme supports the ability of luma and chroma to have separate QTBT structures. Currently, for P slices and B slices, the luma CTB and chroma CTB within one CTU can share the same QTBT structure. However, for I slices, the luma CTB can be partitioned into CUs by the QTBT structure, and the chroma CTB can be partitioned into chroma CUs by a different QTBT structure. This means that the CUs of I slices are composed of coding blocks of the luma component or coding blocks of two chroma components, while the CUs of P slices or B slices are composed of coding blocks of all three color components.
[0203] The encoder 200 applies a rate-distortion optimization (RDO) process to the QTBT structure to determine block partitioning.
[0204] Furthermore, a block partitioning structure named multi-type tree (MTT) has been proposed in U.S. Patent Application Publication No. 2017 / 0208336 to replace the QT, BT, and / or QTBT-based CU structure. The MTT partition structure is still a recursive tree structure. In MTT, multiple different partition structures (such as three or more) are used. For example, according to the MTT technology, three or more different partition structures can be used for each non-leaf node of the tree structure at each depth of the tree structure. The depth of a node in the tree structure can refer to the length of the path from the node to the root of the tree structure (e.g., the number of divisions). The partition structure can generally refer to how many different blocks a block can be divided into. The partition structure may be such that a quadtree partition structure may divide a block into four blocks, a binary tree partition structure may divide a block into two blocks, a ternary tree partition structure may divide a block into three blocks, and furthermore, the ternary tree partition structure may be a structure that may not divide the block in the center. There can be multiple different partition types in the partition structure. The partition type can further define how the block is divided, and the division method includes symmetric or asymmetric partition, uniform or non-uniform partition, and / or horizontal or vertical partition.
[0205] In MTT, at each depth of the tree structure, the encoder 200 can be configured to further divide a subtree using a particular partition type from among three or more partition structures. For example, the encoder 100 can be configured to determine a particular partition type from among QT, BT, ternary tree (TT), and other partition structures. In one example, the QT partition structure can include a quadtree of squares or a quadtree partition type of rectangles. The encoder 200 can partition a square block by using a quadtree partition of squares to divide the block into four equal-sized square blocks along the centers of both the horizontal and vertical directions. Similarly, the encoder 200 can partition a rectangular (e.g., non-square) block by using a quadtree partition of rectangles to divide the rectangular block into four equal-sized rectangular blocks along the centers of both the horizontal and vertical directions.
[0206] The BT partition structure may include at least one of a horizontally symmetric binary tree, a vertically symmetric binary tree, a horizontally asymmetric binary tree, or a vertically asymmetric binary tree partition type. In the case of the horizontally symmetric binary tree partition type, the encoder 200 may be configured to divide a block into two symmetric blocks of the same size horizontally along the center of the block. In the case of the vertically symmetric binary tree partition type, the encoder 200 may be configured to divide a block into two symmetric blocks of the same size vertically along the center of the block. In the case of the horizontally asymmetric binary tree partition type, the encoder 100 may be configured to divide a block into two blocks of different sizes horizontally. For example, similar to the PART_2N×nU or PART_2N×nD partition types, one block may be 1 / 4 the size of the parent block and the other block may be 3 / 4 the size of the parent block. In the case of the vertically asymmetric binary tree partition type, the encoder 100 may be configured to divide a block into two blocks of different sizes vertically. For example, similar to the PART_nL×2N or PART_nR×2N partition types, one block may be 1 / 4 the size of the parent block and the other block may be 3 / 4 the size of the parent block. In other examples, the asymmetric binary tree partition type may divide the parent block into fractions of different sizes. For example, one sub-block may be 3 / 8 of the parent block and the other sub-block may be 5 / 8 of the parent block. Of course, such partition types are either vertical or horizontal.
[0207] The TT partition structure is different from the partitioning of the QT or BT structure in that the TT partition structure does not split the block along the center. The central region of the block remains together in the same sub-block. Different from the QT that generates four blocks or the binary tree that generates two blocks, three blocks are generated when splitting according to the TT partition structure. Examples of partition types according to the TT partition structure include symmetric partition types (both horizontal and vertical), and asymmetric partition types (both horizontal and vertical). Furthermore, the symmetric partition type by the TT partition structure can be irregular / non-uniform, or regular / uniform. The asymmetric partition type by the TT partition structure is irregular / non-uniform. In one example, the TT partition structure may include at least one of the following partition types: horizontal regular / uniform symmetric ternary tree, vertical regular / uniform symmetric ternary tree, horizontal irregular / non-uniform symmetric ternary tree, vertical irregular / non-uniform symmetric ternary tree, horizontal irregular / non-uniform asymmetric ternary tree, or vertical irregular / non-uniform asymmetric ternary tree partition type.
[0208] Generally, the irregular / non-uniform symmetric ternary tree partition type is a partition type that is symmetric with respect to the center line of the block, but at least one of the resulting three blocks is not the same size as the other two. One preferred example is the case where the side blocks are 1 / 4 of the size of the block and the center block is 1 / 2 of the size of the block. The regular / uniform symmetric ternary tree partition type is a partition type that is symmetric with respect to the center line of the block and all of the resulting blocks are the same size. Such a partition is possible when the height or width of the block is a multiple of 3 depending on the vertical or horizontal split. The irregular / non-uniform asymmetric ternary tree partition type is a partition type that is not symmetric with respect to the center line of the block and at least one of the resulting blocks is not the same size as the other two.
[0209] In an example where a block (e.g., at a sub-tree node) is split into an asymmetric ternary partition type, the encoder 200 and / or the decoder 300 applies a restriction such that two of the three partitions have the same size. Such a restriction may correspond to a restriction that the encoder 200 must follow when encoding video data. Further, in some examples, the encoder 200 and the decoder 300 can apply a restriction that the sum of the areas of two partitions is equal to the area of the remaining partition when splitting according to the asymmetric ternary partition type.
[0210] In some examples, the encoder 200 can be configured to select from all of the partition types described above for each of the QT, BT, and TT partition structures. In other examples, the encoder 200 can be configured to determine only the partition type from a subset of the partition types described above. For example, a subset of the partition types discussed above (or other partition types) can be used for a particular block size or a particular depth of the quadtree structure. The subset of supported partition types can be signaled in the bitstream for use by the decoder 200 or can be predefined such that the encoder 200 and the decoder 300 can determine the subset without signaling.
[0211] In other examples, the number of supported partition split types can be fixed for all depths of all CTUs. That is, the encoder 200 and the decoder 300 can be pre-configured to use the same number of partition split types for any depth of the CTU. In other examples, the number of supported partition split types can vary and can depend on depth, slice type, or other previously coded information. In one example, only the QT partition structure is used at depth 0 or depth 1 of the tree structure. At depths greater than 1, each of the QT, BT, and TT partition structures can be used.
[0212] In some examples, the encoder 200 and / or decoder 300 can apply pre-configured constraints to the supported partition types to avoid duplicate partitionings for specific regions of the video image or regions of CTUs. In one example, when a block is partitioned in an asymmetric partition type, the encoder 200 and / or decoder 300 can be configured not to further partition the largest sub-block divided from the current block. For example, when a square block is partitioned according to an asymmetric partition type (similar to the PART_2N×nU partition type), the largest sub-block among all sub-blocks (similar to the largest sub-block of the PART_2N×nU partition type) is the noted leaf node and cannot be further partitioned. However, small sub-blocks (similar to the small sub-blocks of the PART_2N×nU partition type) can be further partitioned.
[0213] As another example of applying constraints to the supported partition types to avoid duplicate partitionings for specific regions, when a block is partitioned in an asymmetric partition type, the largest sub-block divided from the current block cannot be further partitioned in the same direction. For example, when a square block is partitioned into an asymmetric partition type (similar to the PART_2N×nU partition type), the encoder 200 and / or decoder 300 can be configured not to horizontally partition the largest sub-block among all sub-blocks (similar to the largest sub-block of the PART_2N×nU partition type).
[0214] As another example of applying constraints to the supported partition types to avoid further partitioning difficulties, the encoder 200 and / or decoder 300 can be configured not to partition a block either horizontally or vertically when the width / height of the block is not a power of 2 (e.g., when the width / height is not 2, 4, 8, 16, etc.).
[0215] The above example describes how the encoder 200 can be configured to perform MTT partition splitting. Next, the decoder 300 can also apply the same MTT partition splitting as executed by the encoder 200. In some examples, how the image of the video data is partitioned by the encoder 200 can be determined by applying the same set of predefined rules in the decoder 300. However, in many situations, the encoder 200 can determine the specific partition structure and partition type to use based on the rate-distortion criterion of a specific image of the video data being coded. Then, in order for the decoder 300 to determine the partition splitting for a specific image, the encoder 200 can signal to a syntax element in the encoded bitstream indicating how the image and the CTU of the image should be partitioned. The decoder 200 can parse such a syntax element and partition the image and CTU accordingly.
[0216] In one example, the prediction processing unit 260 of the video encoder 200 can be configured to execute any combination of the above partition splitting techniques, particularly for motion estimation, the details of which will be described later.
[0217] Similar to the image 201, the block 203 can also be regarded as a two-dimensional array or matrix of samples having intensity values (sample values), but with dimensions smaller than those of the image 201. In other words, the block 203 can include, for example, one sample array (e.g., the luma array in the case of a monochrome image 201), or three sample arrays (e.g., the luma array and two chroma arrays in the case of a color image 201), or any other number and / or type of arrays depending on the color format applied. The number of samples in the horizontal and vertical (or axial) directions of the block 203 defines the size of the block 203.
[0218] An encoder 200 as shown in FIG. 2 is configured to encode an image 201 block by block. For example, encoding and prediction are performed for each block 203.
[0219] Residual calculation
[0220] A residual calculation unit 204 is configured to calculate a residual block 205 by subtracting the sample values of a prediction block 265 (further details about the prediction block 265 will be provided later) from the sample values of an image block 203, for example, for each sample (for each pixel), based on the image block 203 and the prediction block 265, so as to obtain the residual block 205 in the sample domain.
[0221] Transformation
[0222] A transformation processing unit 206 is configured to apply a transformation, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain transformation coefficients 207 in the transform domain. The transformation coefficients 207 may also be referred to as transform residual coefficients and represent the residual block 205 in the transform domain.
[0223] The conversion processing unit 206 may be configured to apply integer approximations of DCT / DST, such as the conversion specified for HEVC / H.265. Compared to the orthogonal DCT transform, such integer approximations are typically scaled by a particular factor. An additional scaling factor is applied as part of the conversion process to maintain the norm of the residual block processed by the forward and inverse transforms. The scaling factor is typically selected based on specific constraints such as the scaling factor which is a power of two of a shift operation, the bit depth of the transform coefficients, the trade-off between accuracy and implementation cost, etc. For example, a particular scaling factor may be specified, for example, for the inverse transform by the inverse transform processing unit 212 in the decoder 300 (and the corresponding inverse transform, for example, by the inverse transform processing unit 212 in the decoder 300), and the corresponding scaling factor for the forward transform by the conversion processing unit 206 in the encoder 200, for example, can be specified accordingly.
[0224] Quantization
[0225] The quantization unit 208 is configured to obtain a quantized transform coefficient 209 by quantizing the transform coefficient 207, for example, by applying scalar quantization or vector quantization. The quantized transform coefficient 209 may also be referred to as the quantized residual coefficient 209. The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient is truncated to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization can be changed by adjusting the quantization parameter (QP). For example, in the case of scalar quantization, various scalings can be applied to achieve finer or coarser quantization. Reducing the quantization step size corresponds to finer quantization, and increasing the quantization step size corresponds to coarser quantization. The applicable quantization step size can be indicated by the quantization parameter (QP). The quantization parameter can be, for example, an index for a predefined set of applicable quantization step sizes. For example, a small quantization parameter can correspond to finer quantization (small quantization step size), a large quantization parameter can correspond to coarser quantization (large quantization step size), or vice versa. Quantization can include division by the quantization step size and the corresponding inverse dequantization, for example, multiplication by the quantization step size by the inverse quantization 210. Some standards, for example, embodiments according to HEVC, can be configured to determine the quantization step size using the quantization parameter. Generally, the quantization step size can be calculated based on the quantization parameter using a fixed-point approximation of an equation including division. Additional scaling factors may be introduced for quantization and inverse quantization to restore the norm of the residual block, which may be changed for the scaling used in the fixed-point approximation of the equation of the quantization step size and the quantization parameter. In one example embodiment, the inverse transform scaling and inverse quantization can be combined. Alternatively, a customized quantization table can be used and signaled, for example, from the encoder to the decoder in the bitstream. Quantization is an irreversible operation, and the loss increases as the quantization step size increases.
[0226] The inverse quantization unit 210 applies the inverse quantization of the quantization unit 208 to the quantization coefficients to obtain an inverse quantization coefficient 211 based on or using the same quantization step size as the quantization unit 208, for example, by applying the inverse of the quantization scheme applied by the quantization unit 208. The inverse quantization coefficient 211 may also be referred to as the inverse quantized residual coefficient 211 and typically is not identical to the transform coefficient due to loss by quantization, but corresponds to the transform coefficient 207.
[0227] The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain an inverse transform block 213 in the sample domain. The inverse transform block 213 may also be referred to as the inverse transform inverse quantization block 213 or the inverse transform residual block 213.
[0228] The reconstruction unit 214 (e.g., summer 214) is configured to add the inverse transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain a reconstructed block 215 in the sample domain, for example, by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265.
[0229] Optionally, the buffer unit 216 (abbreviated as "buffer" 216), for example, the buffer unit 216, is configured to buffer or store the reconstructed block 215 and its respective sample values, for example, for intra prediction. In a further embodiment, the encoder may be configured to use the unfiltered reconstructed block and / or its respective sample values stored in the buffer unit 216 for any kind of estimation and / or prediction, such as intra prediction.
[0230] Embodiments of the encoder 200 may be configured such that, for example, the buffer unit 216 is used not only to store the block 215 reconstructed for intra prediction 254, but also for a loop filter unit 220 (not shown in FIG. 2), and / or, for example, the buffer unit 216 and the decoded picture buffer unit 230 may be configured to form one buffer. Further embodiments may be configured to use the filtered block 221 and / or a block or sample (both not shown in FIG. 2) from the decoded picture buffer 230 as an input or basis for intra prediction 254.
[0231] The loop filter unit 220 (abbreviated as "loop filter" 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221, for example, to smooth pixel shift or improve video quality in other ways. The loop filter unit 220 is intended to represent one or more loop filters such as a deblocking filter, a sample adaptive offset (SAO) filter, or other filters, for example, a bilateral filter or an adaptive loop filter (ALF), or a sharpening or smoothing filter or a collaborative filter. Although the loop filter unit 220 is shown as an in-loop filter in FIG. 2, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstructed block 221. The decoded picture buffer 230 can store the reconstructed coding block after the loop filter unit 220 performs a filtering operation on the reconstructed coding block.
[0232] Embodiments of the encoder 200 (each loop filter unit 220) may be configured to output loop filter parameters (such as sample adaptive offset information) directly or entropy - encoded via, for example, the entropy encoding unit 270 or any other arbitrary entropy coding unit, such that, for example, the decoder 300 can receive and apply the same loop filter parameters for decoding.
[0233] The decoded picture buffer (DPB) 230 may be a reference picture memory that stores reference picture data for use in encoding video data by the video encoder 20. The DPB 230 may be formed by any of various memory devices such as synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of dynamic random - access memory (DRAM) including other memory devices. The DPB 230 and the buffer 216 may be provided by the same memory device or separate memory devices. In some examples, the decoded picture buffer (DPB) 230 is configured to store the filtered block 221. The decoded picture buffer 230 may be further configured to store other previously filtered blocks of the same current picture or different pictures (e.g., previously reconstructed pictures), such as the previously reconstructed and filtered block 221, and may provide, for example, fully previously reconstructed, i.e., decoded, pictures (and corresponding reference blocks and samples) and / or partially reconstructed current pictures (and corresponding reference blocks and samples) for inter - prediction. In some examples, when the reconstructed block 215 is reconstructed but there is no in - loop filtering, the decoded picture buffer (DPB) 230 is configured to store the reconstructed block 215.
[0234] The prediction processing unit 260, also referred to as the block prediction processing unit 260, receives the block 203 (the current block 203 of the current image 201) and the reconstructed image data, such as reference samples of the same (current) image from the buffer 216 and / or reference image data 231 from one or more previously decoded images from the decoded image buffer 230, and processes such data for prediction, i.e., is configured to provide a prediction block 265 (which can be an inter prediction block 245 or an intra prediction block 255).
[0235] The mode selection unit 262 can be configured to select a prediction mode (e.g., an intra or inter prediction mode) and / or a corresponding prediction block 245 or 255 to be used as the prediction block 265 for the calculation of the residual block 205 and the reconstruction of the reconstructed block 215.
[0236] Embodiments of the mode selection unit 262 can be configured to select a prediction mode (e.g., from the modes supported by the prediction processing unit 260), which provides the best match, in other words, the minimum residual (the minimum residual means better compression for transmission or storage), or the minimum signaling overhead (the minimum signaling overhead means better compression for transmission or storage), or consider or balance both. The mode selection unit 262 can be configured to determine the prediction mode based on rate-distortion optimization (RDO), i.e., provide the minimum rate-distortion optimization, or select a prediction mode whose associated rate distortion at least meets the prediction mode selection criteria.
[0237] In the following, the prediction processing (e.g., the prediction processing unit 260 and mode selection (e.g., by the mode selection unit 262)) performed by the exemplary encoder 200 will be described in more detail.
[0238] As described above, the encoder 200 is configured to determine or select the best or optimal prediction mode from a set of (predetermined) prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0239] The intra prediction mode set may include 35 different intra prediction modes, such as non-directional modes like the DC (i.e., average) mode and the planar mode, or directional modes (such as those defined in H.265), or may include 67 different intra prediction modes, non-directional modes like the DC (i.e., average) mode and the planar mode, or directional modes (such as those defined in the developing H266).
[0240] A series of (or possible) inter prediction modes depend on the available reference images (i.e., at least partially decoded images previously, such as the images stored in the DBP 230) and other inter prediction parameters. For example, whether the entire reference image or only a part of the reference image (e.g., the search window area around the area of the current block of the reference image) is used to search for the most matching reference block, and / or for example, whether pixel interpolation, such as half / semi - pel and / or quarter - pel interpolation, is applied or not.
[0241] In addition to the above - mentioned prediction modes, a skip mode and / or a direct mode can be applied.
[0242] The prediction processing unit 260 may further be configured to repeatedly use, for example, quadtree partitioning (QT), binary tree partitioning (BT), ternary tree partitioning (TT), or any combination thereof, to partition the block 203 into smaller block partitions or sub - blocks and perform predictions for each of the block partitions or sub - blocks. The mode selection includes the selection of the tree structure of the partitioned block 203 and the prediction mode applied to each of the block partitions or sub - blocks.
[0243] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (not shown in FIG. 2). The motion estimation unit is configured to receive or obtain the image block 203 (the current image block 203 of the current image 201) and the decoded image 331, or at least one or a plurality of previously reconstructed blocks, for example, the reconstructed blocks of one or more other / different previously decoded images 331 for motion estimation. For example, the video sequence may include the current image and the previously decoded image 331. In other words, the current image and the previously decoded image 331 are part of or may form a sequence of images forming the video sequence. The encoder 200 may be configured to select a reference block from a plurality of reference blocks of the same or different images among a plurality of other images, and provide the reference image (or reference image index, ···) and / or offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as inter prediction parameters for the motion estimation unit (not shown in FIG. 2). This offset is also called a motion vector (MV). Merging is an important motion estimation tool used in HEVC and carried over to VCC. To perform merge estimation, it is first necessary to create a merge candidate list, where each candidate includes information on whether one or two reference image lists are used, and all motion data including the reference index and motion vector of each list. The merge candidate list is created based on a. up to four spatial merge candidates derived from five spatially adjacent blocks, b. one temporal merge candidate derived from two blocks placed temporarily in the same location, c. additional merge candidates including combined bi-prediction candidates and zero motion vector candidates.
[0244] The intra prediction unit 254 is further configured to determine based on intra prediction parameters, such as a selected intra prediction mode, and an intra prediction block 255. In any case, after selecting the intra prediction mode of a block, the intra prediction unit 254 is also configured to provide intra prediction parameters, that is, information indicating the selected intra prediction mode for the block to the entropy encoding unit 270. In one example, the intra prediction unit 254 may be configured to execute any combination of the intra prediction techniques described below.
[0245] The entropy encoding unit 270 is configured to obtain encoded image data 21 that can be output by an output 272 by applying an entropy encoding algorithm or method (for example, a variable length coding (VLC) method, a context adaptive VLC method (CALVC), an arithmetic coding method, a context adaptive binary arithmetic coding (CABAC), a syntax-based context adaptive binary arithmetic coding (SBAC), a probability interval partition entropy (PIPE) coding, or another entropy encoding method or technique) individually or together (or not at all) to the quantized residual coefficients 209, inter prediction parameters, intra prediction parameters, and / or loop filter parameters, for example, in the form of an encoded bitstream 21. The encoded bitstream 271 may be transmitted to the video decoder 30 or archived for later transmission or retrieval by the video decoder 30. The entropy encoding unit 270 may be further configured to entropy encode other syntax elements of the currently encoded video slice.
[0246] Other structural variations of the video encoder 200 can be used to encode the video stream. For example, a non-transform-based encoder 200 can directly quantize the residual signal for a particular block or frame without the transform processing unit 206. In another embodiment, the encoder 200 can combine the quantization unit 208 and the inverse quantization unit 210 into a single unit.
[0247] FIG. 3 shows an exemplary video decoder 300 configured to implement the technology of the present application. The video decoder 300 is configured to receive, for example, encoded image data (e.g., an encoded bitstream) 271 encoded by the encoder 200 and obtain a decoded image 331. During the decoding process, the video decoder 300 receives video data from the video encoder 200, e.g., an encoded video bitstream representing an encoded video slice and associated syntax elements of an image block.
[0248] In the example of FIG. 3, the decoder 300 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a buffer 316, a loop filter 320, a decoded image buffer 330, and a prediction processing unit 360. The prediction processing unit 360 can include an inter prediction unit 344, an intra prediction unit 354, and a mode selection unit 362. The video decoder 300, in some examples, executes a decoding path that is substantially the reverse of the encoding path described with respect to the video encoder 200 of FIG. 2.
[0249] Entropy decoding unit 304 is configured to perform entropy decoding on the encoded image data 271 to obtain, for example, quantization coefficients 309 and / or decoded coding parameters (not shown in FIG. 3), such as (decoded) inter prediction parameters, intra prediction parameters, loop filter parameters, and / or any or all of other syntax elements. Entropy decoding unit 304 is further configured to transfer inter prediction parameters, intra prediction parameters, and / or other syntax elements to prediction processing unit 360. Video decoder 300 can receive syntax elements at the video slice level and / or the video block level.
[0250] Inverse quantization unit 310 may have the same function as inverse quantization unit 110, inverse transform processing unit 312 may have the same function as inverse transform processing unit 112, reconstruction unit 314 may have the same function as reconstruction unit 114, buffer 316 may have the same function as buffer 116, loop filter 320 may have the same function as loop filter 120, and decoded image buffer 330 may have the same function as decoded image buffer 130.
[0251] Prediction processing unit 360 may include an inter prediction unit 344 and an intra prediction unit 354, where inter prediction unit 344 may have a similar function to inter prediction unit 14 4, and intra prediction unit 354 may have a similar function to intra prediction unit 154. Prediction processing unit 360 is typically configured to perform block prediction and / or obtain prediction block 365 from the encoded data 21, and receive or obtain (explicitly or implicitly) prediction-related parameters and / or information regarding the selected prediction mode, for example, from entropy decoding unit 304.
[0252] When the video slice is coded as an intra-coded (I) slice, the intra prediction unit 354 of the prediction processing unit 360 is configured to generate a prediction block 365 of the image block of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current frame or image. When the video frame is coded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the prediction processing unit 360 is configured to generate a prediction block 365 of the video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. For inter prediction, the prediction block can be generated from one of the reference images within one of the reference image lists. The video decoder 300 can construct the reference frame lists, list 0 and list 1, using default construction techniques based on the reference images stored in the DPB 330.
[0253] The prediction processing unit 360 is configured to determine prediction information for the video blocks of the current video slice by analyzing the motion vector and other syntax elements, and use the prediction information to generate a prediction block of the currently decoded video block. For example, the prediction processing unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra or inter prediction) used to code the video blocks of the video slice, the inter prediction slice type (e.g., B slice, P slice, or GPB slice), the configuration information of one or more reference image lists of the slice, the motion vector of each inter-coded video block of the slice, the inter prediction status of each inter-coded video block of the slice, and other information for decoding the video blocks of the current video slice.
[0254] The inverse quantization unit 310 is configured to inverse quantize, i.e., de-quantize, the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 304. The inverse quantization process may include determining the degree of quantization and, similarly, the degree of inverse quantization to be applied, using the quantization parameter calculated by the video encoder 100 for each video block within the video slice.
[0255] The inverse transform processing unit 312 is configured to apply an inverse transform, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients to generate a residual block in the pixel domain.
[0256] The reconstruction unit 314 (e.g., adder 314) is configured to obtain a reconstructed block 315 in the sample domain by adding the inverse transform block 313 (i.e., the reconstructed residual block 313) to the prediction block 365, for example, by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365.
[0257] The loop filter unit 320 (either within or after the coding loop) filters the reconstructed block 315 to obtain a filtered block 321, and is configured to, for example, smooth pixel shifts or improve video quality in other ways. In one example, the loop filter unit 320 can be configured to perform any combination of the filtering techniques described later. The loop filter unit 320 is intended to represent one or more loop filters such as a deblocking filter, a sample adaptive offset (SAO) filter, or other filters, for example, a bilateral filter or an adaptive loop filter (ALF), or a sharpening or smoothing filter or a collaborative filter. Although the loop filter unit 320 is shown as an in-loop filter in FIG. 3, in other configurations, the loop filter unit 320 can be implemented as a post-loop filter.
[0258] Next, the decoded video block 321 within a given frame or image is stored in a decoded image buffer 330 that stores reference images used for subsequent motion compensation.
[0259] The decoder 300 is configured to output the decoded image 311, for example, via output 312, for presentation or display to the user.
[0260] Other variations of the video decoder 300 can be used to decode the compressed bitstream. For example, the decoder 300 can generate an output video stream without the loop filtering unit 320. For example, a non-transform-based decoder 300 can directly inverse quantize the residual signal for a particular block or frame without the inverse transform processing unit 312. In another embodiment, the video decoder 300 can combine the inverse quantization unit 310 and the inverse transform processing unit 312 into a single unit.
[0261] FIG. 4 is a schematic diagram of a network device 400 (e.g., a coding device) according to an embodiment of the present disclosure. The network device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the network device 400 can be a decoder such as the video decoder 300 of FIG. 1A or an encoder such as the video encoder 200 of FIG. 1A. In one embodiment, the network device 400 can be one or more components of the video decoder 300 of FIG. 1A or the video encoder 200 of FIG. 1A as described above.
[0262] The network device 400 includes an input port 410 and a receiver unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit (CPU) 430 for processing data; a transmitter unit (Tx) 440 and an output port 450 for transmitting data; and a memory 460 for storing data. The network device 400 may also include an optical-to-electrical (OE) conversion component and an electrical-to-optical (EO) conversion component coupled to the input port 410, the receiver unit 420, the transmitter unit 440, and the output port 450 for input or output of optical or electrical signals.
[0263] Processor 430 is implemented by hardware and software. Processor 430 can be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGA, ASIC, and DSP. Processor 430 communicates with input port 410, receiver unit 420, transmitter unit 440, output port 450, and memory 460. Processor 430 includes coding module 470. Coding module 470 implements the above-disclosed embodiments. For example, coding module 470 performs, processes, prepares, or provides various coding operations. Therefore, including coding module 470 provides a substantial improvement to the function of network device 400 and brings about a conversion of network device 400 to different states. Alternatively, coding module 470 is implemented as instructions stored in memory 460 and executed by processor 430.
[0264] Memory 460 includes one or more disks, tape drives, and solid-state drives, is used as an overflow data storage device, stores programs when such programs are selected for execution, and can store instructions and data read during program execution. Memory 460 can be volatile and / or non-volatile and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).
[0265] FIG. 5 is a simplified block diagram of a device 500 that can be used as either or both of the source device 12 and the destination device 14 according to FIG. 1A according to an exemplary embodiment. Device 500 can implement the technology of this application. Device 500 can be in the form of a computing system including a plurality of computing devices, or in the form of a single computing device, such as a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, etc.
[0266] The processor 502 within the device 500 can be a central processing unit. Alternatively, the processor 502 can be any other type of device, or a plurality of devices, capable of manipulating or processing information that exists currently or will be developed in the future. The disclosed embodiments state can be implemented with a single processor, such as processor 502, as shown, but the advantages of speed and efficiency can be achieved using multiple processors.
[0267] The memory 504 within the device 500 can be, in an embodiment, a read-only memory (ROM) device or a random access memory (RAM) device in an embodiment. Other suitable types of storage devices can be used as the memory 504. The memory 504 can include code and data 506 that are accessed by the processor 502 using a bus 512. The memory 504 can further include an operating system 508 and an application program 510, and the application program 510 includes at least one program that enables the processor 502 to execute the methods described herein. For example, the application program 510 can include applications 1 to N, and the applications 1 to N further include video coding applications that execute the methods described herein. The device 500 can also include additional memory in the form of a secondary storage device 514, which can be, for example, a memory card used with a mobile computing device. Since video communication sessions can contain a significant amount of information, those sessions can be stored in whole or in part in the secondary storage device 514 and loaded into the memory 504 as needed for processing.
[0268] Device 500 may also include one or more output devices such as display 518. In one example, display 518 may be a touch-sensing display that combines a display with a touch-sensing element operable to sense touch input. Display 518 can be coupled to processor 502 via bus 512. Other output devices that enable a user to use device 500 in a programmatic or other manner may be provided in addition to, or in place of, display 518. When the output device is, or includes, a display, the display can be implemented in various ways including a liquid crystal display (LCD), a cathode ray tube (CRT) display, a plasma display, or a light-emitting diode (LED) display such as an organic LED (OLED) display.
[0269] Device 500 may also include an image sensing device 520, such as a camera, or any other currently existing or later developed image sensing device 520 capable of sensing an image, such as an image of a user operating device 500, or communicate with it. is Image sensing device 520 can be positioned to face a user operating device 500. In one example, the position and optical axis of image sensing device 520 are such that the field of view includes an area that is directly adjacent to and visible from display 518.
[0270] Device 500 may also include a sound sensing device 522, such as a microphone, or any other currently existing or later developed sound sensing device capable of sensing sound near device 500, or communicate with it. Sensing device 522 can be positioned to face a user operating device 500 and configured to receive sound made by the user while the user is operating device 500, such as speech or other utterances.
[0271] FIG. 5 shows that the processor 502 and the memory 504 of the device 500 are integrated into a single unit, but other configurations can be utilized. The operation of the processor 502 can be distributed among a plurality of machines (each machine having one or more processors) that can be directly coupled or coupled via a local area or other network. The memory 504 can be distributed among a plurality of machines such as network-based memory or memory within a plurality of machines that execute the operation of the device 500. Although shown here as a single bus, the bus 512 of the device 500 can be composed of a plurality of buses. Further, the secondary storage device 514 can be directly coupled to other components of the device 500 or accessed via a network and can include a single integrated unit such as a memory card or a plurality of units such as a plurality of memory cards. Thus, the device 500 can be implemented in a wide variety of configurations.
[0272] In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array. The nominal vertical and horizontal relative positions of the luma and chroma samples within the image are shown in FIG. 6A.
[0273] FIG. 8 is a conceptual diagram showing an exemplary position where scaling parameters used to scale the downsampled, reconstructed luma block are derived. For example, FIG. 8 shows an example of 4:2:0 sampling, and the scaling parameters are α and β.
[0274] Generally, when the LM prediction mode is applied, the video encoder 20 and the video decoder 30 can call the following steps. The video encoder 20 and the video decoder 30 can downsample adjacent luma samples. The video encoder 20 and the video decoder 30 can derive linear parameters (i.e., α and β, also called scaling parameters). The video encoder 20 and the video decoder 30 can downsample the current luma block and derive a prediction (e.g., a prediction block) from the downsampled luma block and the linear parameters. There are various methods for performing downsampling.
[0275] FIG. 6B is a conceptual diagram showing an example of luma positions and chroma positions for downsampling samples of a luma block to generate a prediction block for a chroma block. As shown in FIG. 6B, chroma samples represented by filled (i.e., all - black) triangles are predicted from two luma samples represented by two filled circles by applying a [1, 1] filter. The [1, 1] filter is an example of a 2 - tap filter.
[0276] FIG. 6C is a conceptual diagram showing another example of luma positions and chroma positions for downsampling samples of a luma block to generate a prediction block. As shown in FIG. 6C, chroma samples represented by filled (i.e., all - black) triangles are predicted from six luma samples represented by six filled circles by applying a 6 - tap filter.
[0277] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to search for instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0278] Video compression techniques such as motion compensation, intra prediction, and loop filtering have proven effective and are thus adopted in various video coding standards such as H.264 / AVC and H.265 / HEVC. Intra prediction can be used when there is no available reference picture, or for example in an I-frame or I-slice when inter prediction coding is not being used for the current block or picture. The reference samples for intra prediction are typically derived from previously coded (i.e., reconstructed) adjacent blocks within the same picture. For example, in both H.264 / AVC and H.265 / HEVC, the boundary samples of adjacent blocks are used as references for intra prediction. There are many different intra prediction modes to cover different texture or structural characteristics. In each mode, a different prediction signal derivation method is used. For example, as shown in FIG. 6D, H.265 / HEVC supports a total of 35 intra prediction modes.
[0279] Description of the Intra Prediction Algorithm of H.265 / HEVC
[0280] For intra prediction, the decoded boundary samples of adjacent blocks are used as references. The encoder selects the best luma intra prediction mode (i.e., the mode that provides the most accurate prediction for the current block) for each block from 35 options (33 directional prediction modes, DC mode, and planar mode). The mapping between the intra prediction direction and the intra prediction mode number is specified in FIG. 6D. Note that in the latest video coding technologies, such as VVC (Versatile Video Coding), more than 65 intra prediction modes have been developed, and VVC (Versatile Video Coding) can capture any edge direction presented in natural videos. Among these prediction modes, the modes with a horizontal direction (e.g., mode 10 in FIG. 6D) are also called "horizontal modes", and the modes with a vertical direction (e.g., mode 26 in FIG. 6D) are also called "vertical modes".
[0281] FIG. 7 shows the reference samples of a block. As shown in FIG. 7, the block "CUR" is the current block to be predicted, and the dark samples along the boundary of the current block are the reference samples used to predict the current block. These reference samples are samples within the reconstructed blocks adjacent to the current block, and are also called adjacent blocks. The block "CUR" can be a luma block or a chroma block depending on the type of the block to be predicted. The prediction signal can be derived by mapping the reference samples according to a specific method indicated by the intra prediction mode.
[0282] Replacement of Reference Samples
[0283] Some or all of the reference samples may not be available for intra prediction for several reasons. For example, samples outside the image, slice, or tile are considered unavailable for prediction. Further, when constrained intra prediction is enabled, reference samples belonging to an inter-predicted PU are omitted to avoid error propagation from previous images that may have been incorrectly received and reconstructed. As used herein, a reference sample for the current coding block is available if it is not outside the current image, slice, or tile, if the reference sample has been reconstructed before the current coding block is decoded, and / or if the reference sample has not been omitted for coding decisions at the encoder. In HEVC, all prediction modes can be used after replacing the unavailable reference samples. In the extreme case where there are no available reference samples, all reference samples are replaced with the nominal average sample value for a given bit depth (e.g., 128 for 8-bit data). If there is at least one reference sample marked as available for intra prediction, the unavailable reference samples are replaced using the available reference samples. The unavailable reference samples are replaced by scanning the reference samples in a clockwise order and using the most recent available sample value for the unavailable samples. If the first sample in the clockwise scan is unavailable, when scanning the samples in clockwise order, the unavailable reference sample is replaced with the first available reference sample detected. Here, "replacement" is also referred to as "padding", and the replaced sample may also be referred to as a "padded sample".
[0284] Constrained intra prediction
[0285] Constrained intra prediction is a tool for avoiding spatial noise propagation that occurs by spatial intra prediction using reference pixels of encoder-decoder mismatches. Reference pixels of encoder-decoder mismatches may appear when packet loss occurs during transmission of an inter-coded slice. These may also appear when lossy decoder-side memory compression is used. When constrained intra prediction is valid, inter-predicted samples are marked as not available or unavailable for intra prediction, and these unavailable samples can be padded by the above padding method to perform full intra prediction estimation on the encoder side or intra prediction on the decoder side.
[0286] Cross-component linear model prediction (CCLM)
[0287] Cross-component linear model prediction (CCLM), also called cross-component intra prediction (CCIP), is one type of intra prediction mode used to reduce cross-component redundancy during the intra prediction mode. FIG. 8 (including FIGS. 8A and 8B) is a schematic diagram showing an example of a mechanism for performing CCLM intra prediction. FIG. 8 shows an example of 4:2:0 sampling. FIG. 8 shows an example of the positions of samples of the current block included in the CCLM mode and its adjacent samples on the left and upper sides. The white squares are the samples of the current block, and the shaded circles are the reconstructed samples of the adjacent blocks. FIG. 8A shows an example of adjacent reconstructed pixels of a chroma block. FIG. 8B shows an example of adjacent reconstructed pixels of a luma block arranged in the same place. When the video format is YUV4:2:0, then there is one 16×16 luma block and two 8×8 chroma blocks.
[0288] CCLM intra prediction can be performed by the intra estimation unit 254 of the encoder 200 and / or the intra prediction unit 354 of the decoder 300. CCLM intra prediction predicts the chroma samples 803 within the chroma block 801. The chroma samples 803 appear at integer positions indicated by the squares. The prediction is partially based on the adjacent reference samples indicated by the black circles. The chroma samples 803 are not predicted based only on the adjacent chroma reference samples 805. The chroma samples 803 are also predicted based on the luma reference samples 813 and the adjacent luma reference samples 815. Specifically, a CU includes a luma block 811 and two chroma blocks 801. A model is generated that correlates the chroma samples 803 and the luma reference samples 813 within the same CU. The linear coefficients of the model are determined by comparing the adjacent luma reference samples 815 with the adjacent chroma reference samples 805.
[0289] When the luma reference sample 813 is reconstructed, the luma reference sample 813 is shown as the reconstructed luma sample (Rec’L). When the adjacent chroma reference sample 805 is reconstructed, the adjacent chroma reference sample 805 is shown as the reconstructed adjacent chroma sample (Rec’C).
[0290] As shown, luma block 811 contains four times the samples of chroma block 801. In the example shown in FIG. 8, while chroma block 801 contains N×N samples, luma block 811 contains 2N×2N samples. Therefore, luma block 811 has a resolution four times that of chroma block 801. For prediction operating on luma reference samples 813 and adjacent luma reference samples 815, luma reference samples 813 and adjacent luma reference samples 815 are downsampled to provide an accurate comparison with adjacent chroma reference samples 805 and chroma samples 803. Downsampling is a process that reduces the resolution of a group of sample values. For example, when the YUV4:2:0 format is used, luma samples can be downsampled by a factor of four (e.g., by 2 in width and 2 in height). YUV is a color encoding system that uses a color space from the perspective of luma component Y and two chrominance components U and V.
[0291] In CCLM prediction, chroma samples are predicted based on the corresponding downsampled reconstructed luma samples (current luma block) using a linear model as follows. pred C (i,j)=α·rec L ’(i,j)+β (1) Here, pred C (i,j) represents the predicted chroma sample, and rec’ L (i,j) represents the corresponding downsampled reconstructed luma sample. Parameters α and β can be derived by minimizing the regression error between the reconstructed adjacent luma and chroma samples around the current luma block and current chroma block as follows. α=(N·Σ(L(n)·C(n))-ΣL(n)·ΣC(n)) / (N·Σ(L(n)·L(n))-ΣL(n)·ΣL(n)) (2) β=(ΣC(n)-α·ΣL(n)) / N (3) Here, L(n) represents the downsampled upper and left reconstructed adjacent luma samples, C(n) represents the upper and left reconstructed adjacent chroma samples, and the value of N is equal to the samples used for the derivation of the coefficients. In the case of a square-shaped coding block, the above two equations are directly applicable. Since this regression error minimization calculation is performed not only as an encoder search operation but also as part of the decoding process, no syntax is used to transmit the α and β values.
[0292] In addition to using the above method (also called the least squares (LS) method) to minimize the regression error, the linear model coefficients α and β can also be derived using the maximum and minimum luma sample values. This latter method is also called the MaxMin method. In the MaxMin method, after the downsampled upper and left reconstructed adjacent luma samples, a one-to-one relationship between each of these reconstructed adjacent luma samples and the upper and left reconstructed adjacent chroma samples is obtained. Thus, the linear model coefficient parameters α and β can be derived using pairs of luma and chroma samples based on the one-to-one relationship. The pairs of luma and chroma samples are obtained by identifying the minimum and maximum values of the downsampled upper and left reconstructed adjacent luma samples and then identifying the corresponding samples from the upper and left reconstructed adjacent chroma templates. The pairs of luma and chroma samples are shown as in (A, B) of FIG. 9. The linear model parameters α and β are obtained according to the following equations. α=(y B -y A ) / (x B -x A ) (4) β=y A -αx A 、 (5) Here, (x A 、y A ) are the coordinates of A in FIG. 9, and (x B 、y B ) are the coordinates of B.
[0293] The CCLM luma-to-chroma prediction mode is added as one additional chroma intra prediction mode. On the encoder side, one more rate-distortion (RD) cost check of the chroma component is added to select the chroma intra prediction mode.
[0294] For the sake of brevity, in this document, the term "template" is used to denote the reconstructed adjacent chroma samples and the downsampled reconstructed adjacent luma samples. These reconstructed adjacent chroma samples and downsampled reconstructed adjacent luma samples are also referred to as reference samples within the template. FIG. 10 is a diagram of a template of a chroma block and a corresponding downsampled luma block. In the example shown in FIG. 10, luma '1020 is the downsampled version of the current luma block and has the same spatial resolution as chroma block 1040. In other words, luma '1020 is the juxtaposed downsampled luma block of chroma block 1040. The upper template 1002 includes the upper reconstructed adjacent chroma samples above the current chroma block 1040 and the corresponding downsampled upper reconstructed adjacent luma samples of luma '1020. The downsampled upper reconstructed adjacent luma samples of luma '1020 are obtained based on the adjacent samples above the luma block. As used herein, the adjacent samples above the luma block may include either the adjacent samples immediately above the luma block, or the adjacent samples not adjacent to the luma block, or both. The left template 1004 includes the left reconstructed adjacent chroma samples and the corresponding downsampled left reconstructed adjacent luma samples. The upper reconstructed adjacent chroma samples are also referred to as "upper chroma templates" such as upper chroma template 1006. The corresponding downsampled upper reconstructed adjacent luma samples are referred to as "upper luma templates" such as upper luma template 1008. The left reconstructed adjacent chroma samples are also referred to as "left chroma templates" such as left chroma template 1010. The corresponding downsampled left reconstructed adjacent luma samples are referred to as "left luma templates" such as left luma template 1012. The elements included in the template are referred to as reference samples within that template.
[0295] In an existing CCLM application, when there is one reference sample marked as unavailable in the upper or left template, the entire template is not used. FIG. 11 is a diagram of an example of a template including an unavailable reference sample. In the example shown in FIG. 11, in the case of chroma block 1140, when there is an unavailable reference sample in the upper template such as the reference sample of A2 1102, then the upper template is not used for deriving the linear model coefficients. Similarly, when there is one unavailable reference sample in the left template like the reference sample of B2 1104, then the left template is not used for deriving the linear model coefficients. As a result, the coding performance of intra prediction deteriorates.
[0296] Multi-directional linear model
[0297] In addition to being used to calculate the linear model coefficients together, the reference samples in the upper template and the left template can also be used in two other CCLM modes, namely, the CCLM_T and CCLM_L modes. CCLM_T and CCLM_L can also be collectively referred to as the multi-directional linear model (MDLM). FIG. 12 is a diagram of the reference samples used in the CCLM_T mode, and FIG. 13 is a diagram of the reference samples used in the CCLM_L mode. As shown in FIG. 12, in the CCLM_T mode, only the reference samples in the upper template such as reference samples 1202 and 1204 are used to calculate the linear model coefficients. As shown in FIG. 13, in the CCLM_L mode, only the reference samples in the left template such as reference samples 1212 and 1214 are used to calculate the linear model coefficients. The number of reference samples used in each of these modes is W + H, where W is the width of the chroma block and H is the height of the chroma block.
[0298] The CCLM mode and the MDLM mode (i.e., the CCLM_T mode and the CCLM_L mode) can be used together or alternatively. For example, only the CCLM mode is used in the codec, only the MDLM is used in the codec, or both CCLM and MDLM are used in the codec. In the last case where both CCLM and MDLM are used, three modes (i.e., CCLM, CCLM_T, CCLM_L) are added as three additional chroma intra prediction modes. On the encoder side, in order to select the chroma intra prediction mode, three more RD cost checks of the chroma component are added. In the existing method of MDLM, the model parameters or model coefficients are derived using the LS method. When the number of available reference samples is not sufficient, a padding operation is used to copy the farthest pixel value or fetch the sample value of the available reference samples.
[0299] However, obtaining the linear model coefficients of the MDLM mode using the LS method increases the computational complexity. Furthermore, in the existing MDLM mode, the positions of some template samples may be far from the current block, especially in the case of non-square blocks. For example, the reference sample at the right end of the upper template and the reference sample at the lower part of the left template are far from the current block. Therefore, the correlation between these reference samples and the current block becomes low, and the efficiency of chroma block prediction decreases. The technology presented in this specification can reduce the complexity of MDLM and increase the correlation between the template samples and the current block.
[0300] In one example, when determining the model coefficients using the MaxMin method, in addition to using the reference samples from the upper template and the left template together, only a part of the reference samples of either the left template or the upper template is used. For example, in the MaxMin method, only the reference samples within the upper luma template are inspected, and the maximum luma value and the minimum luma value are determined. Alternatively, only the reference samples within the left luma template are inspected, and the maximum luma value and the minimum luma value are determined. After the sample positions of the maximum and minimum luma values are determined, the corresponding chroma sample values can be obtained based on the positions of the minimum and maximum luma values.
[0301] FIG. 14 is a schematic diagram showing an example of the reference samples used to determine the maximum and minimum luma values. In the example shown in FIG. 14, the number of reference samples within the upper luma template, shown as W1, is larger than the width of the current chroma block, shown as W. The number of reference samples within the left luma template, shown as H1, is larger than the height of the current chroma block, shown as H. FIG. 15 is a schematic diagram showing another example of the reference samples used to determine the maximum and minimum luma values. In the example shown in FIG. 15, the number of upper luma reference samples is equal to the width W of the current chroma block, and the number of left luma reference samples is equal to the height H of the current chroma block.
[0302] In summary, in addition to the LS method, the MaxMin method can also be used in the MDLM mode. In other words, the MaxMin method can be used to derive the model coefficients for the CCLM_T mode and the CCLM_L mode. Since the MaxMin method has a lower computational complexity than the LS method, the proposed method improves the MDLM by reducing its computational complexity. Furthermore, the existing MaxMin method uses both the upper template and the left template in the CCLM mode. The proposed method uses either the upper template or the left template to derive the model coefficients, thereby further reducing the computational complexity of the mode.
[0303] According to a further example of the technology presented in this specification, the reference samples in the template are selected to enhance the correlation between the reference samples and the current block. Up to W2 reference samples are used for the upper template. Up to H2 reference samples are used for the left template. In this way, reference samples that are farther away from the W2 reference samples in the upper template or the H2 reference samples in the left template are not used because their correlation with the current block is low.
[0304] Furthermore, when deriving the model coefficients using the MaxMin method, only the available reference samples are used and no padding is used to replace the unavailable reference samples. For example, when determining the maximum and minimum luma values in the CCLM_T mode, only the available samples in the upper luma template are inspected. Since up to W2 reference samples are used in the upper template, the number of available samples, denoted as W3, is less than or equal to W2. Similarly, when determining the maximum and minimum values in the CCLM_L mode, the available samples in the left luma template are inspected. The number of available samples, denoted as H3, can be less than or equal to H2. The relationship between W2 and W3 and between H2 and H3 is shown in FIG. 16, where W3 <= W2 and H3 <= H2. In one example, W2 = 2 × W and H2 = 2 × H.
[0305] Alternatively, W2 and H2 can each take on the value of W + H. In other words, the model coefficients for the CCLM_T mode can be derived using a maximum of W + H reference samples within the upper luma template, and the model coefficients for the CCLM_L mode can be derived using a maximum of W + H reference samples within the left luma template. Among these reference samples, only the luma template samples available within this range (i.e., W + H for both CCLM_T and CCLM_L) are examined, and the maximum and minimum luma values are determined. In this example, to determine the maximum and minimum values for the CCLM_T mode, the samples available in the upper luma template (up to W + H) are examined. To determine the maximum and minimum values for the CCLM_L mode, the samples available in the left luma template (up to W + H) are examined.
[0306] Compared to the existing MDLM method where exactly W + H reference samples are used to derive the model coefficients for both the CCLM_T and CCLM_L modes, in the proposed method, for deriving the model coefficients for the CCLM_T mode, a maximum of 2×W or W + H reference samples are used, and for deriving the model coefficients for the CCLM_L mode, a maximum of 2×H or W + H reference samples are used. Further, to determine the maximum and minimum values, only the luma reference samples available within the examined sample range (2×W or W + H for CCLM_T, 2×H or W + H for CCLM_L) are used. In the proposed method above, determining the maximum and minimum luma values within the luma template can be speeded up by sampling the reference samples at a step size greater than 1, such as 2, 4, or another value.
[0307] Downsampling method
[0308] As discussed above, since the spatial resolution of the luma component of an image is greater than that of the chroma component, it is necessary to downsample the luma component to the resolution of the chroma portion in the MDLM mode. For example, in the case of the YUV4:2:0 format, in order to match the resolution of the chroma component, it is necessary to downsample the luma component by 4 (by 2 in width and by 2 in height). Using the downsampled luma block corresponding to the chroma block (having the same spatial resolution as the chroma block), prediction of the chroma block using the MDLM mode can be performed. The size of the downsampled luma block is the same as that of the chroma block (since the luma block is downsampled to the size of the chroma block).
[0309] Similarly, in order to derive the linear model coefficients in the MDLM mode, it is necessary to downsample the reference samples in the luma template. In the case of the CCLM_T mode, the upper reconstructed adjacent luma samples are downsampled, and upper template reference samples corresponding to the reference samples in the upper chroma template, that is, upper reconstructed adjacent chroma samples are generated. The downsampling of the upper reconstructed adjacent luma samples typically includes neighboring multiple rows of the upper reconstructed neighborhood luma samples. FIG. 17 is a schematic diagram showing an example of downsampling luma samples using multiple rows or columns of luma samples. As shown in FIG. 17, in the case of a luma block, during downsampling, two upper neighborhood rows A1 and A2 can be used to obtain the downsampled adjacent row A. A[i] as the i-th sample of A, A1[i] as the i-th sample of A1, and A2[i] as the i-th sample of A2 are shown, and a 6-tap downsampling filter can be used as follows. A[i]=(A2[2i]*2+A2[2i-1]+A2[2i+1]+A1[2i]*2+A1[2i-1]+A1[2i+1]+4)>>3; The number of adjacent samples can also be larger than the size of the current block. For example, as shown in FIG. 18, the number of adjacent samples above the downsampled luma block can be M, where M is larger than the width W of the downsampled block.
[0310] In the existing downsampling methods as described above, multiple rows of the reconstructed neighborhood luma samples above are used to generate the reconstructed adjacent luma samples above the downsampled. This increases the size of the line buffer compared to normal intra-mode prediction, thus increasing the memory cost.
[0311] The technique presented in this specification reduces the memory usage of downsampling by using only one row of the reconstructed adjacent luma samples for the CCLM_T mode when the current block is at the upper boundary of the current coding tree unit CTU (i.e., the top row of the current chroma block overlaps with the top row of the current CTU). FIG. 19 is a schematic diagram showing an example of downsampling using a single row of adjacent luma samples for a luma block at the upper boundary of the CTU. As shown in FIG. 19, only A1 (including one row of the reconstructed adjacent luma samples) is used to generate the reconstructed adjacent luma samples above the downsampled at A.
[0312] The above description focuses on the reconstructed adjacent luma samples above the CCLM_T mode, but it should be understood that a similar method can be applied to the reconstructed adjacent luma samples on the left side of the CCLM_L mode. For example, instead of using multiple columns of the reconstructed neighborhood luma samples on the left side for downsampling, when the current block is at the left boundary of the CTU (i.e., when the left row of the current chroma block overlaps with the left row of the CTU), a single column of the reconstructed adjacent luma samples on the left side is used for downsampling to generate reference samples for the left template of the downsampled luma block.
[0313] Determination of Availability of Reference Samples
[0314] In the above example, a luma template is used to determine the availability of reference samples for determining the maximum and minimum luma values. However, in some scenarios, the available luma reference samples do not have corresponding chroma reference samples. For example, luma blocks and chroma blocks can be coded separately. Thus, when the reconstructed luma blocks are available, the corresponding chroma blocks may not yet be available. As a result, the available luma reference samples do not have corresponding available chroma reference samples, which may lead to coding errors.
[0315] The techniques presented in this specification address this problem by determining the availability of reference samples within a template, via an examination of the availability of reference samples within a chroma template. In some examples, a chroma reference sample is available if the chroma reference sample is not outside the current image, slice, or tile and the reference sample is reconstructed. In some examples, a chroma reference sample is available if the chroma reference sample is not outside the current image, slice, or tile, the reference sample is reconstructed, and the reference sample is not omitted based on an encoding decision. An available reference sample for a current chroma block can be a reconstructed adjacent sample for the chroma block. Luma reference samples corresponding to available chroma reference samples are used to determine maximum and minimum luma values. For example, if L reference samples within a chroma template are available, then the reference samples within the luma template corresponding to the L available chroma reference samples are used to determine the maximum and minimum luma values. A luma reference sample corresponding to a chroma reference sample can be determined by identifying the chroma reference sample located within the chroma template and the luma reference sample located at the same position within the luma template (i.e., the luma reference sample located at the same place (position (x, y))), and for example, a luma reference sample corresponding to a chroma reference sample can include adjacent luma reference samples (position (x - 1, y)), a luma reference sample (position (x, y)), and adjacent luma reference samples (position (x + 1, y)). The luma reference samples having the maximum and minimum luma values and the chroma reference samples corresponding to the luma reference samples associated with the maximum and minimum luma values are then used to determine the model coefficients as described above.
[0316] In one example, the model coefficients, e.g., L = 4, can be determined using L available chroma reference samples. In another example, the model coefficients are determined using a part of the L available chroma reference samples or a part of the L available chroma reference samples. For example, a fixed number of available chroma reference samples are selected from the L available chroma reference samples. The luma reference samples corresponding to the selected chroma reference samples are identified and used together with the selected chroma reference samples to determine the model coefficients. For example, when the selected chroma reference samples are 4 available chroma reference samples, 24 reconstructed adjacent luma samples corresponding to the 4 available chroma reference samples are identified. The luma reference samples used to determine the model coefficients are obtained by downsampling the 24 reconstructed adjacent luma samples, where a 6-tap filter is used in the downsampling process.
[0317] In the example shown in FIG. 16, the upper template sample range of the CCLM_T mode is W2. Next, the availability of the reference samples in the upper chroma template of the current chroma block is determined. When W3 chroma reference samples are available (W3 <= W2), then W3 or fewer corresponding luma reference samples are obtained. These obtained luma reference samples and chroma reference samples are used to derive the model coefficients of the CCLM_T mode.
[0318] Similarly, in the example shown in FIG. 16, there is a left template sample range as H2 in the CCLM_L mode. The availability of reference samples within the left chroma template of the current chroma block is examined. If H3 chroma reference samples are available (H3 <= H2), then next, H3 or fewer corresponding luma reference samples are obtained. These obtained luma reference samples and chroma reference samples are used to derive the model coefficients in the CCLM_L mode. The availability of reference samples in the CCLM mode is determined similarly, that is, by determining the availability of reference samples in both the upper chroma template and the left chroma template, and then finding the corresponding luma reference samples and determining the model coefficients as described above.
[0319] The details of the proposed method are described in Table 1 in the format of the specifications of the INTRA_CCLM, INTRA_CCLM_L, or INTRA_CCLM_T intra prediction mode. Table 2 shows an alternative embodiment of the method proposed in this specification.
Table 1-1
Table 1-2
Table 1-3
Table 1-4
Table 1-5
Table 2
[0320] Binarization of the MDLM Mode
[0321] In order to encode the MDLM mode in the bitstream of a video signal, it is necessary to perform binarization of the MDLM mode so that the selected MDLM mode can be encoded in the bitstream and the decoder can determine the mode selected for decoding. Existing binarization methods do not include the two chroma modes of MDLM, namely CCLM_L and CCLM_T. Here, a new chroma mode coding method is proposed.
[0322] Tables 3 and 4 provide details of the binarization of these two chroma modes. In Table 3, 77 indicates the CCLM mode and the intra_chroma_pred_mode index is 4. 78 indicates the CCLM_L mode and the intra_chroma_pred_mode index is 5. 79 indicates the CCLM_T mode and the intra_chroma_pred_mode index is 6. When the intra_chroma_pred_mode index is equal to 7, the selected mode is the DM mode. The remaining index values 0, 1, 2, 3 represent the planar mode, vertical mode, horizontal mode, and DC mode, respectively.
Table 3
Table 4
[0323] Table 4 shows examples of bit strings or syntax elements used for each of the chroma intra prediction modes. As shown in Table 4, the syntax element for the DM mode (index 7) is 0, the syntax element for the CCLM mode (index 4) is 10, the syntax element for the CCLM_L mode (index 5) is 1110, the syntax element for the CCLM_T mode (index 6) is 1111, the syntax element for the planar mode (index 0) is 11000, the syntax element for the vertical mode (index 1) is 11001, the syntax element for the horizontal mode (index 2) is 11010, and the syntax element for the DC mode (index 3) is 11011. Table 5 shows another example of bit strings or syntax elements used for each of the chroma intra prediction modes. Depending on the coding mode selected by the encoder, the corresponding syntax element is included in the bitstream of the encoded video.
Table 5
[0324] When a video encoder such as the video encoder 20 in FIG. 1 performs intra prediction of a chroma block of a video signal based on an intra chroma prediction mode, the video encoder selects an intra chroma prediction mode and includes a syntax element indicating the selected intra chroma prediction mode in the bitstream to generate a bitstream of the video signal. The video encoder can select an intra chroma prediction mode from a plurality of mode sets. For example, the mode can include a first mode set including a derived (DM) mode or a cross-component linear model (CCLM) prediction mode, or both. The mode can also include a second mode set including at least one of the CCLM_L mode or the CCLM_T mode. The mode can further include a third mode set that can include at least one of the vertical mode, the horizontal mode, the DC mode, or the planar mode.
[0325] In some examples, the number of bits of the syntax elements of the intra-chroma prediction mode when the intra-chroma prediction mode is selected from the first mode set is smaller than the number of bits of the syntax elements of the intra-chroma prediction mode when the intra-chroma prediction mode is selected from the second mode set. Further, the number of bits of the syntax elements of the intra-chroma prediction mode when the intra-chroma prediction mode is selected from the second mode set is smaller than the number of bits of the syntax elements of the intra-chroma prediction mode when the intra-chroma prediction mode is selected from the third mode set. In some examples, the syntax elements of the various intra-chroma prediction modes are selected according to the examples shown in Table 4 or Table 5.
[0326] When a decoder receives an encoded bitstream of a video signal and decodes the video, a decoder such as the video decoder 30 in FIG. 1 analyzes syntax elements from the bitstream of the video signal and determines an intra-chroma prediction mode to be used for a chroma block based on the syntax elements selected from the analyzed syntax elements. Based on the determined intra-chroma prediction mode, the decoder performs intra prediction on the current chroma block of the video signal.
[0327] FIG. 20 is a flowchart of a method for performing intra prediction using a linear model according to some aspects of the present disclosure. In block 2002, a luma block (such as luma block 811) corresponding to the current chroma block (such as chroma block 801) is determined.
[0328] In block 2004, the luma reference samples of the luma block are obtained based on determining L available chroma reference samples of the current chroma block. The obtained luma reference samples of the luma block are downsampled luma reference samples. In some examples, the obtained luma reference samples of the luma block are downsampled luma reference samples obtained by downsampling adjacent luma samples selected based on (such as based on some or all of the L available chroma reference samples) the L available chroma reference samples. In other words, the obtained luma reference samples of the luma block are downsampled luma reference samples obtained by downsampling adjacent luma samples corresponding to the available chroma reference samples. In some examples, the obtained luma reference samples correspond to the L available chroma reference samples. In additional examples, the obtained luma reference samples correspond to a portion of the L available chroma reference samples. It can be understood that the correspondence between the obtained luma reference samples (i.e., the downsampled luma reference samples) and the L available chroma reference samples may not be limited to a "one-to-one correspondence", and it can also be understood that the correspondence between the obtained luma reference samples (i.e., the downsampled luma reference samples) and the L available chroma reference samples may be an "M-to-N correspondence". For example, M = 4, N = 4, or M = 4, N>4.
[0329] In some examples, the chroma reference samples of the current chroma block include the reconstructed adjacent samples of the current chroma block. The L available chroma reference samples are determined from the reconstructed adjacent samples. Similarly, the adjacent samples of the luma block are also the reconstructed adjacent samples of the luma block. The obtained luma reference samples of the luma block are obtained by downsampling the reconstructed adjacent samples selected based on the L available chroma reference samples. Such that L = 4.
[0330] In some examples, a chroma reference sample is available when the chroma reference sample is not outside the current image, slice, or title and the reference sample is reconstructed. In some examples, a chroma reference sample is available when the chroma reference sample is not outside the current image, slice, or title, the reference sample is reconstructed, and the reference sample is not omitted based on an encoding decision. The available reference sample for the current chroma block can be the reconstructed adjacent samples available for the chroma block. A luma reference sample corresponding to the available chroma reference sample is obtained.
[0331] In some examples, L available chroma reference samples are determined by determining that the L upper adjacent chroma samples of the current chroma block are available, where 1 <= L <= W2, and L and W2 are positive integers. W2 indicates the upper reference sample range, and the L upper adjacent chroma samples are used as the available chroma reference samples. In some examples, W2 is equal to either 2*W or W + H. Here, W represents the width of the current chroma block, and H represents the height of the current chroma block.
[0332] In other examples, L available chroma reference samples are determined by determining the L left adjacent chroma samples available for the current chroma block. Here, 1 <= L <= H2, and L and H2 are positive integers. H2 indicates the left reference sample range. The L left adjacent chroma samples are used as the available chroma reference samples. In some examples, H2 is equal to either 2*H or W + H. W represents the width of the current chroma block, and H represents the height of the current chroma block.
[0333] In a further example, the L available chroma reference samples are determined by determining the L1 upper adjacent chroma samples and the L2 left adjacent chroma samples available for the current chroma block. Here, 1 <= L1 <= W2, and 1 <= L2 <= H2. W2 indicates the upper reference sample range, and H2 indicates the left reference sample range. L1, L2, W2, and H2 are positive integers, and L1 + L2 = L. In these examples, the L1 upper adjacent chroma samples and the L2 left adjacent chroma samples are used as the available chroma reference samples.
[0334] In one example, the luma reference samples are obtained by downsampling only the adjacent samples that are above the luma block and are selected based on the L available chroma reference samples. In another example, the luma reference samples are obtained by downsampling only the adjacent samples that are to the left of the luma block and are selected based on the L available chroma reference samples.
[0335] In the above examples, the downsampled luma block of the luma block is obtained by downsampling the reconstructed luma block of the luma block corresponding to the current chroma block. In some cases, for example, when the luma reference samples are obtained based only on the adjacent samples above the luma block, and when the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), etc., only one row of the reconstructed adjacent luma samples of the reconstructed version of the luma block is used to obtain the luma reference samples.
[0336] In block 2006, the linear model coefficients used for cross-component prediction are calculated based on the luma reference samples obtained in step 2004 and the chroma reference samples corresponding to the luma reference samples. In some examples, the chroma reference samples corresponding to the luma reference samples are the chroma reference samples arranged at the same location as the luma reference samples.
[0337] In block 2008, the prediction of the current chroma block is generated based on the calculated linear model coefficients and the values of the downsampled luma blocks obtained by downsampling the luma blocks (such as luma block 811).
[0338] FIG. 21 is a flowchart of a method for cross-component linear model (CCLM) prediction according to another aspect of the present disclosure. In block 2102, the luma block (such as luma block 811) corresponding to the current chroma block (such as chroma block 801) is determined.
[0339] In block 2104, the luma reference samples of the luma block are obtained by downsampling the adjacent samples of the luma block. In some examples, the luma reference samples include only the luma reference samples obtained based on the adjacent samples above the luma block. In other examples, the luma reference samples include only the luma reference samples obtained based on the adjacent samples to the left of the luma block.
[0340] In block 2106, the maximum luma value and the minimum luma value are determined based on the luma reference samples.
[0341] In block 2108, the first chroma value is obtained based at least in part on one or more positions of one or more luma reference samples associated with the maximum luma value. The second chroma value is also obtained based at least in part on one or more positions of one or more luma reference samples associated with the minimum luma value.
[0342] In block 2110, the linear model coefficients are calculated based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value.
[0343] In block 2112, the prediction of the current chroma block is generated based on the linear model coefficients and the values of the downsampled luma blocks of the luma block.
[0344] FIG. 22 is a block diagram showing an exemplary structure of a device 2200 for performing intra prediction using a linear model. The device 2200 may include a determination unit 2202 and an intra prediction processing unit 2204. In one example, the device 2200 may correspond to the intra prediction unit 254 of FIG. 2. In another example, the device 2200 may correspond to the intra prediction unit 354 of FIG. 3.
[0345] The determination unit 2202 is configured to determine a luma block (such as block 811) corresponding to a current chroma block (such as chroma block 801). The determination unit 2202 is further configured to obtain luma reference samples of the luma block based on determining L available chroma reference samples of the current chroma block. The obtained luma reference samples of the luma block are downsampled luma reference samples.
[0346] In some examples, the chroma reference samples of the current chroma block include the reconstructed adjacent samples of the current chroma block. The L available chroma reference samples are determined from the reconstructed adjacent samples. Similarly, the adjacent samples of the luma block are also the reconstructed adjacent samples of the luma block. The obtained luma reference samples of the luma block are obtained by downsampling the reconstructed adjacent samples of the luma block. In some examples, the obtained luma reference samples of the luma block are downsampled luma reference samples obtained by downsampling the reconstructed adjacent samples of the luma block selected based on the L available chroma reference samples. In some examples, the obtained luma reference samples of the luma block are downsampled luma reference samples obtained by downsampling the reconstructed adjacent samples corresponding to the L available chroma reference samples.
[0347] In some examples, when the chroma reference samples are not outside the current image, slice, or tile, and the reference samples are reconstructed and not omitted based on the encoding decision, etc., the chroma reference samples are available. The available reference samples for the current chroma block can be the reconstructed adjacent samples available for the chroma block. The luma reference samples corresponding to the available chroma reference samples are obtained.
[0348] In some examples, the L available chroma reference samples are determined by determining that the L upper adjacent chroma samples of the current chroma block are available, where 1 <= L <= W2, and L and W2 are positive integers. W2 indicates the upper reference sample range, and the L upper adjacent chroma samples are used as the available chroma reference samples. In some examples, W2 is equal to either 2*W or W + H. Here, W represents the width of the current chroma block, and H represents the height of the current chroma block.
[0349] In other examples, the L available chroma reference samples are determined by determining the L left adjacent chroma samples available for the current chroma block. Here, 1 <= L <= H2, and L and H2 are positive integers. H2 indicates the left reference sample range. The L left adjacent chroma samples are used as the available chroma reference samples. In some examples, H2 is equal to either 2*H or W + H. W represents the width of the current chroma block, and H represents the height of the current chroma block.
[0350] In a further example, the L available chroma reference samples are determined by determining the L1 upper adjacent chroma samples and the L2 left adjacent chroma samples available for the current chroma block. Here, 1 <= L1 <= W2, and 1 <= L2 <= H2. W2 indicates the upper reference sample range, and H2 indicates the left reference sample range. L1, L2, W2, and H2 are positive integers, and L1 + L2 = L. In these examples, the L1 upper adjacent chroma samples and the L2 left adjacent chroma samples are used as the available chroma reference samples.
[0351] In one example, the luma reference samples are obtained by downsampling only the adjacent samples that are above the luma block and are selected based on the L available chroma reference samples. In another example, the luma reference samples are obtained by downsampling only the adjacent samples that are to the left of the luma block and are selected based on the L available chroma reference samples.
[0352] In the above example, the downsampled luma block of the luma block is obtained by downsampling the reconstructed luma block of the luma block corresponding to the current chroma block. In some cases, when the luma reference samples are obtained based only on the adjacent samples above the luma block, and when the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), etc., only one row of the reconstructed adjacent luma samples of the reconstructed version of the luma block is used to obtain the luma reference samples.
[0353] The intra prediction processing unit 2204 is configured to calculate linear model coefficients (such as α and β) based on the luma reference samples and the chroma reference samples corresponding to the luma reference samples. The intra prediction processing unit 2204 is further configured to obtain a prediction of the current chroma block based on the linear model coefficients and the values of the downsampled luma block of the luma block.
[0354] Figure 23 is a flowchart of a method for coding a chroma intra coding mode in a bitstream of a video signal according to some aspects of the present disclosure.
[0355] In block 2302, intra prediction of a chroma block of a video signal is performed based on an intra chroma prediction mode. The intra chroma prediction mode can be selected from a plurality of modes. In some examples, the plurality of modes include a first mode set including at least one of a derived mode (DM) or a cross-component linear model (CCLM) prediction mode, a second mode set including at least one of a CCLM_L mode or a CCLM_T mode, or a third mode set including at least one of a vertical mode, a horizontal mode, a DC mode, or a planar mode.
[0356] In block 2304, a bitstream of the video signal is generated by including a syntax element indicating the intra chroma prediction mode in the bitstream. In some examples, the number of bits of the syntax element when the intra chroma prediction mode is selected from the first mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set. The number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the third mode set.
[0357] In one example, the syntax element of the DM mode is 0. The syntax element of the CCLM mode is 10. The syntax element of the CCLM_L mode is 1110. The syntax element of the CCLM_T mode is 1111. The syntax element of the planar mode is 11000. The syntax element of the vertical mode is 11001. The syntax element of the horizontal mode is 11010. The syntax element of the DC mode is 11011.
[0358] In another example, the syntax element for the DM mode is 00. The syntax element for the CCLM mode is 10. The syntax element for the CCLM_L mode is 110. The syntax element for the CCLM_T mode is 111. The syntax element for the planar mode is 0100. The syntax element for the vertical mode is 0101. The syntax element for the horizontal mode is 0110. The syntax element for the DC mode is 0111.
[0359] FIG. 24 is a flowchart of a method for decoding a chroma intra coding mode in a bitstream of a video signal according to some aspects of the present disclosure.
[0360] In block 2402, a plurality of syntax elements are parsed from the bitstream of the video signal. In block 2404, the intra chroma prediction mode is determined based on the syntax element indicating the intra chroma prediction mode from the plurality of syntax elements. In some examples, the intra chroma prediction mode is determined from a plurality of modes. For example, the plurality of modes includes at least one of three sets: a first mode set including at least one of a derived mode (DM) or a cross-component linear model (CCLM) prediction mode, a second mode set including at least one of a CCLM_L mode or a CCLM_T mode, or a third mode set including at least one of a vertical mode, a horizontal mode, a DC mode, or a planar mode. Among these intra chroma prediction mode sets, the number of bits of the syntax element when the intra chroma prediction mode is selected from the first mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set. The number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the third mode set.
[0361] In block 2406, an intra prediction of the current chroma block of the video signal is performed based on the intra chroma prediction mode.
[0362] FIG. 25 is a block diagram showing an exemplary structure of a device 2500 for generating a video bitstream. The device 2500 may include an intra prediction processing unit 2502 and a binarization unit 2504. In one example, the intra prediction processing unit 2502 may correspond to the intra prediction unit 254 in FIG. 2. In one example, the binarization unit 2504 may correspond to the entropy encoding unit 270 in FIG. 2.
[0363] The intra prediction processing unit 2502 is configured to perform intra prediction of chroma blocks of a video signal based on an intra chroma prediction mode. The intra chroma prediction mode is selected from a first mode set, a second mode set, or a third mode set. The first mode set includes at least one of a derived mode (DM) or a cross-component linear model (CCLM) prediction mode. The second mode set includes at least one of a CCLM_L mode or a CCLM_T mode. The third mode set includes at least one of a vertical mode, a horizontal mode, a DC mode, or a plane mode.
[0364] The binarization unit 2504 is configured to generate a bitstream of the video signal by including a syntax element indicating the intra chroma prediction mode. The number of bits of the syntax element when the intra chroma prediction mode is selected from the first mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set, and the number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the third mode set.
[0365] In one example, the syntax element for the DM mode is 0. The syntax element for the CCLM mode is 10. The syntax element for the CCLM_L mode is 1110. The syntax element for the CCLM_T mode is 1111. The syntax element for the planar mode is 11000. The syntax element for the vertical mode is 11001. The syntax element for the horizontal mode is 11010. The syntax element for the DC mode is 11011.
[0366] In another example, the syntax element for the DM mode is 00. The syntax element for the CCLM mode is 10. The syntax element for the CCLM_L mode is 110. The syntax element for the CCLM_T mode is 111. The syntax element for the planar mode is 0100. The syntax element for the vertical mode is 0101. The syntax element for the horizontal mode is 0110. The syntax element for the DC mode is 0111.
[0367] FIG. 26 is a block diagram showing an exemplary structure of a device 2600 for decoding a video bit stream. The device may include an analysis unit 2602, a determination unit 2604, and an intra prediction processing unit 2606. In one example, the analysis unit 2602 may correspond to the entropy encoding unit 304 in FIG. 3. In one example, the determination unit 2604 and the intra prediction processing unit 2606 may correspond to the intra prediction unit 354 in FIG. 3.
[0368] The parsing unit 2602 is configured to parse syntax elements from the bitstream of the video signal. The determination unit 2604 is configured to determine an intra chroma prediction mode based on syntax elements from a plurality of syntax elements. The intra chroma prediction mode is determined from one of a first mode set, a second mode set, or a third mode set. The first mode set includes at least one of a derived mode (DM) or a cross-component linear model (CCLM) prediction mode. The second mode set includes at least one of a CCLM_L mode or a CCLM_T mode. The third mode set includes at least one of a vertical mode, a horizontal mode, a DC mode, or a planar mode.
[0369] The number of bits of the syntax element when the intra chroma prediction mode is selected from the first mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set, and the number of bits of the syntax element when the intra chroma prediction mode is selected from the second mode set is smaller than the number of bits of the syntax element when the intra chroma prediction mode is selected from the third mode set.
[0370] In one example, the syntax element of the DM mode is 0. The syntax element of the CCLM mode is 10. The syntax element of the CCLM_L mode is 1110. The syntax element of the CCLM_T mode is 1111. The syntax element of the planar mode is 11000. The syntax element of the vertical mode is 11001. The syntax element of the horizontal mode is 11010. The syntax element of the DC mode is 11011.
[0371] In another example, the syntax element for the DM mode is 00. The syntax element for the CCLM mode is 10. The syntax element for the CCLM_L mode is 110. The syntax element for the CCLM_T mode is 111. The syntax element for the planar mode is 0100. The syntax element for the vertical mode is 0101. The syntax element for the horizontal mode is 0110. The syntax element for the DC mode is 0111.
[0372] The intra prediction processing unit 2606 is configured to perform intra prediction of the current chroma block of the video signal based on the intra chroma prediction mode.
[0373] The following references are incorporated by reference to better understand the present disclosure: JCTVC-H0544, description of MDLM, JVET-G1001, description of CCLM or LM, Section 2.2.4, and JVET-K0204, description for deriving model coefficients using maximum and minimum values.
[0374] The following is an explanation of the encoding method and decoding method applications shown in the above-described embodiments, and systems using those applications.
[0375] FIG. 27 is a block diagram showing a content supply system 3100 for realizing a content delivery service. This content supply system 3100 includes an intake device 3102 and a terminal device 3106, and optionally includes a display 3126. The intake device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the above-described communication channel 13. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination of these types.
[0376] The capturing device 3102 can generate data and encode the data by the encoding method as shown in the above embodiments. Alternatively, the capturing device 3102 can distribute the data to a streaming server (not shown in the figure), and the server encodes the data and transmits the encoded data to the terminal device 3106. The capturing device 3102 includes, but is not limited to, a camera, a smartphone or a tablet, a computer or a laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capturing device 3102 may include the transmitting device 12 as described above. When the data includes video, the video encoder 20 included in the capturing device 3102 can actually perform video encoding processing. When the data includes sound (i.e., audio), the audio encoder included in the capturing device 3102 can actually perform audio encoding processing. In some practical scenarios, the capturing device 3102 distributes the encoded video and audio data by multiplexing them together. In other practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capturing device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.
[0377] In the content supply system 3100, the terminal device 310 receives and plays back the encoded data. The terminal device 3106 can be a smartphone or a tablet 3108, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof that can decode the above-described encoded data. For example, the terminal device 3106 can include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device preferentially performs video decoding. When the encoded data includes audio, the audio decoder included in the terminal device preferentially performs audio decoding processing.
[0378] In the case of a terminal device including its display, such as a smartphone or a tablet 3108, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can supply the decoded data to its display. In the case of a terminal device not equipped with a display, such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 contacts these devices to receive and display the decoded data.
[0379] When each device of this system performs encoding or decoding, an image encoding device or an image decoding device can be used as shown in the above-described embodiments.
[0380] FIG. 28 is a diagram showing the structure of an example of the terminal device 3106. After the terminal device 3106 receives a stream from the capturing device 3102, the protocol processing unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, Real-Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-Time Transport Protocol (RTP), Real-Time Messaging Protocol (RTMP), or any combination of these types.
[0381] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is transmitted to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0382] Through the demultiplexing process, a video elementary stream (ES), an audio ES, and optional subtitles are generated. The video decoder 3206 including the video decoder 30 as described in the above-described embodiment decodes the video ES by the decoding method as shown in the above-described embodiment to generate a video frame, and supplies this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate an audio frame, and supplies this data to the synchronization unit 3212. Alternatively, the video frame can be stored in a buffer (not shown in FIG. 28) before supplying the data to the synchronization unit 3212. Similarly, the audio frame can be stored in a buffer (not shown in FIG. 28) before supplying the data to the synchronization unit 3212.
[0383] The synchronization unit 3212 synchronizes video frames and audio frames and supplies video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation (performance) of video and audio information. The information can be coded in syntax using time stamps related to the presentation (performance) of the coded audio and visual data and time stamps related to the delivery of the data stream itself.
[0384] When subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes the decoded data with video frames and audio frames, and supplies video / audio / subtitles to the video / audio / subtitle display 3216.
[0385] The present invention is not limited to the above-described system, and either the image encoding device or the image decoding device in the above-described embodiment can be incorporated into another system, for example, an automotive system.
[0386] By way of example, and without limitation, such computer-readable storage media can be used to store the desired program code in the form of instructions or data structures and can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be accessed by a computer. Also, all connections are properly called computer-readable media. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transient, tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disk typically magnetically reproduces data while disc optically reproduces data using a laser. The above combinations should also be included within the scope of computer-readable media.
[0387] The commands can be executed by one or more processors such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, or any other structure suitable for an implementation of the techniques described herein. Further, in some embodiments, the functions described herein may be provided within dedicated hardware and / or software modules configured or combined for encoding and decoding, or incorporated in a codec. Also, these techniques can be fully implemented in one or more circuits or logic elements.
[0388] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chip sets). To emphasize the functional aspects of the apparatuses configured to execute the disclosed techniques, various components, modules, or units are described in this disclosure, but the various components, modules, or units do not necessarily require realization by different hardware units. Rather, as described above, the various units can be provided by a set of interoperable hardware units including one or more of the foregoing processors, combined with codec hardware units, or in combination with appropriate software and / or firmware.
[0389] Although several embodiments are provided in the present disclosure, it should be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. This example should be considered illustrative and not restrictive, and the intention should not be limited to the details given herein. For example, various elements or components can be combined or integrated into another system, a particular function can be omitted or not implemented.
[0390] Furthermore, the techniques, systems, subsystems, and methods described and illustrated in various embodiments as discrete or separate can be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as being coupled or directly coupled or communicating with each other can be indirectly coupled or communicating through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of changes, substitutions, and alterations can be ascertained by those skilled in the art and can be made without departing from the spirit and scope disclosed herein.
Claims
1. A method for encoding a chroma intra prediction mode in a bitstream of video data, the method comprising: Performing an intra prediction on a current chroma block of the video data based on the chroma intra prediction mode, wherein the step of performing the intra prediction on the chroma block based on the chroma intra prediction mode comprises: Determining a luma block corresponding to the current chroma block; Determining L available chroma reference samples of the current chroma block by checking the availability of neighboring chroma samples above the current chroma block, where L is a positive integer; Next, obtaining luma reference samples of the luma block based on the determined L available chroma reference samples of the current chroma block, wherein the obtained luma reference samples of the luma block are downsampled luma reference samples corresponding to the available chroma reference samples and obtained by downsampling neighboring luma samples with respect to the luma block; Calculating linear model coefficients based on the luma reference samples and chroma reference samples corresponding thereto; and Obtaining a prediction of the current chroma block based on the linear model coefficients and the value of the downsampled luma block of the luma block. Generating a bitstream of the video data by including a syntax element indicating the chroma intra prediction mode. Method.
2. A method for decoding video data implemented by a decoding device, the method comprising: Parsing a syntax element from a bitstream, the syntax element indicating a chroma intra prediction mode; Performing intra prediction on the current chroma block of the video data based on the chroma-intra prediction mode, and The step of performing the intra prediction of the chroma block based on the chroma-intra prediction mode includes: Determining a luma block corresponding to the current chroma block; Determining L available chroma reference samples of the current chroma block by checking the availability of upper neighboring chroma samples of the current chroma block, where L is a positive integer; Next, obtaining luma reference samples of the luma block based on the determination of the L available chroma reference samples of the current chroma block, where the obtained luma reference samples of the luma block are downsampled luma reference samples corresponding to the available chroma reference samples and obtained by downsampling neighboring luma samples related to the luma block; Calculating linear model coefficients based on the luma reference samples and chroma reference samples corresponding to the luma reference samples; Obtaining a prediction of the current chroma block based on the linear model coefficients and the values of the downsampled luma blocks of the luma block. Method. **Claim 3** The step of determining the L available chroma reference samples includes: Determining that the L upper neighboring chroma samples of the current chroma block are available by checking the availability of the upper neighboring chroma samples within an upper reference sample range; The method according to claim 1 or 2, wherein 1 < L < W2, W2 represents the upper reference sample range, L and W2 are positive integers, and the L upper neighboring chroma samples are used as available chroma reference samples. **Claim 4** W2 is equal to either 2*W or W+H, where W represents the width of the current chroma block and H represents the height of the current chroma block, the method according to claim 3.
5. The luma reference sample is above the luma block and is obtained by downsampling only the neighboring samples selected based on the determined L available chroma reference samples, the method according to any one of claims 1 to 4.
6. The downsampled luma block of the luma block is obtained by downsampling the reconstructed luma block of the luma block corresponding to the current chroma block, the method according to any one of claims 1 to 5.
7. When the luma reference sample is obtained based only on the neighboring samples above the luma block and the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), only one row of the reconstructed neighboring luma samples of the reconstructed version of the luma block is used to obtain the luma reference sample, the method according to claim 6.
8. The step of calculating the linear model coefficients based on the luma reference sample and the chroma reference sample corresponding to the luma reference sample includes determining a maximum luma value and a minimum luma value based on the luma reference sample; and obtaining a first chroma value based at least in part on the position of the luma reference sample associated with the maximum luma value; and obtaining a second chroma value based at least in part on the position of the luma reference sample associated with the minimum luma value; and A step of calculating a linear model coefficient based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value, the method according to any one of claims 1 to 7.
9. The step of obtaining the first chroma value based at least in part on the position of the luma reference sample associated with the maximum luma value includes the step of obtaining the first chroma value based at least in part on one or more positions of one or more luma reference samples associated with the maximum luma value. The step of obtaining the second chroma value based at least in part on the position of the luma reference sample associated with the minimum luma value includes the step of obtaining the second chroma value based at least in part on one or more positions of one or more luma reference samples associated with the minimum luma value, the method according to claim 8.
10. The chroma intra prediction mode is the CCLLM_T mode or the CCIP_T mode, the method according to any one of claims 1 to 9.
11. An encoder including a processing circuit for executing the method according to any one of claims 1 and 3 to 10.
12. A decoder including a processing circuit for executing the method according to any one of claims 2 to 10.
13. A non-transitory computer-readable medium including program instructions, which when executed by a computer device or a processor, cause the computer device or the processor to execute the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Intra-Frame Prediction and Decoding Methods and Apparatuses for Image Signal
US20140233650A1
Linear model prediction mode with sample accessing for video coding
US20180176594A1