Method and apparatus for intra-prediction

By determining and reconstructing chroma reference samples within specific boundaries, the method addresses coding errors in video compression, enhancing efficiency and compression ratios.

JP7860302B2Active Publication Date: 2026-05-15HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-04-17
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing video compression methods face challenges in determining reference samples for chroma blocks, leading to coding errors due to the absence of corresponding chroma reference samples, which affects coding efficiency.

Method used

The method determines the availability of reference samples by examining chroma reference samples within certain image, slice, or tile boundaries and reconstructing them if necessary, using downsampling techniques to derive linear model coefficients for improved intra-prediction.

Benefits of technology

This approach enhances coding efficiency by reducing coding errors and improving compression ratios without sacrificing image quality, aligning with the demands of modern telecommunications networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007860302000010
    Figure 0007860302000010
  • Figure 0007860302000011
    Figure 0007860302000011
  • Figure 0007860302000012
    Figure 0007860302000012
Patent Text Reader

Abstract

To provide an intra prediction device and method for encoding and decoding an image.SOLUTION: An intra prediction method by using cross component liner prediction mode (CCLM), includes: determining a luma block corresponding to a current chroma block; obtaining luma reference samples of the luma block based on determining L available chroma reference samples of the current chroma block, wherein the obtained luma reference samples of the luma block are down-sampled luma reference samples; calculating linear model coefficients based on the luma reference samples and chroma reference samples that correspond to the luma reference samples; and obtaining a prediction for the current chroma block based on the linear model coefficients and values of a down-sampled luma block of the luma block.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-references to related applications This patent application claims priority to U.S. Provisional Patent Application No. 62 / 742,266 filed on October 5, 2018, U.S. Provisional Patent Application No. 62 / 742,355 filed on October 6, 2018, U.S. Provisional Patent Application No. 62 / 742,275 filed on October 6, 2018, and U.S. Provisional Patent Application No. 62 / 742,356 filed on October 6, 2018. The aforementioned patent applications are incorporated herein by reference in their entirety.

[0002] Embodiments of this disclosure generally relate to the field of video coding, and more specifically to the field of intra-prediction using cross-component linear model prediction (CCLM). [Background technology]

[0003] Even relatively short videos can require a considerable amount of video data to depict, which can create difficulties when streaming or otherwise transmitting data over communication networks with limited bandwidth. Thus, video data is typically compressed before being transmitted over modern telecommunications networks. Video size can also be a concern when storing video in storage, as memory resources may be limited. Video compressors typically use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompressor that decodes the video data. With limited network resources and increasing demands for higher video quality, improved compression and decompression techniques that increase compression ratios with little to no sacrifice of image quality are desired. High Efficiency Video Coding (HEVC), published by the ISO / IEC Video Expert Group and the ITU-T Video Coding Expert Group as ISO / IEC 23008-2 MPEG-H Part 2, also known as ITU-T H.265, is the latest video compression method that approximately doubles the data compression ratio at the same level of video quality, or significantly improves video quality at the same bitrate. [Overview of the project]

[0004] Examples of this disclosure provide intra-predictive devices and methods for encoding and decoding images, which can improve the efficiency of cross-component linear model prediction (CCLM), thereby improving the coding efficiency of video signals. This disclosure is described in detail in the examples and claims contained in this file.

[0005] The aforementioned and other objectives are achieved by the subject matter of the independent claims. Further embodiments are evident from the dependent claims, the description of the specification and the drawings.

[0006] Specific embodiments are outlined in the attached independent claims, and other embodiments are outlined in the dependent claims.

[0007] According to a first aspect, the disclosure relates to a method for performing intraprediction using a linear model. The method includes: determining a luma block corresponding to a current chroma block; obtaining luma reference samples of a luma block based on determining L available chroma reference samples of the current chroma block, wherein the obtained luma reference samples of the luma block are downsampled luma reference samples; calculating linear model coefficients based on the luma reference samples and the chroma reference samples corresponding to the luma reference samples; and obtaining a prediction of the current chroma block based on the linear model coefficients and the downsampled luma block values ​​of the luma block. The chroma reference samples of the current chroma block include reconstructed neighbor samples of the current chroma block. The L available chroma reference samples are determined from the reconstructed neighbor samples. Similarly, the neighbor samples of a luma block are also reconstructed neighbor samples of the luma block (i.e., reconstructed neighbor luma samples). In one example, the acquired luma reference sample for a luma block is obtained by downsampling a reconstructed adjacent luma sample selected based on the available chroma reference sample.

[0008] In existing methods, chroma reference samples are used to determine the availability of reference samples in order to determine linear model coefficients. However, in some scenarios, there is no corresponding chroma reference sample for an available chroma reference sample, which can lead to coding errors. The technique presented herein addresses this problem by determining the availability of reference samples by examining the availability of chroma reference samples. In some cases, a chroma reference sample is available if the chroma reference sample is not outside the current image, slice, or title, and the reference sample has been reconstructed. In some cases, a chroma reference sample is available if the chroma reference sample is not outside the current image, slice, or title, the reference sample has been reconstructed, and the reference sample has not been omitted based on the coding decision. The available reference samples for the current chroma block may be the available reconstructed adjacent samples for the chroma block. The chroma reference samples corresponding to the available chroma reference samples are used to determine the linear model coefficients.

[0009] In a possible embodiment of the method according to the first aspect itself, the step of determining L available chroma reference samples includes determining that L upper adjacent chroma samples of the current chroma block are available, where 1 <= L <= W2, where W2 represents the upper reference sample range, L and W2 are positive integers, and the L upper adjacent chroma samples are used as available chroma reference samples.

[0010] In possible embodiments of the method according to the first embodiment itself or any prior embodiment of the first embodiment, W2 is equal to either 2*W or W+H, where W is the width of the current chroma block and H is the height of the current chroma block.

[0011] In possible embodiments of the method according to the first embodiment itself or any prior embodiment of the first embodiment, the step of determining L available chroma reference samples includes determining that L left-side adjacent chroma samples of the current chroma block are available, where 1 <= L <= H2, where H2 represents the left-side reference sample range, and L and H2 are positive integers, and the L left-side adjacent chroma samples are used as available chroma reference samples.

[0012] In possible embodiments of the method according to the first embodiment itself or any prior embodiment of the first embodiment, H2 is equal to either 2*H or W+H, where W is the width of the current chroma block and H is the height of the current chroma block.

[0013] In possible embodiments of the method according to the first embodiment itself or any preceding embodiment of the first embodiment, the step of determining L available chroma reference samples includes determining that L1 upper adjacent chroma samples and L2 left adjacent chroma samples of the current chroma block are available, where 1 <= L1 <= W2, 1 <= L2 <= H2, where W2 indicates the upper reference sample range and H2 indicates the left reference sample range, where L1, L2, W2, and H2 are positive integers, where L1 + L2 = L, and the L1 upper adjacent chroma samples and L2 left adjacent chroma samples are used as available chroma reference samples.

[0014] In possible embodiments of the method according to the first embodiment itself or any prior embodiment of the first embodiment, the luma reference sample is obtained by downsampling only adjacent samples that are located above the luma block and selected based on L available chroma reference samples, or by downsampling only adjacent samples that are located to the left of the luma block and selected based on L available chroma reference samples. For example, when L is 4, the luma reference sample is obtained by downsampling 24 adjacent samples that are located above the luma block and selected based on 4 available chroma reference samples, or by downsampling 24 adjacent samples that are located to the left of the luma block and selected based on 4 available chroma reference samples, with a 6-tap filter used in the downsampling process.

[0015] In possible embodiments of the method according to the first embodiment itself or any prior embodiment of the first embodiment, the downsampled lumablock of the lumablock is obtained by downsampling a reconfigured lumablock of the lumablock that corresponds to the current chromablock.

[0016] In possible embodiments of the method according to the first embodiment itself or any prior embodiment of the first embodiment, when a luma reference sample is obtained based only on adjacent samples on a luma block, and the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), only one row of the reconstructed adjacent luma sample of the reconfigured version of the luma block is used to obtain the luma reference sample.

[0017] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the step of calculating the linear model coefficients based on the luma reference sample and the chroma reference sample corresponding to the luma reference sample includes: determining a maximum luma value and a minimum luma value based on the luma reference sample; obtaining a first chroma value based at least in part on the position of the luma reference sample associated with the maximum luma value; obtaining a second chroma value based at least in part on the position of the luma reference sample associated with the minimum luma value; and calculating the linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value.

[0018] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the step of obtaining a first chroma value based at least in part on the position of the luma reference sample associated with the maximum luma value includes obtaining the first chroma value based at least in part on one or more positions of one or more luma reference samples associated with the maximum luma value, and the step of obtaining a second chroma value based at least in part on the position of the luma reference sample associated with the minimum luma value includes obtaining the second chroma value based at least in part on one or more positions of one or more luma reference samples associated with the minimum luma value.

[0019] In a possible embodiment of the method according to the first aspect itself or any preceding embodiment of the first aspect, the linear model coefficients α and β are calculated based on the following: α = (y A , B , B , A , A , B , A , B , A , , A , - y A ) / (x B - x A ), β = y A - αx A where x B represents the maximum luma value, y B represents the first chroma value, x A represents the minimum luma value, and y A represents the second chroma value.

[0020] In possible embodiments of the method according to the first embodiment itself or any prior embodiment of the first embodiment, the prediction of the current chromablock is obtained based on the following: Nod C (i,j)=α·rec' L (i,j)+β, Here, pred C (i,j) represents the predicted value of the chroma sample in the current chroma block, and rec' L (i,j) represents the sample value of the corresponding luma sample of the downsampled luma block of the reconstructed luma block.

[0021] In possible embodiments of the method according to the first embodiment itself or any prior embodiment of the first embodiment, the luma reference sample is acquired based only on adjacent samples to the left of the luma block, and when the current chroma block is at the left boundary of the current coding tree unit (CTU), only one row of reconstructed adjacent luma samples of the reconstructed luma block is used to acquire the luma reference sample.

[0022] In possible embodiments of the method according to the first embodiment itself or any prior embodiment of the first embodiment, the linear model includes a multidirectional linear model (MDLM).

[0023] According to a second aspect, the Disclosure relates to a method for performing intraprediction using a linear model. The method includes: determining a luma block corresponding to a current chroma block; obtaining luma reference samples of a luma block based on determining L available chroma reference samples of the current chroma block, wherein the obtained luma reference samples of a luma block are downsampled luma reference samples obtained by downsampling adjacent samples of the luma block (i.e., reconstructed adjacent luma samples) corresponding to the L available chroma reference samples; calculating linear model coefficients based on the luma reference samples and the chroma reference samples corresponding to the luma reference samples; and obtaining a prediction of the current chroma block based on the linear model coefficients and the downsampled luma block values ​​of the luma block. The chroma reference samples of the current chroma block include reconstructed adjacent samples of the current chroma block. The L available chroma reference samples are determined from the reconstructed adjacent samples. Similarly, the adjacent samples of a luma block are also reconstructed adjacent samples of the luma block (i.e., reconstructed adjacent luma samples). The acquired luma reference sample for the luma block is obtained by downsampling a reconstructed adjacent luma sample corresponding to an available chroma reference sample.

[0024] Regarding "reconstructed adjacent chroma samples corresponding to available chroma reference samples," it can be understood that the correspondence between the reconstructed adjacent chroma samples and the available chroma reference samples is not limited to a "one-to-one correspondence," but can also be an "M-to-N correspondence." For example, when a 6-tap filter is used for downsampling, M=24 and N=4.

[0025] According to a third aspect, the present invention relates to an apparatus for encoding video data. The apparatus has a video data memory and a video encoder, the video encoder is configured to: obtain a luma reference sample of a luma block based on (or by determining L available chroma reference samples of the current chroma block), where the obtained luma reference sample of the luma block is a downsampled luma reference sample obtained by downsampling adjacent samples of the luma block (i.e., reconstructed adjacent luma samples) corresponding to the L available chroma reference samples; calculate linear model coefficients of a linear model based on the luma reference sample and the chroma reference sample corresponding to the luma reference sample; and obtain a prediction of the current chroma block based on the linear model coefficient and the values ​​of the downsampled version of the luma block. For example, if there are 4 Ls, the acquired luma reference samples from the luma block are 4 downsampled luma reference samples obtained by downsampling the 24 adjacent samples (i.e., reconstructed adjacent luma samples) of the luma block, corresponding to the 4 available chroma reference samples, and a 6-tap filter is used in the downsampling process.

[0026] According to a fourth aspect, the present invention relates to an apparatus for decoding video data. The apparatus has a video data memory and a video decoder, the video decoder is configured to obtain luma reference samples of a luma block based on determining a luma block corresponding to the current chroma block and determining L available chroma reference samples of the current chroma block, wherein the obtained luma reference samples of the luma block are downsampled luma reference samples; calculate linear model coefficients based on the luma reference samples and the chroma reference samples corresponding to the luma reference samples; and obtain a prediction of the current chroma block based on the linear model coefficients and the downsampled luma block values ​​of the luma block.

[0027] In existing approaches, chroma reference samples are used to determine the availability of reference samples in order to determine linear model coefficients. However, in some scenarios, there is no corresponding chroma reference sample for an available chroma reference sample, which can lead to coding errors. The technique presented herein addresses this problem by determining the availability of reference samples by examining the availability of chroma reference samples. In some cases, a chroma reference sample is available if the chroma reference sample is not outside the current image, slice, or title, and the reference sample has been reconstructed. In some cases, a chroma reference sample is available if the chroma reference sample is not outside the current image, slice, or title, the reference sample has been reconstructed, and the reference sample has not been omitted based on the coding decision. The available chroma reference sample for the current chroma block may be an available reconstructed adjacent sample for the chroma block (i.e., an available reconstructed adjacent chroma sample). The chroma reference sample corresponding to the available chroma reference sample is used to determine the linear model coefficients.

[0028] In possible embodiments of the apparatus according to the third or fourth aspect itself, determining L available chroma reference samples involves determining that L upper adjacent chroma samples of the current chroma block are available, where 1 <= L <= W2, where W2 represents the upper reference sample range, L and W2 are positive integers, and the L upper adjacent chroma samples are used as available chroma reference samples.

[0029] In possible embodiments of the apparatus according to the third or fourth embodiment itself or any prior embodiment of the third or fourth embodiment, W2 is equal to either 2*W or W+H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0030] In possible embodiments of the apparatus according to the third or fourth embodiment itself or any preceding embodiment of the third or fourth embodiment, determining L available chroma reference samples involves determining that L left-side adjacent chroma samples of the current chroma block are available, where 1 <= L <= H2, where H2 represents the left-side reference sample range, L and H2 are positive integers, and the L left-side adjacent chroma samples are used as available chroma reference samples.

[0031] In possible embodiments of the apparatus according to the third or fourth embodiment itself or any prior embodiment of the third or fourth embodiment, H2 is equal to either 2*H or W+H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0032] In possible embodiments of the apparatus according to the third or fourth embodiment itself or any preceding embodiment of the third or fourth embodiment, determining L available chroma reference samples involves determining that L1 upper adjacent chroma samples and L2 left adjacent chroma samples of the current chroma block are available, where 1 <= L1 <= W2, 1 <= L2 <= H2, where W2 indicates the upper reference sample range and H2 indicates the left reference sample range, where L1, L2, W2, and H2 are positive integers, where L1 + L2 = L, and L1 upper adjacent chroma samples and L2 left adjacent chroma samples are used as available chroma reference samples.

[0033] In possible embodiments of the apparatus according to the third or fourth embodiment itself or any prior embodiment of the third or fourth embodiment, the chroma reference sample is obtained by downsampling only adjacent samples selected based on L available chroma reference samples that are located on the chroma block, or by downsampling only adjacent samples selected based on L available chroma reference samples that are located to the left of the chroma block.

[0034] In possible embodiments of the apparatus according to the third or fourth embodiment itself or any prior embodiment of the third or fourth embodiment, the downsampled lumablock of the lumablock is obtained by downsampling a reconfigured lumablock of the lumablock that corresponds to the current chromablock.

[0035] In possible embodiments of the apparatus according to the third or fourth embodiment itself or any prior embodiment of the third or fourth embodiment, when a luma reference sample is obtained based only on adjacent samples on a luma block, and the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), only one row of the reconfigured adjacent luma sample of the reconfigured version of the luma block is used to obtain the luma reference sample.

[0036] In possible embodiments of the apparatus according to the third or fourth embodiment itself or any preceding embodiment of the third or fourth embodiment, the linear model coefficients α and β are calculated based on the following: α=(y B -y A ) / (x B -x A ), β=y A -αx A Here, x B represents the maximum luma value, y B represents the first chroma value, x A represents the minimum luma value, y A This represents the second chroma value.

[0037] In possible embodiments of the apparatus according to the third or fourth embodiment itself or any prior embodiment of the third or fourth embodiment, the current chromablock prediction is obtained based on the following: Nod C (i,j)=α·rec' L (i,j)+β, Here, pred C (i,j) represents the predicted value of the chroma sample in the current chroma block, and rec' L (i,j) represents the sample value of the corresponding luma sample of the downsampled luma block of the reconstructed luma block.

[0038] In possible embodiments of the apparatus according to the third or fourth embodiment itself or any prior embodiment of the third or fourth embodiment, the luma reference sample is obtained based only on adjacent samples to the left of the luma block, and when the current chroma block is at the left boundary of the current coding tree unit (CTU), only one row of reconstructed adjacent luma samples of the reconstructed luma block is used to obtain the luma reference sample.

[0039] In possible embodiments of the apparatus according to the third or fourth aspect itself or any prior embodiment of the third or fourth aspect, the linear model includes a multidirectional linear model (MDLM).

[0040] According to a fifth aspect, the present invention relates to a method for coding an intrachroma prediction mode in a bitstream of a video signal. The method includes: making an intra-prediction of a chroma block of a video signal based on an intrachroma prediction mode, wherein the intrachroma prediction mode is selected from a first mode set, a second mode set including at least one of CCLM_L mode or CCLM_T mode, or a third mode set; and generating a bitstream of a video signal by including a syntax element indicating the intrachroma prediction mode, wherein the number of bits in the syntax element when the intrachroma prediction mode is selected from the first mode set is less than the number of bits in the syntax element when the intrachroma prediction mode is selected from the second mode set, and the number of bits in the syntax element when the intrachroma prediction mode is selected from the second mode set is less than the number of bits in the syntax element when the intrachroma prediction mode is selected from the third mode set.

[0041] The proposed method for encoding intrachroma prediction modes allows CCLM_L and CCLM_T to be represented using binary strings and included in the bitstream of the video signal.

[0042] In possible embodiments of the fifth aspect of the method, the first mode set includes at least one derived mode (DM) or cross-component linear model (CCLM) prediction mode, and the third mode set includes at least one vertical mode, horizontal mode, DC mode, or planar mode.

[0043] In possible embodiments of the fifth embodiment itself or of any prior embodiment of the fifth embodiment, the syntax elements for DM mode are 0; the syntax elements for CCLM mode are 10; the syntax elements for CCLM_L mode are 1110; the syntax elements for CCLM_T mode are 1111; the syntax elements for planar mode are 11000; the syntax elements for vertical mode are 11001; the syntax elements for horizontal mode are 11010; and the syntax elements for DC mode are 11011.

[0044] In possible embodiments of the fifth embodiment itself or of any prior embodiment of the fifth embodiment, the syntax element for DM mode is 00; the syntax element for CCLM mode is 10; the syntax element for CCLM_L mode is 110; the syntax element for CCLM_T mode is 111; the syntax element for planar mode is 0100; the syntax element for vertical mode is 0101; the syntax element for horizontal mode is 0110; and the syntax element for DC mode is 0111.

[0045] According to a sixth aspect, the present invention relates to a method for decoding an intrachroma prediction mode in a bitstream of a video signal. The method includes: analyzing a plurality of syntax elements from a bitstream of a video signal; determining an intrachroma prediction mode based on the syntax elements from the plurality of syntax elements, wherein the intrachroma prediction mode is determined from one of a first mode set, a second mode set including at least one of CCLM_L mode or CCLM_T mode, or a third mode set, wherein the number of bits of the syntax element when the intrachroma prediction mode is selected from the first mode set is less than the number of bits of the syntax element when the intrachroma prediction mode is selected from the second mode set, and the number of bits of the syntax element when the intrachroma prediction mode is selected from the second mode set is less than the number of bits of the syntax element when the intrachroma prediction mode is selected from the third mode set; and performing an intra-prediction of the current chroma block of the video signal based on the intrachroma prediction mode.

[0046] According to a seventh aspect, the present invention relates to an apparatus for encoding video data. The apparatus includes a video data memory and a video encoder, the video encoder is configured to perform intra-prediction of chroma blocks of a video signal based on an intra-chroma prediction mode, wherein the intra-chroma prediction mode is selected from a first mode set, a second mode set including at least one of CCLM_L mode or CCLM_T mode, or a third mode set; and to generate a bitstream of the video signal by including a syntax element indicating the intra-chroma prediction mode, wherein the number of bits in the syntax element when the intra-chroma prediction mode is selected from the first mode set is less than the number of bits in the syntax element when the intra-chroma prediction mode is selected from the second mode set, and the number of bits in the syntax element when the intra-chroma prediction mode is selected from the second mode set is less than the number of bits in the syntax element when the intra-chroma prediction mode is selected from the third mode set.

[0047] According to the eighth aspect, the present invention relates to an apparatus for decoding video data. The device has a video data memory and a video decoder, the video decoder is configured to: analyze a plurality of syntax elements from a bitstream of a video signal; determine an intrachroma prediction mode based on the plurality of syntax elements, wherein the intrachroma prediction mode is determined from one of a first mode set, a second mode set including at least one of CCLM_L mode or CCLM_T mode, or a third mode set, wherein the number of bits of the syntax element when the intrachroma prediction mode is selected from the first mode set is less than the number of bits of the syntax element when the intrachroma prediction mode is selected from the second mode set, and the number of bits of the syntax element when the intrachroma prediction mode is selected from the second mode set is less than the number of bits of the syntax element when the intrachroma prediction mode is selected from the third mode set; and perform an intra-prediction of the current chroma block of the video signal based on the intrachroma prediction mode.

[0048] The proposed apparatus for encoding video data and the apparatus for decoding video data represent CCLM_L and CCLM_T using binary strings, which can be included in the bitstream of the video signal in the encoding apparatus and decoded in the decoding apparatus. In possible embodiments of the apparatus according to the seventh and eighth aspects themselves, the first mode set includes at least one of the derivation mode (DM) or the cross-component linear model (CCLM) prediction mode, and the third mode set includes at least one of the vertical mode, horizontal mode, DC mode, or planar mode.

[0049] In possible embodiments of the apparatus according to the seventh and eighth aspects themselves or any prior embodiments of the seventh and eighth aspects, the syntax elements for DM mode are 0; the syntax elements for CCLM mode are 10; the syntax elements for CCLM_L mode are 1110; the syntax elements for CCLM_T mode are 1111; the syntax elements for planar mode are 11000; the syntax elements for vertical mode are 11001; the syntax elements for horizontal mode are 11010; and the syntax elements for DC mode are 11011.

[0050] In possible embodiments of the methods according to the seventh and eighth aspects themselves or any prior embodiments of the seventh and eighth aspects, the syntax element for DM mode is 00; the syntax element for CCLM mode is 10; the syntax element for CCLM_L mode is 110; the syntax element for CCLM_T mode is 111; the syntax element for planar mode is 0100; the syntax element for vertical mode is 0101; the syntax element for horizontal mode is 0110; and the syntax element for DC mode is 0111.

[0051] According to a ninth aspect, the present invention relates to a method for performing intraprediction using a cross-component linear model (CCLM) prediction mode. The method includes: determining a luma block corresponding to the current chroma block; obtaining luma reference samples of a luma block by downsampling adjacent samples of the luma block, wherein the luma reference samples include only luma reference samples obtained based on adjacent samples above the luma block, or only luma reference samples obtained based on adjacent samples to the left of the luma block; determining a maximum luma value and a minimum luma value based on the luma reference samples; obtaining a first chroma value based at least partially on one or more locations of one or more luma reference samples associated with the maximum luma value; obtaining a second chroma value based at least partially on one or more locations of one or more luma reference samples associated with the minimum luma value; calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; and generating a prediction of the current chroma block based on the linear model coefficients and the values ​​of the downsampled version of the luma block.

[0052] In a possible embodiment of the method according to the ninth aspect itself, the number of chroma reference samples is greater than or equal to the width of the current chroma block or greater than or equal to the height of the current chroma block.

[0053] In possible embodiments of the method according to the ninth aspect itself or any prior embodiment of the ninth aspect, the luma reference sample used in determining the maximum luma value and the minimum luma value is an available luma reference sample of the luma block.

[0054] In possible embodiments of the method according to the ninth aspect itself or any prior embodiment of the ninth aspect, the available luma reference samples for the luma block are determined based on the available chroma reference samples for the current chroma block.

[0055] In possible embodiments of the method according to the ninth aspect itself or any preceding embodiment of the ninth aspect, up to 2*W chroma reference samples are used to derive linear model coefficients when the chroma reference samples are obtained based only on adjacent samples on the chroma block, where W represents the width of the current chroma block.

[0056] In possible embodiments of the method according to the ninth aspect itself or any prior embodiment of the ninth aspect, up to 2*H chroma reference samples are used to derive linear model coefficients when the chroma reference samples are obtained based only on adjacent samples to the left of the chroma block, where H represents the height of the current chroma block.

[0057] In possible embodiments of the method according to the ninth aspect itself or any prior embodiment of the ninth aspect, up to N available chroma reference samples are used to derive linear model coefficients, where N is the sum of W and H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0058] In possible embodiments of the method according to the ninth aspect itself or any prior embodiment of the ninth aspect, the linear model coefficients α and β are calculated based on the following: α=(y B -y A ) / (x B -x A ), β=y A -αx A Here, x B represents the maximum luma value, y B represents the first chroma value, x A represents the minimum luma value, y A This represents the second chroma value.

[0059] In possible embodiments of the method according to the ninth aspect itself or any prior embodiment of the ninth aspect, the prediction of the current chromablock is obtained based on the following: Nod C (i,j)=α·rec'L (i,j)+β, Here, pred C (i,j) represents the predicted value of the chroma sample in the current chroma block, and rec' L (i,j) represents the sample value of the corresponding luma sample in the downsampled version of the reconfigured version of the luma block.

[0060] In possible embodiments of the method according to the ninth aspect itself or any prior embodiment of the ninth aspect, a downsampled version of the luma block is obtained by downsampling a reconfigured version of the luma block that corresponds to the current chroma block.

[0061] In possible embodiments of the method according to the ninth aspect itself or any prior embodiment of the ninth aspect, when a luma reference sample is obtained based only on adjacent samples on a luma block, and the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), or the current chroma block is at the top boundary of the current coding tree unit (CTU), then only one row of the reconfigured adjacent luma sample of the reconfigured version of the luma block is used to obtain the luma reference sample.

[0062] In possible embodiments of the method according to the ninth aspect itself or any prior embodiment of the ninth aspect, the luma reference sample is acquired based only on adjacent samples to the left of the luma block, and when the current chroma block is at the left boundary of the current coding tree unit (CTU), only one row of reconstructed adjacent luma samples of the reconstructed luma block is used to acquire the luma reference sample.

[0063] In possible embodiments of the ninth aspect itself or the method according to any preceding embodiment of the ninth aspect, the CCLM includes a multi-directional linear model (MDLM).

[0064] According to a tenth aspect, the present invention relates to an encoder configured to perform a method according to the ninth aspect itself or any prior embodiment of the ninth aspect.

[0065] According to the eleventh aspect, the present invention relates to a decoder configured to perform a method according to the ninth aspect itself or any prior embodiment of the ninth aspect.

[0066] According to an eleventh aspect, the present invention relates to an intra prediction method using a cross-component linear prediction mode (CCLM). The method includes the steps of: acquiring a reference sample of the current chroma block, wherein the reference sample belongs only to the upper template of the current chroma block; acquiring a maximum chroma value and a minimum chroma value based on the reference sample; acquiring a first chroma value based on the sample location of the maximum chroma value; acquiring a second chroma value based on the sample location of the minimum chroma value; calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum chroma value, and the minimum chroma value; and acquiring a predictor of the current chroma block based on the linear model coefficients, wherein the current chroma block corresponds to the current chroma block.

[0067] In a possible embodiment of the eleventh aspect of the method, the number of reference samples is greater than or equal to the width of the current chroma block.

[0068] Reference samples are available in the 11th aspect itself or in possible embodiments of the method according to any prior embodiment of the 11th aspect.

[0069] In possible embodiments of the method according to the 11th aspect itself or any prior embodiment of the 11th aspect, the method further includes a step of checking the availability of a reference sample within a range, wherein the length of the range is 2*W, or the length of the range is the sum of W and H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0070] In possible embodiments of the method according to the 11th aspect itself or any prior embodiment of the 11th aspect, up to 2*W available reference samples are used to derive linear model coefficients, where W represents the width of the current chroma block.

[0071] In possible embodiments of the method according to the 11th aspect itself or any prior embodiment of the 11th aspect, up to N available reference samples are used to derive linear model coefficients, where N is the sum of W and H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0072] In possible embodiments of the method according to the 11th aspect itself or any prior embodiment of the 11th aspect, the linear model coefficients α and β are calculated based on the following: α=(y B -y A ) / (x B -x A ), β=y A -αx A Here, x B represents the maximum luma value, y B represents the first chroma value, x A represents the minimum luma value, y A This represents the second chroma value.

[0073] In possible embodiments of the method according to the 11th aspect itself or any prior embodiment of the 11th aspect, the predictor of the current chromablock is obtained based on the following: Nod C (i,j)=α·rec L '(i,j)+β, Here, pred C (i,j) represents a chroma sample, rec L (i,j) represents the corresponding reconstructed luma sample.

[0074] In possible embodiments of the method according to the 11th aspect itself or any prior embodiment of the 11th aspect, the number of reference samples is greater than or equal to the size of the current chroma block.

[0075] In possible embodiments of the method according to the 11th aspect itself or any prior embodiment of the 11th aspect, the reference sample is a downsampled Luma sample.

[0076] In possible embodiments of the method according to the 11th aspect itself or any prior embodiment of the 11th aspect, when the current chroma block is at the upper boundary, only one row of the reconstructed adjacent chroma sample is used to obtain the reference sample.

[0077] In possible embodiments of the method according to the 11th aspect itself or any prior embodiment of the 11th aspect, the CCLM is a multi-directional linear model (MDLM), and linear model coefficients are used to obtain the MDLM.

[0078] In possible embodiments of the method according to the 11th aspect itself or any prior embodiment of the 11th aspect, this method is referred to as CCIP_T.

[0079] According to a twelfth aspect, the present invention relates to an intra-prediction method using a cross-component linear prediction mode (CCLM). This method is A step to retrieve a reference sample for the current Luma block, wherein the reference sample belongs only to the left-hand template of the current Luma block; Steps include obtaining the maximum and minimum luma values ​​based on a reference sample; The steps include: obtaining a first chroma value based on the sample position of the maximum chroma value; The steps include: obtaining a second chroma value based on the sample position of the minimum chroma value; The steps include: calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; The steps include obtaining a predictor for the current chroma block based on linear model coefficients, where the current chroma block corresponds to the current chroma block.

[0080] In a possible embodiment of the method according to the twelfth aspect itself, the number of reference samples is greater than or equal to the height of the current chroma block.

[0081] Reference samples are available in the 12th aspect itself or in possible embodiments of the method according to any prior embodiment of the 12th aspect.

[0082] In possible embodiments of the method according to the 12th aspect itself or any prior embodiment of the 12th aspect, the method further includes a step of checking the availability of a reference sample within a range, wherein the length of the range is 2*H, or the length of the range is the sum of W and H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0083] In possible embodiments of the method according to the 12th aspect itself or any prior embodiment of the 12th aspect, up to 2*H available reference samples are used to derive linear model coefficients, where H represents the height of the current chroma block.

[0084] In possible embodiments of the method according to the 12th aspect itself or any prior embodiment of the 12th aspect, up to N available reference samples are used to derive linear model coefficients, where N is the sum of W and H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0085] In possible embodiments of the method according to the 12th aspect itself or any prior embodiment of the 12th aspect, the linear model coefficients α and β are calculated based on the following: α=(y B -y A ) / (x B -x A ), β=y A -αxA Here, x B represents the maximum luma value, y B represents the first chroma value, x A represents the minimum luma value, y A This represents the second chroma value.

[0086] In possible embodiments of the method according to the 12th aspect itself or any prior embodiment of the 12th aspect, the predictor of the current chromablock is obtained based on the following: Nod C (i,j)=α·rec L '(i,j)+β Here, pred C (i,j) represents a chroma sample, rec L (i,j) represents the corresponding reconstructed luma sample.

[0087] In possible embodiments of the method according to the 12th aspect itself or any prior embodiment of the 12th aspect, the number of reference samples is greater than or equal to the size of the current chroma block.

[0088] In possible embodiments of the method according to the 12th aspect itself or any prior embodiment of the 12th aspect, the reference sample is a downsampled Luma sample.

[0089] In possible embodiments of the method according to the 12th aspect itself or any prior embodiment of the 12th aspect, when the current block of the current chroma block is at the left boundary, only one column of the reconstructed adjacent chroma sample is used to obtain the reference sample.

[0090] In possible embodiments of the method according to the 12th aspect itself or any prior embodiment of the 12th aspect, the CCLM is a multi-directional linear model (MDLM), and linear model coefficients are used to obtain the MDLM.

[0091] In possible embodiments of the method according to the 12th aspect itself or any prior embodiment of the 12th aspect, this method is referred to as CCIP_L.

[0092] According to a thirteenth aspect, the present invention relates to an intra-prediction method using a cross-component linear prediction mode (CCLM). This method is A step of obtaining a reference sample for the current Luma block, wherein the reference sample belongs only to the top template of the current Luma block, or only to the left template of the current Luma block; A step to obtain a chroma sample of the current chroma block, wherein the current chroma block corresponds to the current chroma block; The steps include: calculating linear model coefficients based on a reference sample and a chromatic sample; This includes the step of obtaining the predictor for the current chroma block based on linear model coefficients;

[0093] In a possible embodiment of the method according to the 13th aspect itself, up to N reference samples are used to derive linear model coefficients, where N is the sum of W and H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0094] In possible embodiments of the method according to the 13th aspect itself or any prior embodiment of the 13th aspect, the number of reference samples is greater than or equal to the width of the current chroma block, when the reference samples belong only to the upper template of the current chroma block.

[0095] In possible embodiments of the method according to the 13th aspect itself or any prior embodiment of the 13th aspect, up to 2*W reference samples are used to derive linear model coefficients, where W represents the width of the current chroma block, if the reference samples belong only to the upper template of the current chroma block.

[0096] In possible embodiments of the method according to the 13th aspect itself or any prior embodiment of the 13th aspect, the number of reference samples is greater than or equal to the height of the current luma block, when the reference samples belong only to the left-hand template of the current luma block.

[0097] In possible embodiments of the method according to the 13th aspect itself or any prior embodiment of the 13th aspect, up to 2*H reference samples are used to derive linear model coefficients, where W represents the height of the current chroma block, if the reference samples belong only to the left-hand template of the current chroma block.

[0098] In possible embodiments of the method according to the 13th aspect itself or any prior embodiment of the 13th aspect, the number of reference samples is greater than or equal to the size of the current chroma block.

[0099] In possible embodiments of the method according to the 13th aspect itself or any prior embodiment of the 13th aspect, the reference sample is a downsampled Luma sample.

[0100] In possible embodiments of the method according to the 13th aspect itself or any prior embodiment of the 13th aspect, if the reference sample belongs only to the upper template of the current luma block and the current block of the current chroma block is at the upper boundary, then only one row of the reconstructed adjacent luma sample is used to obtain the reference sample.

[0101] In possible embodiments of the method according to the 13th aspect itself or any prior embodiment of the 13th aspect, if the reference sample belongs only to the left-hand template of the current luma block and the current block of the current chroma block is on the left-hand boundary, then only one column of the reconstructed adjacent luma sample is used to acquire the reference sample.

[0102] In possible embodiments of the method according to the 13th aspect itself or any prior embodiment of the 13th aspect, the CCLM is a multi-directional linear model (MDLM), and linear model coefficients are used to obtain the MDLM.

[0103] Reference samples are available in possible embodiments of the method according to the 13th aspect itself or any prior embodiment of the 13th aspect.

[0104] According to the 14th aspect, the present invention relates to a decoder including a processing circuit for performing a method according to the 11th aspect itself or one of the preceding embodiments of the 11th aspect.

[0105] According to the 15th aspect, the present invention relates to a decoder including a processing circuit for performing a method according to either the 12th aspect itself or one of the preceding embodiments of the 12th aspect.

[0106] According to the sixteenth aspect, the present invention relates to a decoder including a processing circuit for performing a method according to either the thirteenth aspect itself or one of the preceding embodiments of the thirteenth aspect.

[0107] According to the seventeenth aspect, the present invention relates to an intra-prediction method using a cross-component linear prediction mode (CCLM). This method is A step to retrieve a reference sample for the current Luma block, wherein the reference sample belongs only to the top template of the current Luma block; Steps include obtaining the maximum and minimum luma values ​​based on a reference sample; The steps include obtaining a first chroma value and a second chroma value based on the maximum luma value and the minimum luma value; The steps include: calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; This includes the step of obtaining the predictor for the current block based on linear model coefficients;

[0108] In a possible embodiment of the method according to the 17th aspect itself, the number of reference samples is greater than or equal to the width of the current chroma block.

[0109] Reference samples are available in possible embodiments of the method according to the 17th aspect itself or any prior embodiment of the 17th aspect.

[0110] In possible embodiments of the method according to the 17th aspect itself or any prior embodiment of the 17th aspect, up to 2*W reference samples are used to derive the model coefficients.

[0111] In possible embodiments of the method according to the 17th aspect itself or any prior embodiment of the 17th aspect, this method is referred to as CCIP_T.

[0112] According to the 18th aspect, the present invention relates to an intra-prediction method using a cross-component linear prediction mode (CCLM). This method is A step to retrieve a reference sample for the current Luma block, wherein the reference sample belongs only to the left-hand template of the current Luma block; Steps include obtaining the maximum and minimum luma values ​​based on a reference sample; The steps include obtaining a first chroma value and a second chroma value based on the maximum luma value and the minimum luma value; The steps include: calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; This includes the step of obtaining the predictor for the current block based on linear model coefficients;

[0113] In a possible embodiment of the method according to the 18th aspect itself, the number of reference samples is greater than or equal to the height of the current chroma block.

[0114] Reference samples are available in possible embodiments of the 18th aspect itself or of the method according to any prior embodiment of the 18th aspect.

[0115] In possible embodiments of the method according to the 18th aspect itself or any prior embodiment of the 18th aspect, up to 2*H reference samples are used to derive the model coefficients.

[0116] In possible embodiments of the method according to the 18th aspect itself or any prior embodiment of the 18th aspect, this method is referred to as CCIP_L.

[0117] According to the 19th aspect, the present invention relates to a decoder for performing a method according to the 17th aspect itself or any preceding embodiment of the 17th aspect.

[0118] According to the 20th aspect, the present invention relates to a decoder for performing a method according to the 18th aspect itself or any prior embodiment of the 18th aspect.

[0119] According to a 21st aspect, the present invention relates to a method for intraprediction using a linear model. This method includes the steps of: obtaining a reference sample of the current luma block; obtaining a maximum luma value and a minimum luma value based on the reference sample; obtaining a first chroma value and a second chroma value based on the location of the luma sample having the maximum luma value and the location of the luma sample having the minimum luma value; calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; and obtaining a predictor of the current block based on the linear model coefficients, wherein the step of obtaining a reference sample of the current luma block includes the step of determining L available chroma template samples of the current luma block, wherein the reference sample of the current luma block is L luma template samples corresponding to L available chroma template samples, or the step of determining L available adjacent chroma samples of the current chroma block, wherein the reference sample of the current luma block is L adjacent luma samples corresponding to L available adjacent chroma samples, where L >= 1 and L is a positive integer.

[0120] In a possible embodiment of the method according to the 21st aspect, the step of determining L available chroma template samples for the current chroma block includes checking the availability of upper adjacent chroma samples for the current chroma block, where if L upper adjacent chroma samples are available, the reference samples for the current chroma block are L adjacent chroma samples corresponding to the L upper adjacent chroma samples, where L>=1 and L<=W2, where W2 represents the upper template sample range, and L and W2 are positive integers.

[0121] In possible embodiments of the method according to the 21st aspect itself or any prior embodiment of the 21st aspect, the step of determining L available chroma template samples for the current chroma block includes checking the availability of left-side adjacent chroma samples for the current chroma block, where if L left-side adjacent chroma samples are available, the reference samples for the current chroma block are L adjacent chroma samples corresponding to the L left-side adjacent chroma samples, where L>=1 and L<=H2, where H2 indicates the left-side template sample range, and L and H2 are positive integers.

[0122] In possible embodiments of the method according to the 21st embodiment itself or any prior embodiment of the 21st embodiment, the step of determining L available chroma template samples for the current chroma block includes checking the availability of upper adjacent chroma samples and left adjacent chroma samples for the current chroma block, such that if L1 upper adjacent chroma samples are available and L2 left adjacent chroma samples are available, the reference sample for the current chroma block includes or comprises L1 adjacent chroma samples corresponding to L1 upper adjacent chroma samples and L2 adjacent chroma samples corresponding to L2 left adjacent chroma samples, where L2>=1 and L2<=H2, where H2 indicates the left template sample range, L and H2 are positive integers, where L1>=1 and L1<=W2, where W2 indicates the upper template sample range, L1 and W2 are positive integers, and L=L1+L2.

[0123] In possible embodiments of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, if L adjacent chromatic samples are available in the template range, then L template chromatic samples and L template chromatic samples are used to obtain model coefficients.

[0124] Reference samples are available in the 21st aspect itself or in possible embodiments of the method according to any prior embodiment of the 21st aspect.

[0125] In possible embodiments of the method according to the 21st aspect itself or any prior embodiment of the 21st aspect, the linear model coefficients α and β are calculated based on the following: α=(y B -y A ) / (x B -x A ), β=y A -αx A Here, x B represents the maximum luma value, y Brepresents the first chroma value, x A represents the minimum luma value, y A This represents the second chroma value.

[0126] In possible embodiments of the method according to the 21st aspect itself or any prior embodiment of the 21st aspect, the predictor of the current chromablock is obtained based on the following: Nod C (i,j)=α·rec L '(i,j)+β, Here, pred C (i,j) represents a chroma sample, rec L (i,j) represents the corresponding reconstructed luma sample.

[0127] In possible embodiments of the method according to the 21st aspect itself or any prior embodiment of the 21st aspect, the number of reference samples is greater than or equal to the size of the current luma block.

[0128] In possible embodiments of the method according to the 21st aspect itself or any prior embodiment of the 21st aspect, the reference sample is a downsampled Luma sample.

[0129] In possible embodiments of the method according to the 21st aspect itself or any prior embodiment of the 21st aspect, when the current block of the current chroma block is at the upper boundary, only one row of the reconstructed adjacent chroma sample is used to obtain the reference sample.

[0130] In possible embodiments of the method according to the 21st aspect itself or any prior embodiment of the 21st aspect, the linear model is a multi-directional linear model (MDLM), and the linear model coefficients are used to obtain the MDLM.

[0131] In possible embodiments of the method according to the 21st aspect itself or any preceding embodiment of the 21st aspect, this method is referred to as CCIP_T or CCIP_L.

[0132] In possible embodiments of the method according to the 21st aspect itself or any prior embodiment of the 21st aspect, the reference sample belongs only to the upper template of the current luma block, or only to the left template of the current luma block, or the reference sample belongs to both the upper template and the left template of the current luma block.

[0133] According to the 22nd aspect, the present invention relates to a decoder including a processing circuit for performing the method according to the 21st aspect itself or a method according to a preceding embodiment of the 21st aspect.

[0134] According to the 23rd aspect, the present invention relates to an encoder including a processing circuit for performing the method according to the 21st aspect itself or a method according to a preceding embodiment of the 21st aspect.

[0135] According to a 24th aspect, the present invention relates to a method for binarizing chroma modes. This method is Steps include: performing intra-prediction using a linear model (multidirectional linear model, MDLM, etc.); A step of generating a bitstream containing multiple syntax elements, wherein the multiple syntax elements indicate or include CCLM mode, CCIP_L mode, or CCIP_T mode;

[0136] In a possible embodiment of the method according to the 24th aspect, The first indicator (77) indicates CCLM mode, and the intra_chroma_pred_mode index is 4. The second indicator (78) indicates CCIP_L mode, and the intra_chroma_pred_mode index is 5. The third indicator (79) indicates CCIP_T mode, and the intra_chroma_pred_mode index is 6.

[0137] In possible embodiments of the method according to the 24th aspect itself or any prior embodiment of the 24th aspect, when sps_cclm_enabled_flag is equal to 1, IntraPredModeC[xCb][yCb] depends on intra_chroma_pred_mode[xCb][yCb] and IntraPredModeY[xCb][yCb].

[0138] According to the 25th aspect, the present invention relates to a method of decoding carried out by a decoding device. A step of analyzing multiple syntax elements from a bitstream, wherein the multiple syntax elements indicate or include CCLM mode, CCIP_L mode, or CCIP_T mode; This includes the step of performing intra-prediction using the given linear model;

[0139] In a possible embodiment of the method according to the 25th aspect, The first indicator (77) indicates CCLM mode, and the intra_chroma_pred_mode index is 4. The second indicator (78) indicates CCIP_L mode, and the intra_chroma_pred_mode index is 5. The third indicator (79) indicates CCIP_T mode, and the intra_chroma_pred_mode index is 6.

[0140] In possible embodiments of the method according to the 25th aspect itself or any prior embodiment of the 25th aspect, when sps_cclm_enabled_flag is equal to 1, IntraPredModeC[xCb][yCb] depends on intra_chroma_pred_mode[xCb][yCb] and IntraPredModeY[xCb][yCb].

[0141] According to the 26th aspect, the present invention relates to a decoder including a processing circuit for performing the methods according to the 24th and 25th aspects themselves or any prior embodiment of the 24th and 25th aspects.

[0142] According to the 27th aspect, the present invention relates to an encoder including a processing circuit for performing the methods according to the 24th and 25th aspects themselves or any prior embodiment of the 24th and 25th aspects.

[0143] According to the 28th aspect, the present invention relates to a computer-readable medium for storing instructions, wherein when an instruction is executed on a processor, the processor is caused to execute the methods according to the 24th and 25th aspects themselves or any prior embodiment of the 24th and 25th aspects.

[0144] According to the 28th aspect, the present invention relates to a decoder. This decoder is One or more processors, The system includes a non-temporary computer-readable storage medium coupled to a processor and storing a program to be executed by the processor, wherein the program, when executed by the processor, configures the decoder to perform the 24th and 25th embodiments themselves or the methods according to any prior embodiments of the 24th and 25th embodiments.

[0145] According to the 28th aspect, the present invention relates to an encoder. This encoder is One or more processors, The encoder includes a non-temporary computer-readable storage medium coupled to a processor and storing a program executed by the processor, wherein the program, when executed by the processor, configures the encoder to perform the 24th and 25th embodiments themselves or a method according to any prior embodiment of the 24th and 25th embodiments.

[0146] According to the 29th aspect of the present invention, the present invention relates to an intra-prediction method using a cross-component linear prediction mode (CCLM). This method is Steps include obtaining a reference sample for the current Luma block; Steps include obtaining the maximum and minimum luma values ​​based on a reference sample; The steps include obtaining a first chroma value and a second chroma value based on the maximum luma value and the minimum luma value; The steps include: calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value; The steps include: obtaining the predictor for the current block based on linear model coefficients; The availability of template samples can be checked by examining adjacent chroma samples.

[0147] According to the 30th aspect, the present invention relates to a decoder for carrying out the method according to the 28th aspect itself.

[0148] According to the 31st aspect, the present invention relates to a decoder for carrying out the method according to the 29th aspect itself.

[0149] According to the 32nd aspect, the present invention relates to a decoder for carrying out a method according to the 28th or 29th aspect.

[0150] According to the 33rd aspect, a device is provided that includes a module / unit / component / circuit for performing at least a portion of the steps of the above method according to any preceding aspect itself or any preceding embodiment of any preceding aspect.

[0151] The apparatus according to the 33rd embodiment can be extended to embodiments corresponding to embodiments of any prior embodiment of the method. Therefore, embodiments of the apparatus include features of the corresponding embodiments of any prior embodiment of the method.

[0152] The advantages of any prior embodiment of the apparatus are the same as the advantages of the corresponding embodiment of the method according to any prior embodiment.

[0153] For clarity, any one of the aforementioned examples may be combined with one or more of the other aforementioned examples to create new examples within the scope of this disclosure.

[0154] These and other features will be more clearly understood from the following detailed description, which will be interpreted in conjunction with the attached drawings and claims. [Brief explanation of the drawing]

[0155] For a more complete understanding of this disclosure, the following brief description, to be interpreted in relation to the accompanying drawings and detailed description, is to be referenced, and similar reference numerals represent similar parts. [Figure 1A] A block diagram illustrating an exemplary coding system that can implement examples of the present invention. [Figure 1B] A block diagram shows another exemplary coding system that can implement an example of the present invention. [Figure 2] A block diagram showing an exemplary video encoder that can carry out examples of the present invention. [Figure 3] This block diagram shows an example of a video decoder that can implement an example of the present invention. [Figure 4] This is a schematic diagram of a video coding device. [Figure 5] This is a simplified block diagram of equipment 500 that can be used as either or both of the source device 12 and destination device 14 shown in Figure 1A, as an exemplary example. [Figure 6A] This is a conceptual diagram showing the nominal relative positions of luma and chromatic samples in the vertical and horizontal directions. [Figure 6B] This is a conceptual diagram showing examples of luma and chroma positions for downsampling luma block samples to generate prediction blocks. [Figure 6C] This is a conceptual diagram showing another example of luma and chroma positions for downsampling luma block samples to generate prediction blocks. [Figure 6D]This is a diagram of the intra-prediction mode for H.265 / HEVC. [Figure 7] This is a diagram of a reference sample for the current block. [Figure 8] This diagram shows the positions of the reference samples on the left and top of the current luma and chroma blocks involved in CCLM mode. [Figure 9] This is a diagram of the straight line between the minimum and maximum luma values. [Figure 10] This is a diagram of the chroma block (including the reference sample) and the corresponding downsampled luma block template. [Figure 11] This is an example diagram of a template that includes an unavailable reference sample. [Figure 12] This is a diagram of a reference sample used in CCLM_T mode. [Figure 13] This is a reference sample diagram used in CCLM_L mode. [Figure 14] This schematic diagram illustrates an example of determining model coefficients using either an upper template larger than the width of the downsampled Lumablock of the current Lumablock, or a left-side template larger than the height of the downsampled Lumablock of the current Lumablock. The upper template contains reference samples located above the downsampled Lumablock, and the left-side template contains reference samples located to the left of the downsampled Lumablock. [Figure 15] This schematic diagram illustrates an example of determining model coefficients using either an upper template with the same size as the width of the downsampled Lumablock of the current Lumablock, or a left-side template with the same size as the height of the downsampled Lumablock of the current Lumablock. [Figure 16] This is a schematic diagram illustrating an example of determining model coefficients for intra-prediction using available reference samples. [Figure 17] This is a schematic diagram illustrating an example of downsampling a luma sample using multiple rows or columns of a neighboring luma sample. [Figure 18]This is a schematic diagram illustrating another example of downsampling a luma sample using multiple rows or columns of a neighboring luma sample. [Figure 19] This is a schematic diagram illustrating an example of downsampling using a single row of adjacent luma samples for luma blocks at the upper boundary of the CTU. [Figure 20] This is a flowchart of a method for performing intra-prediction using a linear model according to some aspects of this disclosure. [Figure 21] This is a flowchart of a method for performing intra-prediction using a linear model according to other aspects of this disclosure. [Figure 22] This is a block diagram illustrating an example structure of equipment for performing intra-prediction using a linear model. [Figure 23] This is a flowchart of a method for coding a chromatic intracoding mode in a bitstream of a video signal according to some aspects of the present disclosure. [Figure 24] This is a flowchart of a method for decoding the chromatic intracoding mode in a bitstream of a video signal according to some aspects of the present disclosure. [Figure 25] A block diagram illustrating an exemplary structure of equipment for generating a video bitstream. [Figure 26] This is a block diagram illustrating the exemplary structure of a device for decoding a video bitstream. [Figure 27] This is a block diagram illustrating an example structure of a content supply system that provides content distribution services. [Figure 28] A block diagram showing the structure of an example terminal device. [Modes for carrying out the invention]

[0156] While one or more exemplary embodiments are provided below, it should be understood from the outset that the systems and / or methods disclosed may be implemented using any number of techniques, whether currently known or existing. This disclosure is not to be limited in any way to the exemplary embodiments, drawings, and techniques shown below, including the exemplary designs and embodiments shown and described herein, and may be modified with their full-range equivalents within the scope of the appended claims.

[0157] Figure 1A is a block diagram illustrating an exemplary coding system 10 that can utilize bidirectional prediction technology. As shown in Figure 1A, the coding system 10 includes a source device 12 that provides encoded video data, which is later decoded by a destination device 14. In particular, the source device 12 can provide the video data to the destination device 14 via a computer-readable medium 16. The source device 12 and destination device 14 may include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source device 12 and destination device 14 may be equipped for wireless communication.

[0158] The destination device 14 can receive and decode the encoded video data via the computer-readable medium 16. The computer-readable medium 16 may include any type of medium or device that can move the encoded video data from the source device 12 to the destination device 14. For example, the computer-readable medium 16 may include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that can help facilitate communication from the source device 12 to the destination device 14.

[0159] In some examples, encoded data may be output to a storage device via the output interface 22. Similarly, encoded data may be accessed from the storage device via the input interface. The storage device may include any of the various distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, digital video disc (DVD), compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or other suitable digital storage media for storing encoded video data. In further examples, the storage device may correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device 12. The destination device 14 may access the stored video data from the storage device via streaming or download. The file server may be any type of server that stores encoded video data and can transmit that encoded video data to the destination device 14. Examples of file servers include a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data via any standard data connection, including an Internet connection. This may include wireless channels (e.g., Wi-Fi connection), wired connections (such as digital subscriber lines (DSL), cable modems, etc.), or a combination of both, suitable for accessing encoded video data stored on a file server. Transmission of encoded video data from storage devices may be via streaming, download, or a combination thereof.

[0160] The technology of this disclosure is not necessarily limited to wireless applications or settings. This technology can be applied to video coding when supporting any of a variety of multimedia applications, such as terrestrial television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission such as dynamic adaptive streaming via HTTP (DASH), digital video encoded to data storage media, decoding of digital video stored on data storage media, or other applications. In some examples, the coding system 10 may be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video conferencing.

[0161] In the example shown in Figure 1A, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The destination device 14 includes an input interface 28, a video decoder 300, and a display device 32. In accordance with this disclosure, the video encoder 200 of the source device 12 and / or the video decoder 300 of the destination device 14 may be configured to apply techniques for bidirectional prediction. In other examples, the source and destination devices may include other components or configurations. For example, the source device 12 may receive video data from an external video source such as an external camera. Similarly, the destination device 14 may interface with an external display device rather than including an integrated display device.

[0162] The coding system 10 shown in Figure 1A is merely an example. The technique for bidirectional prediction can be performed by any digital video coding and / or decoding device. While the technique of this disclosure is generally performed by a video coding device, the technique can also be performed by a video encoder / decoder, typically called a “codec.” Furthermore, the technique of this disclosure can also be performed by a video preprocessor. The video encoder and / or decoder may be a graphics processing unit (GPU) or a similar device.

[0163] The source device 12 and destination device 14 are merely examples of such a coding device that generates coded video data for the source device 12 to transmit to the destination device 14. In some examples, the source device 12 and destination device 14 may operate in a substantially symmetrical manner, such that each of the source and destination devices 12 and 14 includes video coding and decoding components. Thus, the coding system 10 can support one-way or two-way video transmission between the video devices 12 and 14 for, for example, video streaming, video playback, video broadcasting, or videophone.

[0164] The video source 18 of the source device 12 may include video capture devices such as a video camera, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source 18 may generate computer graphics-based data as source video, or as a combination of live video, archived video, and computer-generated video.

[0165] In some cases, when the video source 18 is a video camera, the source device 12 and destination device 14 may form a so-called camera phone or video phone. However, as stated above, the technology described herein may be applicable to video coding in general and to wireless and / or wired applications. In any case, captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video information may then be output onto a computer-readable medium 16 via the output interface 22.

[0166] The computer-readable medium 16 may include temporary media such as wireless broadcasting or wired network transmission, or storage media (i.e., non-temporary storage media) such as hard disks, flash drives, compact discs, digital video discs, Blu-ray discs, or other computer-readable media. In some examples, a network server (not shown) can receive encoded video data from a source device 12 and provide the encoded video data to a destination device 14, for example, via network transmission. Similarly, a computing device in media production equipment such as a disc stamping machine can receive encoded video data from a source device 12 and manufacture a disc containing the encoded video data. Thus, the computer-readable medium 16 can be understood to include one or more computer-readable media in various forms in various examples.

[0167] The input interface 28 of the destination device 14 receives information from the computer-readable medium 16. The information on the computer-readable medium 16 may include syntax information defined by the video encoder 20, which is also used by the video decoder 30, and the syntax information includes syntax elements that describe the characteristics and / or processing of blocks and other coded units, such as groups of photographs (GOPs). The display device 32 displays the decoded video data to the user and may include a cathode ray tube (CRT), liquid crystal display (LDC), plasma display, organic light-emitting diode (OLED) display, or another type of display device.

[0168] The video encoder 200 and video decoder 300 can operate in accordance with video coding standards such as the High Efficiency Video Coding (HEVC) standard currently under development, and can comply with the HEVC Test Model (HM). Alternatively, the video encoder 200 and video decoder 300 can operate in accordance with other company or industry standards such as the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) H.264 standard, or MPEG (Motion Picture Expert Group)-4, Part 10, AVC (Advanced Video Coding), H.265 / HEVC, or extensions of such standards. However, the technology of this disclosure is not limited to any particular coding standard. Other examples of video coding standards include MPEG-2 and ITU-T H.263. Although not shown in Figure 1A, in some embodiments, the video encoder 200 and video decoder 300 may be integrated with an audio encoder and decoder, respectively, and may include a suitable multiplexer-demultiplexer (MUX-DEMUX) unit or other hardware and software to process both audio and video encoding on a common data stream or separate data streams. Where applicable, the MUX-DEMUX unit may comply with the ITU H.223 multiplexer protocol or other protocols such as the User Datagram Protocol (UDP).

[0169] The video encoder 200 and video decoder 300 may each be implemented as one or more suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. Where the technology is partially implemented in software, the device can implement the technology of this disclosure by storing software instructions in a suitable non-temporary computer-readable medium and executing the instructions in hardware using one or more processors. Each of the video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) within the respective device. A device including the video encoder 200 and / or video decoder 300 may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a mobile phone.

[0170] Figure 1B is an exemplary diagram of an exemplary video coding system 40, including the encoder 200 of Figure 2 and / or the decoder 300 of Figure 3 according to an exemplary embodiment. The system 40 can perform the technology of this application, for example, merge estimation with interpretation. In the illustrated embodiment, the video coding system 40 may include an imaging device 41, a video encoder 20, a video decoder 300 (and / or a video coder implemented via the logic circuit 47 of a processing unit 46), an antenna 42, one or more processors 43, one or more memory stores 44, and / or a display device 45.

[0171] As illustrated, the imaging device 41, antenna 42, processing unit 46, logic circuit 47, video encoder 20, video decoder 30, processor 43, memory store 44, and / or display device 45 may communicate with each other. Although both the video encoder 200 and video decoder 30(0) are shown as described, the video coding system 40 may include only the video encoder 200 or only the video decoder 300 in various practical scenarios.

[0172] As shown, in some examples, the video coding system 40 may include an antenna 42. The antenna 42 may be configured, for example, to transmit or receive an encoded bitstream of video data. Furthermore, in some examples, the video coding system 40 may include a display device 45. The display device 45 may be configured to present video data. As shown, in some examples, the logic circuit 47 may be implemented via a processing unit 46. The processing unit 46 may include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. The video coding system 40 may also include an optional processor 43, which may similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. In some examples, the logic circuit 47 may be implemented via hardware, video coding-specific hardware, etc., and the processor 43 may be implemented with general-purpose software, an operating system, etc. Furthermore, the memory store 44 may be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory). In non-restrictive examples, the memory store 44 may be implemented by cache memory. In some examples, the logic circuit 47 may have access to the memory store 44 (for example, for the implementation of an image buffer). In other examples, the logic circuit 47 and / or the processing unit 46 may include a memory store (e.g., a cache) for the implementation of an image buffer, etc.

[0173] In some examples, the video encoder 200 implemented via logic circuits may include an image buffer (e.g., via either a processing unit 46 or a memory store 44) and a graphics processing unit (e.g., via processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video encoder 200 implemented via logic circuits 47 to embody various modules and / or any other encoder systems or subsystems described herein, such as those described with respect to Figure 2. The logic circuits may be configured to perform various operations, as discussed herein.

[0174] The video decoder 300 can be implemented in a manner similar to that implemented via logic circuits 47 to embody various modules and / or any other decoder systems or subsystems described herein, such as those described with respect to the decoder 300 in Figure 3. In some examples, the video decoder 300 may be implemented via logic circuits and may include an image buffer (e.g., via either a processing unit 46 or a memory store 44) and a graphics processing unit (e.g., via the processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video decoder 300 implemented via logic circuits 47 to embody various modules and / or any other decoder systems or subsystems described herein, such as those described with respect to Figure 3.

[0175] In some examples, the antenna 42 of the video coding system 40 may be configured to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data associated with encoding video frames as discussed herein, such as data associated with coding partitions (e.g., conversion coefficients or quantization conversion coefficients, optional indicators (as discussed), and / or data defining coding partitions), indicators, index values, mode selection data, etc. The video coding system 40 may also include a video decoder 300 coupled to the antenna 42 and configured to decode the encoded bitstream. The display device 45 is configured to present video frames.

[0176] Figure 2 is a block diagram showing an example of a video encoder 200 capable of implementing the technology of the present invention. The video encoder 200 can perform intra-coding and inter-coding of video blocks within a video slice. Intra-coding relies on spatial prediction to reduce or eliminate spatial redundancy of video in a given video frame or image. Inter-coding relies on temporal prediction to reduce or eliminate temporal redundancy of video in adjacent frames or images of a video sequence. Intra-mode (I-mode) may refer to any of several spatial-based coding modes. Inter-mode, such as unidirectional prediction (P-mode) or bidirectional prediction (B-mode), may refer to any of several temporal-based coding modes.

[0177] Figure 2 shows a schematic / conceptual block diagram of an exemplary video encoder 200 configured to implement the technology of the present disclosure. In the example of Figure 2, the video encoder 200 includes a residual calculation unit 204, a transformation processing unit 206, a quantization unit 208, an inverse quantization unit 210, and an inverse transformation processing unit 212, a reconstruction unit 210, a buffer 216, a loop filter unit 220, a decoded image buffer (DPB) 230, a prediction processing unit 260, and an entropy coding unit 270. The prediction processing unit 260 may include an inter-estimation unit 242, an inter-prediction unit 244, an intra-estimation unit 252, an intra-prediction unit 254, and a mode selection unit 262. The inter-prediction unit 244 may further include a motion compensation unit (not shown). The video encoder 200 as shown in Figure 2 may also be called a hybrid video encoder or a video encoder with a hybrid video codec.

[0178] For example, the residual calculation unit 204, the transformation processing unit 206, the quantization unit 208, the prediction processing unit 260, and the entropy coding unit 270 form the forward signal path of the encoder 200, while the inverse quantization unit 210, the inverse transformation processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded image buffer (DPB) 230, and the prediction processing unit 260 form the reverse signal path of the encoder, where the reverse signal path of the encoder corresponds to the signal path of the decoder (see decoder 300 in Figure 3).

[0179] The encoder 200 is configured to receive, for example, an image of a series of images forming a video or video sequence, via input 202, image 201, or a block 203 of image 201. The image block 203 may also be called the current image block or the image block to be coded, and image 201 may be called the current image or the image to be coded (particularly in video coding to distinguish the current image from other images (e.g., previously coded and / or decoded images of the same video sequence (i.e., the video sequence which also includes the current image))).

[0180] Partitioning

[0181] Embodiments of encoder 200 may include a partitioning unit (not shown in Figure 2) configured to partition an image 201 into multiple blocks, such as block 203, typically into multiple non-overlapping blocks. The partitioning unit can be configured to use the same block size and corresponding grid defining the block size for all images in the video sequence, or to change the block size between images or subsets or groups of images, and to partition each image into a corresponding block.

[0182] In HEVC and other video coding specifications, a set of coding tree units (CTUs) may be generated to produce an encoded representation of an image. Each CTU may contain a coding tree block for a luminous sample, two corresponding coding tree blocks for a chroma sample, and a syntax structure used to code the samples in the coding tree block. For monochrome images or images with three distinct color planes, a CTU may contain a single coding tree block and a syntax structure used to code the samples in the coding tree block. A coding tree block can be an N×N block of samples. A CTU may also be called a “tree block” or “maximum coding unit” (LCU). HEVC CTUs may be roughly similar to macroblocks in other standards such as H.264 / AVC. However, a CTU is not necessarily limited to a specific size and may contain one or more coding units (CUs). A slice may contain an integer number of CTUs that are sequentially ordered in the raster scan order.

[0183] In HEVC, the CTU is divided into CUs using a quadtree structure, represented as a coding tree, to adapt to various local characteristics. The decision of whether to code an image region using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. A CU may include a coding block for luma samples, two corresponding coding blocks for chroma samples in an image having luma sample arrays, Cb sample arrays, and Cr sample arrays, and a syntax structure used to code the samples in the coding block. For monochrome images or images with three distinct color planes, a CU may include a single coding block and a syntax structure used to code the samples in the coding block. A coding block is an N×N block of samples. In some examples, a CU may be the same size as the CTU. Each CU is coded in one coding mode, which may be, for example, intra-coding mode or inter-coding mode. Other coding modes are also possible. Encoder 200 receives video data. The encoder 200 can encode each CTU within an image slice of video data. As part of the CTU encoding, the predictive processing unit 260 of the encoder 200 or another processing unit (including, but not limited to, the units of the encoder 200 shown in Figure 2) can partition the CTU's CTB into progressively smaller blocks 203. The smaller blocks may be coding blocks for the CUs.

[0184] Syntax data within a bitstream can also define the size of a CTU. A slice contains a number of consecutive CTUs in coding order. A video frame or image or picture can be partitioned into one or more slices. As mentioned above, each tree block can be divided into coding units (CUs) according to a quadtree. Generally, a quadtree data structure contains one node per CU, with the root node corresponding to a tree block (e.g., a CTU). If a CU is divided into four subCUs, the node corresponding to the CU contains four child nodes, each child node corresponding to one of the subCUs. Multiple nodes in a quadtree structure include leaf nodes and non-leaf nodes. A leaf node has no child nodes in the tree structure (i.e., a leaf node cannot be further divided). A non-leaf node contains the root node of the tree structure. For each non-root node of multiple nodes, each non-root node corresponds to a subCU of the CU corresponding to the parent node in the tree structure of that non-root node. Each non-leaf node has one or more child nodes in the tree structure.

[0185] Each node in a quadtree data structure can provide syntax data to its corresponding CU. For example, a node in a quadtree might include a partitioning flag indicating whether the CU corresponding to the node is partitioned into subCUs. The syntax elements of a CU can be defined recursively and may depend on whether the CU is partitioned into subCUs. If a CU is not further partitioned, it is called a leaf CU. If a block of CU is further partitioned, it can generally be called a non-leaf CU. Each level of partitioning is a quadtree partitioned into four subCUs. The black CU is an example of a leaf node (i.e., a block that is not further partitioned).

[0186] A CU serves a similar purpose to a macroblock in the H.264 standard, except that CUs do not have size distinctions. For example, a tree block can be divided into four child nodes (also called subCUs), and then each child node can be divided into another four child nodes, with each child node being the parent node. The last undivided child node, called the leaf node of a quadtree, contains a coding node, also called a leaf CU. Syntax data associated with a coded bitstream can define the maximum number of times the tree block can be divided (called the maximum CU depth) and the minimum size of the coding node. Thus, a bitstream can also define the minimum coding unit (SCU). The term "block" is used to refer to any of CUs, PUs, or TUs in the context of HEVC, or similar data structures in the context of other standards (e.g., macroblocks and their subblocks in H.264 / AVC).

[0187] In HEVC, each CU can be further divided into one, two, or four PUs according to the PU partitioning type. Within a single PU, the same prediction process is applied, and the relevant information is sent to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU partitioning type, the CU can be partitioned into Transform Units (TUs) according to another quadtree structure similar to the coding tree of the CU. One of the key features of the HEVC structure is the concept of multiple partitions, including CUs, PUs, and TUs. PUs can be partitioned so that their shape is non-square. Syntax data associated with a CU may, for example, describe partitioning the CU into one or more PUs. TUs can be square or non-square (e.g., rectangular), and syntactic data associated with a CU may, for example, describe partitioning the CU into one or more TUs according to a quadtree. The partitioning mode may differ depending on whether the CU is encoded in skip or direct mode, intra-predictive mode, or inter-predictive mode.

[0188] VVC (Versatile Video Coding) eliminates the separation of PU and TU concepts while supporting greater flexibility in CU partition shapes. The size of a CU corresponds to the size of the coding node and can be square or non-square (e.g., rectangular). CU sizes can range from 4x4 pixels (or 8x8 pixels) to tree block sizes of up to 128x128 pixels or larger (e.g., 256x256 pixels).

[0189] After encoder 200 has generated prediction blocks for CU (e.g., Luma, Cb, and Cr prediction blocks), encoder 200 can generate residual blocks for CU. For example, encoder 100 can generate Luma residual blocks for CU. Each sample in the Luma residual block for CU represents the difference between a Luma sample in the prediction Luma block for CU and the corresponding sample in the original Luma coding block for CU. Furthermore, encoder 200 can generate Cb residual blocks for CU. Each sample in the Cb residual block for CU may represent the difference between a Cb sample in the prediction Cb block for CU and the corresponding sample in the original Cb coding block for CU. Encoder 100 can also generate Cr residual blocks for CU. Each sample in the Cr residual block for CU may represent the difference between a Cr sample in the prediction Cr block for CU and the corresponding sample in the original Cr coding block for CU.

[0190] In some examples, encoder 100 skips applying the transformation to the transformation block. In such examples, encoder 200 can process the residual sample values ​​in the same way as the transformation coefficients. Thus, in examples where encoder 100 skips applying the transformation, the following descriptions of the transformation coefficients and coefficient block may be applicable to the transformation block of the residual samples.

[0191] After generating a coefficient block (e.g., a Luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the encoder 200 can quantize the coefficient block to potentially reduce the amount of data used to represent the coefficient block and provide further compression. Quantization generally refers to the process of compressing a range of values ​​into a single value. After the encoder 200 has quantized the coefficient block, the encoder 200 can entropy encode the syntax elements representing the quantization transformation coefficients. For example, the encoder 200 can perform context-adaptive binary arithmetic coding (CABAC) or other entropy coding techniques on the syntax elements representing the quantization transformation coefficients.

[0192] The encoder 200 can output a bitstream of encoded image data 271, which includes a sequence of bits that form a representation of the encoded image and associated data. Thus, the bitstream includes the encoded representation of the video data.

[0193] In "Block partitioning structure for next generation video coding" by J. An et al., International Telecommunication Union, COM16-C966, September 2015 (hereinafter referred to as "VCEG proposal COM16-C966"), a quadtree-binary tree (QTBT) partitioning technique was proposed for future video coding standards beyond HEVC. Simulations show that the proposed QTBT structure is more efficient than the HEVC quadtree structure used. In HEVC, inter-prediction of small blocks is restricted to reduce memory access for motion compensation, so bidirectional prediction is not supported for 4x8 and 8x4 blocks, and inter-prediction is not supported for 4x4 blocks. These restrictions are removed in JEM's QTBT.

[0194] In QTBT, a CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is initially partitioned by a quadtree structure. The quadtree leaf nodes can be further partitioned by a binary tree structure. There are two types of binary tree partitioning: horizontal symmetric partitioning and vertical symmetric partitioning. In either case, the node is partitioned by dividing it horizontally or vertically down the middle. The binary tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transformation processing without further partitioning. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. A CU may consist of coding blocks (CBs) of different color components; for example, in the case of P and B slices in a 4:2:0 chroma format, one CU may contain one luma CB and two chroma CBs. A CU may also consist of CBs of a single component; for example, in the case of an I slice, one CU may contain only one luma CB(b) or only two chroma CBs.

[0195] The following parameters are specified for the QTBT partition scheme.

[0196] - CTU size: The size of the root node in a quad tree; the same concept as HEVC.

[0197] - MinQT size: Minimum allowable quadtree leaf node size.

[0198] - MaxBT size: Maximum allowable size of the root node in a binary tree.

[0199] - MaxBT depth: Maximum allowed 2-bit tree depth.

[0200] MinBT size: Minimum allowable 2-minute tree leaf node size.

[0201] In an example of a QTBT partitioning structure, the CTU size is set to 128x128 chromasamples containing two corresponding 64x64 blocks of chromasamples, the MinQT size is set to 16x16, the MaxBT size is set to 64x64, the MinBT size (both width and height) is set to 4x4, and the MaxBT depth is set to 4. A quadtree partition is applied to the CTU to initially generate a quadtree leaf node. A quadtree leaf node can have a size ranging from 16x16 (i.e., MinQT size) to 128x128 (i.e., CTU size). If a quadtree node has a size equal to the MinQT size, no further quadtrees are considered. If a quadtree leaf node is 128x128, it will not be further partitioned by a binary tree because its size exceeds the MaxBT size (i.e., 64x64). Otherwise, a leaf quadtree node can be further partitioned by a binary tree. Therefore, a quad tree leaf node is also the root node of a binary tree, which has a binary tree depth of 0. When the binary tree depth reaches the MaxBT depth (i.e., 4), no further partitioning is considered. No further horizontal partitioning is considered when a binary tree node has a width equal to the MinBT size (i.e., 4). Similarly, no further vertical partitioning is considered when a binary tree node has a height equal to the MinBT size. Binary tree leaf nodes are further processed by prediction and transformation processes without further partitioning. In JEM, the maximum CTU size is 256 × 256 luma samples. Binary tree (CU) leaf nodes can be further processed (e.g., by performing prediction and transformation processes) without further partitioning.

[0202] Furthermore, the QTBT scheme supports the ability for luma and chroma to have separate QTBT structures. Currently, in the case of P-slice and B-slice, the luma CTB and chroma CTB within a single CTU may share the same QTBT structure. However, in the case of I-slice, the luma CTB may be partitioned into CUs by a QTBT structure, and the chroma CTB may be partitioned into chroma CUs by a separate QTBT structure. This means that the CU in an I-slice consists of a coding block for the luma component or a coding block for the two chroma components, while the CU in a P-slice or B-slice consists of a coding block for all three color components.

[0203] The encoder 200 applies a rate-distortion optimization (RDO) process to the QTBT structure to determine the block partitioning.

[0204] Furthermore, a block partitioning structure called Multi-Type Tree (MTT) has been proposed in U.S. Patent Application Publication 2017 / 0208336 to replace QT, BT, and / or QTBT-based CU structures. The MTT partitioning structure is still a recursive tree structure. In MTT, multiple different partitioning structures (three or more, for example) are used. For example, according to MTT technology, three or more different partitioning structures may be used for each non-leaf node of the tree structure at each depth of the tree structure. The depth of a node in the tree structure can refer to the length of the path from the node to the root of the tree structure (e.g., the number of partitions). The partitioning structure can generally refer to how many different blocks a block can be divided into. A partitioning structure may be a quadtree partitioning structure which divides a block into four blocks, a binary tree partitioning structure which divides a block into two blocks, a ternary tree partitioning structure which divides a block into three blocks, and furthermore, a ternary tree partitioning structure which may not divide a block in the middle. A partition structure can have several different partition types. The partition type can further specify how the blocks are divided, and the methods of division include symmetric or asymmetric partitioning, uniform or non-uniform partitioning, and / or horizontal or vertical partitioning.

[0205] In MTT, at each depth of the tree structure, the encoder 200 may be configured to further partition the subtree using a specific partition type from among three or more partitioning structures. For example, the encoder 100 may be configured to determine a specific partition type from QT, BT, ternary tree (TT), and other partitioning structures. In one example, the QT partitioning structure may include a square quadtree or a rectangular quadtree partitioning type. The encoder 200 can partition a square block using a square quadtree partitioning by dividing the block into four equally sized square blocks along both the horizontal and vertical centers. Similarly, the encoder 200 can partition a rectangular (e.g., non-square) block using a rectangular quadtree partitioning by dividing the rectangular block into four equally sized rectangular blocks along both the horizontal and vertical centers.

[0206] The BT partition structure may include at least one of the following partition types: horizontally symmetric binary tree, vertically symmetric binary tree, horizontally asymmetric binary tree, or vertically asymmetric binary tree. In the case of the horizontally symmetric binary tree partition type, encoder 200 may be configured to divide a block horizontally along the center of the block into two symmetric blocks of the same size. In the case of the vertically symmetric binary tree partition type, encoder 200 may be configured to divide a block vertically along the center of the block into two symmetric blocks of the same size. In the case of the horizontally asymmetric binary tree partition type, encoder 100 may be configured to divide a block horizontally into two blocks of different sizes. For example, similar to the PART_2N×nU or PART_2N×nD partition type, one block may be 1 / 4 the size of the parent block and the other block may be 3 / 4 the size of the parent block. In the case of the vertically asymmetric binary tree partition type, encoder 100 may be configured to divide a block vertically into two blocks of different sizes. For example, similar to the PART_nL×2N or PART_nR×2N partition types, one block may be 1 / 4 the size of the parent block, and the other block may be 3 / 4 the size of the parent block. In another example, an asymmetric binary tree partition type can divide the parent block into fractions of different sizes. For example, one subblock may be 3 / 8 the size of the parent block, and the other subblock may be 5 / 8 the size of the parent block. Naturally, such partition types can be either vertical or horizontal.

[0207] The TT partition structure differs from the partitioning of QT or BT structures in that it does not divide blocks along the center. The central region of a block remains together in the same subblock. Unlike QT, which produces four blocks, or binary trees, which produce two blocks, partitioning according to the TT partition structure produces three blocks. Examples of partition types according to the TT partition structure include symmetric partition types (both horizontal and vertical) and asymmetric partition types (both horizontal and vertical). Furthermore, symmetric partition types according to the TT partition structure can be irregular / non-uniform or regular / uniform. Asymmetric partition types according to the TT partition structure are irregular / non-uniform. For example, a TT partition structure may contain at least one of the following partition types: horizontally regular / uniform symmetric ternary, vertically regular / uniform symmetric ternary, horizontally irregular / non-uniform symmetric ternary, vertically irregular / non-uniform symmetric ternary, horizontally irregular / non-uniform asymmetric ternary, or vertically irregular / non-uniform asymmetric ternary partition types.

[0208] Generally, an irregular / heterogeneously symmetric ternary tree partition type is one in which the partition is symmetric with respect to the centerline of the blocks, but at least one of the resulting three blocks is not the same size as the other two. One preferred example is when the side blocks are 1 / 4 the size of the block and the center block is 1 / 2 the size of the block. A regular / heterogeneously symmetric ternary tree partition type is one in which the partition is symmetric with respect to the centerline of the blocks, and all the resulting blocks are the same size. Such partitions are possible when the height or width of the blocks is a multiple of 3, depending on the vertical or horizontal partitioning. An irregular / heterogeneously asymmetric ternary tree partition type is one in which the partition is not symmetric with respect to the centerline of the blocks, and at least one of the resulting blocks is not the same size as the other two.

[0209] In an example where a block (e.g., in a subtree node) is partitioned into an asymmetric ternary partition type, the encoder 200 and / or decoder 300 impose a restriction that two of the three partitions must be the same size. Such a restriction may correspond to a restriction that the encoder 200 must adhere to when encoding video data. Furthermore, in some examples, the encoder 200 and decoder 300 may impose a restriction that, when partitioning according to an asymmetric ternary partition type, the sum of the areas of two partitions is equal to the area of ​​the remaining partition.

[0210] In some examples, the encoder 200 may be configured to select from all of the partition types described above for each of the QT, BT, and TT partition structures. In other examples, the encoder 200 may be configured to determine only a partition type from a subset of the partition types described above. For example, a subset of the partition types (or other partition types) discussed above may be used for a particular block size or a particular depth of a quadtree structure. The supported subset of partition types may be signaled in the bitstream for use by the decoder 200, or it may be predefined so that the encoder 200 and decoder 300 can determine the subset without signaling.

[0211] In other examples, the number of supported partition types may be fixed for all depths of all CTUs. That is, the encoder 200 and decoder 300 may be pre-configured to use the same number of partition types for any depth of the CTU. In other examples, the number of supported partition types may vary and depend on the depth, slice type, or other previously coded information. In one example, at tree structure depths of 0 or 1, only the QT partition structure is used. At depths greater than 1, the QT, BT, and TT partition structures can each be used.

[0212] In some examples, the encoder 200 and / or decoder 300 can apply pre-configured constraints to supported partition types to avoid overlapping partitioning of specific regions of a video image or regions of a CTU. For example, when a block is partitioned with an asymmetric partition type, the encoder 200 and / or decoder 300 may be configured not to further partition the largest subblock that is partitioned from the current block. For example, when a square block is partitioned according to an asymmetric partition type (similar to the PART_2N×nU partition type), the largest subblock among all subblocks (similar to the largest subblock of the PART_2N×nU partition type) is the noted leaf node and cannot be further partitioned. However, smaller subblocks (similar to the smaller subblocks of the PART_2N×nU partition type) can be further partitioned.

[0213] Another example of how constraints can be applied to supported partition types to avoid overlapping partitioning of a particular region is that when a block is partitioned with an asymmetric partition type, the largest subblock partitioned from the current block cannot be further partitioned in the same direction. For example, when a square block is partitioned with an asymmetric partition type (similar to the PART_2N×nU partition type), the encoder 200 and / or decoder 300 may be configured not to partition the largest subblock among all subblocks (similar to the largest subblock of the PART_2N×nU partition type) horizontally.

[0214] To avoid further partitioning difficulties, another example of applying constraints to supported partitioning types is that the encoder 200 and / or decoder 300 may be configured not to partition a block horizontally or vertically if the block width / height is not a power of 2 (for example, if the width / height is not 2, 4, 8, 16, etc.).

[0215] The above example illustrates how encoder 200 can be configured to perform MTT partitioning. Decoder 300 can then apply the same MTT partitioning performed by encoder 200. In some examples, how the video data images are partitioned by encoder 200 can be determined by decoder 300 applying the same set of predefined rules. However, in many situations, encoder 200 can determine a specific partition structure and partition type to use based on a rate-distortion criterion for a particular image of the video data being coded. Thus, for decoder 300 to determine the partitioning for a particular image, encoder 200 can signal syntax elements in the encoded bitstream that indicate how the image and the CTU of the image should be partitioned. Decoder 200 can parse such syntax elements and partition the image and CTU accordingly.

[0216] In one example, the prediction processing unit 260 of the video encoder 200 can be configured to perform any combination of the above partitioning techniques, particularly for motion estimation, which will be explained in detail later.

[0217] Similar to image 201, block 203 can also be considered a two-dimensional array or matrix of samples having intensity values ​​(sample values), although it has fewer dimensions than image 201. In other words, block 203 may include, for example, one sample array (e.g., a luminar array in the case of monochrome image 201), or three sample arrays (e.g., a luminar array and two chromar arrays in the case of color image 201), or any other number and / or type of arrays depending on the applied color format. The number of samples in the horizontal and vertical (or axial) directions of block 203 defines the size of block 203.

[0218] An encoder 200 as shown in FIG. 2 is configured to encode an image 201 block by block. For example, encoding and prediction are performed for each block 203.

[0219] Residual calculation

[0220] Based on the image block 203 and the prediction block 265 (more details about the prediction block 265 will be provided later), the residual calculation unit 204 calculates the residual block 205, for example, for each sample (for each pixel), by subtracting the sample value of the prediction block 265 from the sample value of the image block 203, so as to obtain the residual block 205 in the sample domain.

[0221] Transformation

[0222] The transformation processing unit 206 is configured to apply a transformation, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain the transformation coefficients 207 in the transform domain. The transformation coefficients 207 can also be called transform residual coefficients and represent the residual block 205 in the transform domain.

[0223] The conversion processing unit 206 may be configured to apply an integer approximation of the DCT / DST, such as the conversion specified for HEVC / H.265. Compared to the orthogonal DCT conversion, such an integer approximation is typically scaled by a specific coefficient. An additional scaling coefficient is applied as part of the conversion process to maintain the norm of the residual blocks processed by the forward and reverse conversions. The scaling coefficient is typically selected based on specific constraints such as a scaling coefficient that is a power of 2 of the shift operation, the bit depth of the conversion coefficient, and the trade-off between precision and implementation cost. For example, a specific scaling coefficient may be specified, for example, for the reverse conversion by the reverse conversion processing unit 212 in the decoder 300 (and the corresponding reverse conversion, for example, by the reverse conversion processing unit 212 in the decoder 300), and the corresponding scaling coefficient for the forward conversion by the conversion processing unit 206 in the encoder 200 may be specified accordingly.

[0224] Quantization

[0225] The quantization unit 208 is configured to obtain quantized conversion coefficients 209 by quantizing the conversion coefficients 207 and applying, for example, scalar quantization or vector quantization. The quantized conversion coefficients 209 may also be called quantized residual coefficients 209. The quantization process can reduce the bit depth associated with some or all of the conversion coefficients 207. For example, an n-bit conversion coefficient is truncated to an m-bit conversion coefficient during quantization, where n is greater than m. The degree of quantization can be changed by adjusting the quantization parameter (QP). For example, in the case of scalar quantization, various scalings can be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, and a larger quantization step size corresponds to coarser quantization. Applicable quantization step sizes may be indicated by the quantization parameter (QP). The quantization parameter may be, for example, an index to a predetermined set of applicable quantization step sizes. For example, small quantization parameters may correspond to fine quantization (small quantization step size), large quantization parameters may correspond to coarse quantization (large quantization step size), or vice versa. Quantization may include division by the quantization step size and the corresponding inverse inverse quantization, for example, multiplication by the quantization step size by inverse quantization 210. Some standards, e.g., embodiments of HEVC, can be configured to determine the quantization step size using the quantization parameters. Generally, the quantization step size can be calculated based on the quantization parameters using a fixed-point approximation of the equations involving division. Additional scaling factors may be introduced into quantization and inverse quantization to restore the norm of the residual block, which may be modified for the scaling used in the fixed-point approximation of the equations for quantization step size and quantization parameters. In one example embodiment, scaling of the inverse transform and inverse quantization can be combined. Alternatively, a customized quantization table can be used to signal from encoder to decoder, for example, in a bitstream. Quantization is an irreversible operation, and the loss increases as the quantization step size increases.

[0226] The inverse quantization unit 210 obtains the inverse quantization coefficients 211 by applying the inverse quantization of the quantization unit 208 to the quantization coefficients, based on or using the same quantization step size as the quantization unit 208, for example, by applying the inverse of the quantization scheme applied by the quantization unit 208. The inverse quantization coefficients 211 may also be called the inverse quantized residual coefficients 211, and correspond to the transformation coefficients 207, although they are typically not identical to the transformation coefficients due to quantization losses.

[0227] The inverse transformation processing unit 212 is configured to obtain an inverse transformation block 213 in the sample domain by applying the inverse transformation of the transformation applied by the transformation processing unit 206, for example, the inverse discrete cosine transform (DCT) or the inverse discrete sine transform (DST). The inverse transformation block 213 may also be called the inverse transformation inverse quantization block 213 or the inverse transformation residual block 213.

[0228] The reconstruction unit 214 (e.g., the adder 214) is configured to obtain the reconstructed block 215 in the sample domain by adding the inverse transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example by adding the sample values ​​of the reconstructed residual block 213 to the sample values ​​of the prediction block 265.

[0229] Optionally, a buffer unit 216 (abbreviated as "buffer" 216), for example, the buffer unit 216 is configured to buffer or store the reconstructed block 215 and its respective sample values, for example, for intra-prediction. In a further embodiment, the encoder may be configured to use the unfiltered reconstructed block and / or its respective sample values ​​stored in the buffer unit 216 for any kind of estimation and / or prediction, for example, intra-prediction.

[0230] Embodiments of the encoder 200 may be configured such that, for example, the buffer unit 216 is used not only for storing the reconstructed block 215 for the intra-prediction 254 but also for the loop filter unit 220 (not shown in Figure 2), and / or for example, the buffer unit 216 and the decoded image buffer unit 230 form a single buffer. Further embodiments may be configured to use the filtered block 221 and / or block or sample (neither shown in Figure 2) from the decoded image buffer 230 as input or basis for the intra-prediction 254.

[0231] The loop filter unit 220 (abbreviated as "loop filter" 220) is configured to filter the reconstructed block 215 to obtain the filtered block 221, for example, to smooth pixel transitions or to improve video quality in other ways. The loop filter unit 220 is intended to represent one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or other filters, e.g., a bilateral filter or adaptive loop filter (ALF), or a sharpening or smoothing filter or a co-filter. Although the loop filter unit 220 is shown as an in-loop filter in Figure 2, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be called the filtered reconstructed block 221. The decoding image buffer 230 can store the reconstructed coding block after the loop filter unit 220 has performed a filtering operation on the reconstructed coding block.

[0232] Embodiments of the encoder 200 (each a loop filter unit 220) may be configured to output loop filter parameters (such as sample adaptive offset information) directly or entropically encoded via, for example, an entropy coding unit 270 or any other entropy coding unit, so that, for example, a decoder 300 can receive and apply the same loop filter parameters for decoding.

[0233] The decoded image buffer (DPB) 230 may be a reference image memory that stores reference image data for use in encoding video data by the video encoder 20. The DPB 230 may be formed by any of various memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The DPB 230 and buffer 216 may be provided by the same memory device or separate memory devices. In some examples, the decoded image buffer (DPB) 230 is configured to store filtered blocks 221. The decoded image buffer 230 may be further configured to store other previously filtered blocks of the same current image or a different image (e.g., a previously reconstructed image), e.g., previously reconstructed and filtered blocks 221, and may provide, for example, a fully previously reconstructed, i.e., decoded image (and corresponding reference blocks and samples) and / or a partially reconstructed current image (and corresponding reference blocks and samples) for interpretation. In some examples, when the reconstructed block 215 is reconstructed but there is no in-loop filtering, the decoded image buffer (DPB) 230 is configured to store the reconstructed block 215.

[0234] The prediction processing unit 260, also called the block prediction processing unit 260, is configured to receive block 203 (the current block 203 of the current image 201) and reconstructed image data, for example, reference samples of the same (current) image from buffer 216 and / or reference image data 231 from one or more previously decoded images from the decoded image buffer 230, and to process such data for prediction, i.e., to provide a prediction block 265 (which may be an inter-prediction block 245 or an intra-prediction block 255).

[0235] The mode selection unit 262 may be configured to select a prediction mode (e.g., intra or inter-prediction mode) and / or the corresponding prediction block 245 or 255 to be used as the prediction block 265 for calculating the residual block 205 and for reconstructing the reconstructed block 215.

[0236] Embodiments of the mode selection unit 262 may be configured to select a prediction mode (for example, from modes supported by the prediction processing unit 260) that provides the best match, in other words, minimum residual (minimum residual means better compression for transmission or storage), or minimum signaling overhead (minimum signaling overhead means better compression for transmission or storage), or considers or balances both. The mode selection unit 262 may be configured to determine the prediction mode based on rate-distortion optimization (RDO), i.e., to provide minimum rate-distortion optimization, or to select a prediction mode in which the associated rate distortion satisfies at least the prediction mode selection criteria.

[0237] The following describes in more detail the prediction process performed by the exemplary encoder 200 (e.g., the prediction processing unit 260 and mode selection (e.g., by the mode selection unit 262)).

[0238] As described above, the encoder 200 is configured to determine or select the best or optimal prediction mode from a set of (predetermined) prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.

[0239] The intra prediction mode set may include 35 different intra prediction modes, such as non-directional modes like DC (i.e., average) mode and planar mode, or directional modes (such as those defined in H.265), or may include 67 different intra prediction modes, non-directional modes like DC (i.e., average) mode and planar mode, or directional modes (such as those defined in the yet-to-be-released H266).

[0240] A series of (or possible) inter prediction modes depend on available reference images (i.e., at least previously partially decoded images, such as images stored in DBP230) and other inter prediction parameters. For example, whether the entire reference image or only a part of the reference image (e.g., the search window area around the current block area of the reference image) is used to search for the most matching reference block, and / or, for example, whether pixel interpolation, such as half / semi - pel and / or quarter - pel interpolation, is applied or not.

[0241] In addition to the above prediction modes, a skip mode and / or a direct mode can be applied.

[0242] The prediction processing unit 260 may be further configured to repeatedly use, for example, quadtree partition (QT), binary tree partition (BT), ternary tree partition (TT), or any combination thereof, to partition block 203 into smaller block partitions or sub - blocks and perform respective predictions for each of the block partitions or sub - blocks. The mode selection includes the selection of the tree structure of the partitioned block 203 and the prediction mode applied to each of the block partitions or sub - blocks.

[0243] The interprediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (not shown in Figure 2). The motion estimation unit is configured to receive or acquire an image block 203 (the current image block 203 of the current image 201) and a decoded image 331, or at least one or more previously reconstructed blocks, for example, a reconstructed block of one or more other / different previously decoded images 331 for motion estimation. For example, a video sequence may include the current image and a previously decoded image 331, in other words, the current image and a previously decoded image 331 are part of, or may form part of, a sequence of images that make up the video sequence. The encoder 200 may be configured, for example, to select a reference block from multiple reference blocks of the same or different images among multiple other images, and to provide a reference image (or reference image index, ...) and / or offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as an interprediction parameter to the motion estimation unit (not shown in Figure 2). This offset is also called a motion vector (MV). Merging is a key motion estimation tool used in HEVC and carried over to VCC. To perform merge estimation, a merge candidate list must first be created, where each candidate includes information on whether one or two reference image lists are used, as well as all motion data, including the reference index and motion vectors for each list. The merge candidate list is created based on a. up to four spatial merge candidates derived from five spatially adjacent blocks, b. one temporary merge candidate derived from two blocks temporarily placed in the same location, and c. additional merge candidates including the combined biprediction candidate and zero motion vector candidate.

[0244] The intra-prediction unit 254 is further configured to make decisions based on intra-prediction parameters, such as the selected intra-prediction mode, and the intra-prediction block 255. In any case, after selecting the intra-prediction mode for the block, the intra-prediction unit 254 is also configured to provide information indicating the selected intra-prediction mode for the block to the intra-prediction parameters, i.e., to the entropy coding unit 270. In one example, the intra-prediction unit 254 may be configured to perform any combination of intra-prediction techniques described later.

[0245] The entropy coding unit 270 is configured to obtain encoded image data 21 that can be output by output 272, for example, in the form of an encoded bitstream 21, by applying an entropy coding algorithm or scheme (e.g., variable-length coding (VLC) scheme, context-adaptive VLC scheme (CALVC), arithmetic coding scheme, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partition entropy (PIPE) coding, or another entropy coding method or technique) individually or together (or in whole) to the quantized residual coefficients 209, inter-prediction parameters, intra-prediction parameters, and / or loop filter parameters. The encoded bitstream 271 can be transmitted to the video decoder 30 or archived by the video decoder 30 for later transmission or retrieval. The entropy coding unit 270 may be further configured to entropy code other syntax elements of the current video slice being coded.

[0246] Other structural variations of the video encoder 200 may be used to encode video streams. For example, a non-conversion-based encoder 200 can directly quantize the residual signal for a given block or frame without a conversion processing unit 206. In another embodiment, the encoder 200 can combine a quantization unit 208 and an inverse quantization unit 210 into a single unit.

[0247] Figure 3 shows an exemplary video decoder 300 configured to implement the technology of the present invention. The video decoder 300 is configured to receive encoded image data (e.g., encoded bitstream) 271 encoded by, for example, an encoder 200, and to obtain a decoded image 331. During the decoding process, the video decoder 300 receives from the video encoder 200 an encoded video bitstream representing video data, e.g., an encoded video slice and an image block of associated syntax elements.

[0248] In the example in Figure 3, the decoder 300 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transformation unit 312, a reconstruction unit 314 (e.g., an adder 314), a buffer 316, a loop filter 320, a decoded image buffer 330, and a prediction processing unit 360. The prediction processing unit 360 may include an inter-prediction unit 344, an intra-prediction unit 354, and a mode selection unit 362. In some examples, the video decoder 300 performs a decoding path that is roughly the reverse of the encoding path described for the video encoder 200 in Figure 2.

[0249] The entropy decoding unit 304 is configured to perform entropy decoding on the encoded image data 271 to obtain, for example, quantization coefficients 309 and / or decoded coding parameters (not shown in Figure 3), such as (decoded) inter-prediction parameters, intra-prediction parameters, loop filter parameters, and / or other syntax elements, or any or all of them. The entropy decoding unit 304 is further configured to transfer the inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the prediction processing unit 360. The video decoder 300 can receive the syntax elements at the video slice level and / or video block level.

[0250] The inverse quantization unit 310 may have the same function as the inverse quantization unit 110, the inverse transformation processing unit 312 may have the same function as the inverse transformation processing unit 112, the reconstruction unit 314 may have the same function as the reconstruction unit 114, the buffer 316 may have the same function as the buffer 116, the loop filter 320 may have the same function as the loop filter 120, and the decoded image buffer 330 may have the same function as the decoded image buffer 130.

[0251] The prediction processing unit 360 may include an inter-prediction unit 344 and an intra-prediction unit 354, where the inter-prediction unit 344 may function similarly to the inter-prediction unit 1 / 44, and the intra-prediction unit 354 may function similarly to the intra-prediction unit 154. The prediction processing unit 360 is typically configured to perform block prediction and / or obtain prediction blocks 365 from encoded data 21, and to receive or obtain (explicitly or implicitly) information regarding prediction-related parameters and / or a selected prediction mode from, for example, an entropy decoding unit 304.

[0252] When a video slice is coded as an intra-coded (I) slice, the intra-prediction unit 354 of the prediction processing unit 360 is configured to generate a prediction block 365 for the image block of the current video slice based on the signaled intra-prediction mode and data from a previously decoded block of the current frame or image. When a video frame is coded as an inter-coded (i.e., B, or P) slice, the inter-prediction unit 344 (e.g., motion compensation unit) of the prediction processing unit 360 is configured to generate a prediction block 365 for the video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. For inter-prediction, the prediction block may be generated from one of the reference images in one of the reference image lists. The video decoder 300 can configure the reference frame lists, list 0 and list 1, using default configuration techniques based on reference images stored in the DPB 330.

[0253] The prediction processing unit 360 is configured to determine prediction information for the video blocks of the current video slice by analyzing motion vectors and other syntax elements, and uses the prediction information to generate prediction blocks for the current video block being decoded. For example, the prediction processing unit 360 uses some of the received syntax elements to determine the prediction mode used to code the video blocks of the video slice (e.g., intra or interpredict), the interpredict slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more reference image lists of the slice, the motion vector for each intercoded video block of the slice, the interpredict status for each intercoded video block of the slice, and other information for decoding the video blocks of the current video slice.

[0254] The inverse quantization unit 310 is provided in a bitstream and is configured to inverse quantize, or dequantize, the quantization conversion coefficients decoded by the entropy decoding unit 304. The inverse quantization process may include determining the degree of quantization, and similarly the degree of dequantization to be applied, using the quantization parameters calculated by the video encoder 100 for each video block in the video slice.

[0255] The inverse transformation processing unit 312 is configured to apply an inverse transformation, such as an inverse DCT, an inverse integer transformation, or a conceptually similar inverse transformation process, to the transformation coefficients in order to generate residual blocks in the pixel domain.

[0256] The reconstruction unit 314 (for example, an adder 314) is configured to obtain the reconstructed block 315 in the sample domain by adding the inverse transform block 313 (i.e., the reconstructed residual block 313) to the prediction block 365, for example by adding the sample values ​​of the reconstructed residual block 313 and the sample values ​​of the prediction block 365.

[0257] The loop filter unit 320 (either within or after the coding loop) is configured to filter the reconstructed block 315 to obtain the filtered block 321, for example, to smooth pixel transitions or otherwise improve video quality. In one example, the loop filter unit 320 may be configured to perform any combination of filtering techniques described later. The loop filter unit 320 is intended to represent one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or other filters, e.g., a bilateral filter or adaptive loop filter (ALF), or a sharpening or smoothing filter or a co-filter. Although the loop filter unit 320 is shown as an in-loop filter in Figure 3, in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.

[0258] Next, the decoded video block 321 within a given frame or image is stored in a decoded image buffer 330 that stores a reference image used for subsequent motion compensation.

[0259] The decoder 300 is configured to output the decoded image 311, for example via the output 312, for presentation or display to the user.

[0260] Other variations of the video decoder 300 can be used to decode a compressed bitstream. For example, the decoder 300 can generate an output video stream without a loop filtering unit 320. For example, a non-conversion-based decoder 300 can directly dequantize the residual signal for a particular block or frame without an inverse conversion processing unit 312. In another embodiment, the video decoder 300 can combine the inverse quantization unit 310 and the inverse conversion processing unit 312 into a single unit.

[0261] Figure 4 is a schematic diagram of a network device 400 (e.g., a coding device) according to one embodiment of the present disclosure. The network device 400 is suitable for carrying out the embodiments disclosed herein. In one embodiment, the network device 400 may be a decoder such as the video decoder 300 in Figure 1A or an encoder such as the video encoder 200 in Figure 1A. In one embodiment, the network device 400 may be one or more components of the video decoder 300 in Figure 1A or the video encoder 200 in Figure 1A, as described above.

[0262] The network device 400 includes an input port 410 and a receiver unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit (CPU) 430 for processing data; a transmitter unit (Tx) 440 and an output port 450 for transmitting data; and memory 460 for storing data. The network device 400 may also include optical-to-electrical (OE) and electrical-to-optical (EO) components coupled to the input port 410 for inputting or outputting optical or electrical signals, the receiver unit 420, the transmitter unit 440, and the output port 450.

[0263] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. The processor 430 communicates with input port 410, receiver unit 420, transmitter unit 440, output port 450, and memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the embodiments disclosed above. For example, the coding module 470 performs, processes, prepares, or provides various coding operations. Thus, including the coding module 470 provides a substantial improvement to the functionality of the network device 400 and results in the conversion of the network device 400 to different states. Alternatively, the coding module 470 may be stored in memory 460 and implemented as instructions executed by the processor 430.

[0264] Memory 460 includes one or more disks, tape drives, and solid-state drives, and is used as an overflow data storage device, which can store the program when such a program is selected for execution, and can store instructions and data to be read during the execution of the program. Memory 460 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), tri-level associative memory (TCAM), and / or static random access memory (SRAM).

[0265] Figure 5 is a simplified block diagram of a device 500 that can be used as either or both of the source device 12 and destination device 14 shown in Figure 1A, according to an exemplary embodiment. The device 500 can carry out the technology of this application. The device 500 may take the form of a computing system including multiple computing devices, or the form of a single computing device, such as a mobile phone, tablet computer, laptop computer, notebook computer, or desktop computer.

[0266] The processor 502 within the device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device, or multiple devices, capable of manipulating or processing information that currently exists or will be developed in the future. The disclosed implementations can be implemented with a single processor, e.g., processor 502, as shown, but speed and efficiency advantages can be achieved using multiple processors.

[0267] The memory 504 within the device 500 may, in an embodiment, be a read-only memory (ROM) device or a random access memory (RAM) device. Other suitable types of storage devices may be used as memory 504. Memory 504 may contain code and data 506 accessed by the processor 502 using the bus 512. Memory 504 may further include an operating system 508 and an application program 510, the application program 510 including at least one program that enables the processor 502 to perform the method described herein. For example, the application program 510 may include applications 1 to N, the applications 1 to N further including a video coding application that performs the method described herein. The device 500 may also include additional memory in the form of a secondary storage device 514, which may be, for example, a memory card used with a mobile computing device. Since video communication sessions may contain a considerable amount of information, those sessions may be stored whole or partially in the secondary storage device 514 and loaded into memory 504 as needed for processing.

[0268] The device 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines the display with a touch-sensing element capable of operating to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512. Other output devices that enable a user to program or otherwise use the device 500 may be provided in addition to, or as a replacement for, the display 518. Where the output device is or includes a display, the display can be implemented in various ways, including liquid crystal displays (LCDs), cathode ray tube (CRT) displays, plasma displays, or light-emitting diode (LED) displays such as organic LED (OLED) displays.

[0269] The device 500 may also include, or communicate with, any other image sensing device 520, which is currently existing or may be developed in the future, capable of sensing images such as a camera or an image of the user operating the device 500. The image sensing device 520 may be positioned to face the user operating the device 500. In one example, the position and optical axis of the image sensing device 520 include a field of view that is directly adjacent to the display 518 and from which the display 518 is visible.

[0270] The device 500 may also include, or communicate with, a sound sensing device 522, such as a microphone, or any other sound sensing device that is currently existing or may be developed in the future, capable of sensing sounds near the device 500. The sensing device 522 may be positioned to face a user operating the device 500 and may be configured to receive sounds made by the user while the user is operating the device 500, such as voice or other utterances.

[0271] Figure 5 shows that the processor 502 and memory 504 of device 500 are integrated into a single unit, but other configurations are possible. The operation of the processor 502 can be distributed to multiple machines (each machine having one or more processors) that can be directly connected or connected via a local area network or other network. The memory 504 can be distributed to multiple machines, such as network-based memory or memory in multiple machines running the operation of device 500. Although shown here as a single bus, the bus 512 of device 500 can consist of multiple buses. Furthermore, the secondary storage device 514 can be directly connected to other components of device 500 or can be accessed via a network, and can include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, device 500 can be implemented in a wide variety of configurations.

[0272] In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array. The nominal vertical and horizontal relative positions of the luma and chroma samples in the image are shown in Figure 6A.

[0273] Figure 8 is a conceptual diagram illustrating the derivative location of the scaling parameters used to scale the downsampled, reconstructed Lumablock. For example, Figure 8 shows an example of 4:2:0 sampling, where the scaling parameters are α and β.

[0274] Generally, when the LM prediction mode is applied, the video encoder 20 and video decoder 30 can invoke the following steps: The video encoder 20 and video decoder 30 can downsample adjacent luma samples. The video encoder 20 and video decoder 30 can derive linear parameters (i.e., α and β) (also called scaling parameters). The video encoder 20 and video decoder 30 can downsample the current luma block and derive a prediction (e.g., a prediction block) from the downsampled luma block and linear parameters. There are various methods for performing downsampling.

[0275] Figure 6B is a conceptual diagram showing examples of luma and chroma locations for downsampling luma block samples to generate predicted blocks of chroma blocks. As shown in Figure 6B, chroma samples represented by filled (i.e., solid black) triangles are predicted from two luma samples represented by two filled circles by applying a [1,1] filter. The [1,1] filter is an example of a two-tap filter.

[0276] Figure 6C is a conceptual diagram showing another example of luma and chroma positions for downsampling luma block samples to generate a predicted block. As shown in Figure 6C, the chroma samples, represented by filled (i.e., solid black) triangles, are predicted from six luma samples, represented by six filled circles, by applying a six-tap filter.

[0277] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. Where implemented in software, the functions may be stored or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include computer-readable storage media corresponding to tangible media such as data storage media, or communication media including any medium that facilitates the transfer of computer programs from one location to another, for example, according to a communication protocol. Thus, the computer-readable medium may generally correspond to (1) non-transient tangible computer-readable storage media, or (2) communication media such as signals or carrier waves. The data storage medium may be any available medium accessible by one or more computers or one or more processors for retrieving instructions, codes, and / or data structures for implementing the technologies described herein. Computer program products may include computer-readable media.

[0278] Video compression techniques such as motion compensation, intra-prediction, and loop filtering have proven effective and are therefore adopted in various video coding standards such as H.264 / AVC and H.265 / HEVC. Intra-prediction can be used when no reference image is available, or when inter-prediction coding is not used for the current block or image, for example, in an I-frame or I-slice. The reference sample for intra-prediction is usually derived from a previously coded (i.e., reconstructed) adjacent block within the same image. For example, in both H.264 / AVC and H.265 / HEVC, boundary samples from adjacent blocks are used as references for intra-prediction. There are many different intra-prediction modes to cover different texture or structural characteristics. Each mode uses a different prediction signal derivation method. For example, as shown in Figure 6D, H.265 / HEVC supports a total of 35 intra-prediction modes.

[0279] Description of the H.265 / HEVC intra-prediction algorithm.

[0280] For intra-prediction, decoded boundary samples from adjacent blocks are used as reference. The encoder selects the best intra-prediction mode for each block from 35 options (33 directional prediction modes, DC mode, and planar mode) (i.e., the mode that provides the most accurate prediction for the current block). The mapping between intra-prediction direction and intra-prediction mode number is shown in Figure 6D. Note that in modern video coding techniques, such as VVC (Variable Video Coding), more than 65 intra-prediction modes have been developed, and VVC can capture any edge direction presented in natural video. Of these prediction modes, modes with a horizontal direction (e.g., mode 10 in Figure 6D) are also called "horizontal modes," and modes with a vertical direction (e.g., mode 26 in Figure 6D) are also called "vertical modes."

[0281] Figure 7 shows the reference samples for a block. As shown in Figure 7, block "CUR" is the current block to be predicted, and the dark samples along the boundary of the current block are the reference samples used to predict the current block. These reference samples are samples in reconstructed blocks adjacent to the current block, and are also called adjacent blocks. Block "CUR" can be a luminous block or a chroma block, depending on the type of block to be predicted. The prediction signal can be derived by mapping the reference samples according to a specific method indicated by the intra-prediction mode.

[0282] Replacement of reference sample

[0283] Some or all of the reference samples may be unavailable for intra-prediction for several reasons. For example, samples outside an image, slice, or tile are considered unavailable for prediction. Furthermore, when constrained intra-prediction is enabled, reference samples belonging to the inter-predicted PU are omitted to avoid error propagation from previous images that may have been misreceived and reconstructed. When used herein, a reference sample of the current coding block is available if it is not outside the current image, slice, or title, if the reference sample has been reconstructed before the current coding block is decoded, and / or if the reference sample has not been omitted for coding decisions in the encoder. HEVC allows all prediction modes to be used after replacing unavailable reference samples. In the extreme case where no reference samples are available, all reference samples are replaced with the nominal average sample value for a given bit depth (e.g., 128 for 8-bit data). If there is at least one reference sample marked as available for intra-prediction, the unavailable reference sample is replaced with an available reference sample. Unavailable reference samples are replaced by scanning the reference samples clockwise and using the most recent available sample value for the unavailable sample. If the first sample is unavailable in the clockwise scan, that unavailable reference sample is replaced with the first available reference sample found when scanning the samples in the clockwise order. Here, "replacement" is also called padding, and the replaced sample may also be called a padded sample.

[0284] Constrained Intra Prediction

[0285] Constrained intra-prediction is a tool to avoid spatial noise propagation caused by spatial intra-prediction using encoder-decoder mismatched reference pixels. Encoder-decoder mismatched reference pixels can appear if packet loss occurs during transmission of the intercoded slice. They can also appear if lossy decoder-side memory compression is used. When constrained intra-prediction is enabled, inter-predicted samples are marked as unavailable or not available for intra-prediction, and these unavailable samples can be padded using the padding method described above to perform full intra-prediction estimation on the encoding side or intra-prediction on the decoding side.

[0286] Cross-component linear model prediction (CCLM)

[0287] Cross-component linear model prediction (CCLM), also known as cross-component intra-prediction (CCIP), is a type of intra-prediction mode used to reduce cross-component redundancy during intra-prediction mode. Figure 8 (including Figures 8A and 8B) is a schematic diagram illustrating an example of the mechanism for performing CCLM intra-prediction. Figure 8 shows an example of 4:2:0 sampling. Figure 8 shows an example of the location of the current block sample and its adjacent samples to the left and above it, included in CCLM mode. White squares are the current block samples, and shaded circles are the reconstructed samples of adjacent blocks. Figure 8A shows an example of adjacent reconstructed pixels in a chroma block. Figure 8B shows an example of adjacent reconstructed pixels in a luma block placed in the same location. When the video format is YUV4:2:0, then we have one 16x16 luma block and two 8x8 chroma blocks.

[0288] CCLM intra-prediction can be performed by the intra-estimation unit 254 of the encoder 200 and / or the intra-prediction unit 354 of the decoder 300. CCLM intra-prediction predicts chroma sample 803 within chroma block 801. Chroma sample 803 appears at integer positions indicated by squares. The prediction is partially based on adjacent reference samples indicated by black circles. Chroma sample 803 is not predicted solely based on adjacent chroma reference sample 805. Chroma sample 803 is also predicted based on luma reference sample 813 and adjacent luma reference sample 815. Specifically, the CU includes luma block 811 and two chroma blocks 801. A model is generated that correlates chroma sample 803 and luma reference sample 813 within the same CU. The linear coefficients of the model are determined by comparing adjacent luma reference sample 815 with adjacent chroma reference sample 805.

[0289] When the luma reference sample 813 is reconstructed, it is shown as the reconstructed luma sample (Rec'L). When the adjacent chromat reference sample 805 is reconstructed, it is shown as the reconstructed adjacent chromat reference sample (Rec'C).

[0290] As shown, luma block 811 contains four times the number of samples as chroma block 801. In the example shown in Figure 8, chroma block 801 contains N×N samples, while luma block 811 contains 2N×2N samples. Therefore, luma block 811 has four times the resolution of chroma block 801. For predictions that operate on luma reference sample 813 and adjacent luma reference sample 815, luma reference sample 813 and adjacent luma reference sample 815 are downsampled to provide an accurate comparison with adjacent chroma reference sample 805 and chroma sample 803. Downsampling is the process of reducing the resolution of a group of sample values. For example, when the YUV4:2:0 format is used, luma samples may be downsampled by a factor of 4 (e.g., 2 in width and 2 in height). YUV is a color coding system that uses a color space in terms of a luma component Y and two chrominance components U and V.

[0291] In CCLM prediction, the chroma sample is predicted based on the downsampled, corresponding reconstructed luma sample (the current luma block) using a linear model as follows: Nod C (i,j)=α·rec L '(i,j)+β (1) Here, pred C (i,j) represents the predicted chroma sample, and rec' L (i,j) represents the downsampled corresponding reconstructed luma sample. The parameters α and β can be derived by minimizing the regression error between reconstructed adjacent luma and chroma samples around the current luma block and the current chroma block, as follows: α=(N·Σ(L(n)·C(n))-ΣL(n)·ΣC(n)) / (N·Σ(L(n)·L(n))-ΣL(n)·ΣL(n)) (2) β = (ΣC(n) - α·ΣL(n)) / N (3) Here, L(n) represents the downsampled upper and left reconstructed adjacent luma samples, C(n) represents the upper and left reconstructed adjacent chroma samples, and the value of N is equal to the samples used for the derivation of the coefficients. In the case of a square-shaped coding block, the above two equations are directly applicable. Since this regression error minimization calculation is executed not only as an encoder search operation but also as part of the decoding process, no syntax for transmitting the α value and β value is used.

[0292] In addition to using the above method (also called the least squares (LS) method) to minimize the regression error, the linear model coefficients α and β can also be derived using the maximum and minimum luma sample values. This latter method is also called the MaxMin method. In the MaxMin method, after the downsampling of the upper and left reconstructed adjacent luma samples, a one-to-one relationship between each of these reconstructed adjacent luma samples and the upper and left reconstructed adjacent chroma samples is obtained. Thus, the linear model coefficient parameters α and β can be derived using pairs of luma and chroma samples based on the one-to-one relationship. The pairs of luma and chroma samples are obtained by identifying the minimum and maximum values of the downsampled upper and left reconstructed adjacent luma samples and then identifying the corresponding samples from the upper and left reconstructed adjacent chroma templates. The pairs of luma and chroma samples are shown as in (A, B) of FIG. 9. The linear model parameters α and β are obtained according to the following equations. α=(y B -y A ) / (x B -x A ) (4) β=y A -αx A (X Here, (x A , y A ) are the coordinates of A in FIG. 9, and (x B , y B ) are the coordinates of B.

[0293] The CCLM luma-to-chroma prediction mode is added as an additional chroma-intra prediction mode. On the encoder side, another rate-distortion (RD) cost check of the chroma component is added to select the chroma-intra prediction mode.

[0294] For brevity, in this document the term “template” is used to refer to reconstructed adjacent chroma samples and downsampled reconstructed adjacent luma samples. These reconstructed adjacent chroma samples and downsampled reconstructed adjacent luma samples are also referred to as reference samples within the template. Figure 10 is a diagram of the templates for a chroma block and its corresponding downsampled luma block. In the example shown in Figure 10, luma '1020' is a downsampled version of the current luma block and has the same spatial resolution as chroma block 1040. In other words, luma '1020' is a downsampled luma block juxtaposed with chroma block 1040. The upper template 1002 contains the upper reconstructed adjacent chroma sample above the current chroma block 1040 and the corresponding downsampled upper reconstructed adjacent luma sample for luma '1020'. The downsampled upper reconstructed adjacent luma sample for luma '1020' is obtained based on the adjacent sample above the luma block. In this specification, adjacent samples above a luma block may include either adjacent samples immediately above the luma block, adjacent samples not adjacent to the luma block, or both. Left-side template 1004 includes a reconstructed adjacent chroma sample on the left and a corresponding downsampled reconstructed adjacent luma sample on the left. An upper reconstructed adjacent chroma sample is also called an "upper chroma template," such as upper chroma template 1006. A corresponding downsampled upper reconstructed adjacent luma sample is called an "upper luma template," such as upper luma template 1008. A left-side reconstructed adjacent chroma sample is also called a "left-side chroma template," such as left-side chroma template 1010. A corresponding downsampled left reconstructed adjacent luma sample is called a "left-side luma template," such as left luma template 1012. Elements included in a template are called reference samples within that template.

[0295] In existing CCLM applications, if there is one reference sample marked as unavailable in the upper or left-hand template, the entire template is not used. Figure 11 shows an example of a template with an unavailable reference sample. In the example shown in Figure 11, for chroma block 1140, if there is an unavailable reference sample in the upper template, such as reference sample A2 1102, the upper template will not be used to derive the linear model coefficients. Similarly, if there is one unavailable reference sample in the left-hand template, such as reference sample B2 1104, the left-hand template will not be used to derive the linear model coefficients. This degrades the coding performance of intra-prediction.

[0296] Multidirectional linear model

[0297] In addition to being used to calculate linear model coefficients together, the reference samples in the upper and left templates can also be used in two other CCLM modes, namely CCLM_T and CCLM_L modes. CCLM_T and CCLM_L can collectively be called multi-directional linear models (MDLM). Figure 12 shows the reference samples used in CCLM_T mode, and Figure 13 shows the reference samples used in CCLM_L mode. As shown in Figure 12, in CCLM_T mode, only the reference samples in the upper template, such as reference samples 1202 and 1204, are used to calculate linear model coefficients. As shown in Figure 13, in CCLM_L mode, only the reference samples in the left template, such as reference samples 1212 and 1214, are used to calculate linear model coefficients. The number of reference samples used in each of these modes is W + H, where W is the width of the chroma block and H is the height of the chroma block.

[0298] CCLM and MDLM modes (i.e., CCLM_T mode and CCLM_L mode) can be used together or interchangeably. For example, the codec may use only CCLM mode, only MDLM mode, or both CCLM and MDLM modes. In the last case, where both CCLM and MDLM modes are used, three modes (i.e., CCLM, CCLM_T, CCLM_L) are added as three additional chromatic intraprediction modes. On the encoder side, three additional RD cost checks of the chroma component are added to select the chromatic intraprediction mode. In the existing MDLM method, the model parameters or model coefficients are derived using the LS method. If there are not enough available reference samples, a padding operation is used to copy the furthest pixel values ​​or fetch the sample values ​​of the available reference samples.

[0299] However, obtaining linear model coefficients for the MDLM mode using the LS method increases computational complexity. Furthermore, in existing MDLM modes, the positions of some template samples may be far from the current block, especially in the case of non-square blocks. For example, the reference sample at the right edge of the upper template and the reference sample at the bottom of the left template are far from the current block. As a result, these reference samples have a low correlation with the current block, reducing the efficiency of chroma block prediction. The technique presented herein can reduce the complexity of MDLM and increase the correlation between template samples and the current block.

[0300] For example, when determining model coefficients using the MaxMin method, in addition to using reference samples from both the upper and left templates, only a portion of the template (either the left or upper template) is used as reference samples. For instance, in the MaxMin method, only the reference samples within the upper chroma template are examined to determine the maximum and minimum chroma values. Alternatively, only the reference samples within the left chroma template are examined to determine the maximum and minimum chroma values. After the sample locations for the maximum and minimum chroma values ​​are determined, the corresponding chroma sample values ​​can be obtained based on the locations of the minimum and maximum chroma values.

[0301] Figure 14 is a schematic diagram showing an example of reference samples used to determine the maximum and minimum luma values. In the example shown in Figure 14, the number of reference samples in the upper luma template, indicated as W1, is greater than the width of the current chroma block, indicated as W. The number of reference samples in the left luma template, indicated as H1, is greater than the height of the current chroma block, indicated as H. Figure 15 is a schematic diagram showing another example of reference samples used to determine the maximum and minimum luma values. In the example shown in Figure 15, the number of upper luma reference samples is equal to the width of the current chroma block W, and the number of left luma reference samples is equal to the height of the current chroma block H.

[0302] In summary, in addition to the LS method, the MaxMin method can also be used for MDLM modes. In other words, the model coefficients for CCLM_T and CCLM_L modes can be derived using the MaxMin method. Since the MaxMin method is less computationally complex than the LS method, the proposed method improves MDLM by reducing its computational complexity. Furthermore, the existing MaxMin method uses both the upper and left-hand templates for the CCLM mode. The proposed method uses either the upper or left-hand template to derive the model coefficients, which further reduces the computational complexity of the modes.

[0303] In a further example of the technology presented herein, reference samples within a template are selected to increase the correlation between the reference sample and the current block. The upper template uses a maximum of W2 reference samples. The left template uses a maximum of H2 reference samples. In this way, reference samples further away than W2 reference samples in the upper template or H2 reference samples in the left template are not used because they have a low correlation with the current block.

[0304] Furthermore, when deriving model coefficients using the MaxMin method, only available reference samples are used, and no padding is used to replace unavailable reference samples. For example, to determine the maximum and minimum luma values ​​for the CCLM_T mode, only available samples in the upper luma template are examined. Since the upper template uses W2 reference samples, the number of available samples, indicated as W3, will be less than or equal to W2. Similarly, to determine the maximum and minimum values ​​for the CCLM_L mode, the available samples in the left luma template are examined. The number of available samples, indicated as H3, may be less than or equal to H2. The relationship between W2 and W3 and H2 and H3 is shown in Figure 16, where W3 <= W2 and H3 <= H2. In one example, W2 = 2 × W and H2 = 2 × H.

[0305] Alternatively, W2 and H2 can each take values ​​equal to W+H. In other words, the model coefficients for the CCLM_T mode can be derived using a maximum of W+H reference samples in the upper luma template, and the model coefficients for the CCLM_L mode can be derived using a maximum of W+H reference samples in the left luma template. Of these reference samples, only the luma template samples available within this range (i.e., W+H for both CCLM_T and CCLM_L) are examined to determine the maximum and minimum luma values. In this example, to determine the maximum and minimum values ​​for the CCLM_T mode, the samples available in the upper luma template (less than or equal to W+H) are examined. To determine the maximum and minimum values ​​for the CCLM_L mode, the samples available in the left luma template (less than or equal to W+H) are examined.

[0306] Compared to existing MDLM methods that use exactly W+H reference samples to derive model coefficients for CCLM_T and CCLM_L modes, the proposed method uses a maximum of 2×W or W+H reference samples to derive model coefficients for CCLM_T mode, and a maximum of 2×H or W+H reference samples to derive model coefficients for CCLM_L mode. Furthermore, to determine the maximum and minimum values, only the available luma reference samples within the examined sample range (2×W or W+H for CCLM_T, and 2×H or W+H for CCLM_L) are used. In the proposed method described above, determining the maximum and minimum luma values ​​in the luma template can be sped up by sampling reference samples with a step size greater than 1, such as 2, 4, or another value.

[0307] Downsampling method

[0308] As discussed above, the spatial resolution of the luminous component of an image is greater than that of the chroma component, so the luminous component needs to be downsampled to the resolution of the chroma portion in MDLM mode. For example, in the YUV4:2:0 format, the luminous component needs to be downsampled by 4 (by 2 in width and 2 in height) to match the resolution of the chroma component. The downsampled luminous block (which has the same spatial resolution as the chroma block) corresponding to the chroma block can be used to predict the chroma block using MDLM mode. The size of the downsampled luminous block is the same as the size of the chroma block (because the luminous block is downsampled to the size of the chroma block).

[0309] Similarly, to derive the linear model coefficients for MDLM mode, it is necessary to downsample the reference samples in the chroma template. In CCLM_T mode, the upper reconstructed neighboring chroma samples are downsampled to generate the upper template reference samples corresponding to the reference samples in the upper chroma template, i.e., the upper reconstructed neighboring chroma samples. Downsampling of the upper reconstructed neighboring chroma samples typically involves multiple rows of the upper reconstructed neighboring chroma samples. Figure 17 is a schematic diagram showing an example of downsampling a chroma sample using multiple rows or columns of neighboring chroma samples. As shown in Figure 17, in the case of a chroma block, two upper neighboring rows A1 and A2 can be used during downsampling to obtain the downsampled neighboring row A. A[i] is shown as the i-th sample of A, A1[i] as the i-th sample of A1, and A2[i] as the i-th sample of A2, and a 6-tap downsampling filter can be used as follows: A[i]=(A2[2i]*2+A2[2i-1]+A2[2i+1]+A1[2i]*2+A1[2i-1]+A1[2i+1]+4)>>3; The number of adjacent samples may also be larger than the current block size. For example, as shown in Figure 18, the number of adjacent samples at the top of the downsampled Luma block may be M, where M is greater than the width W of the downsampled block.

[0310] In existing downsampling methods like the one described above, multiple rows of the reconstructed neighboring luma samples from the top are used to generate downsampled reconstructed adjacent luma samples from the top. This results in a larger line buffer size compared to normal intra-mode prediction, thus increasing memory costs.

[0311] The technique presented herein reduces the memory usage of downsampling by using only one row of the top reconstructed adjacent luma samples for CCLM_T mode when the current block is at the top boundary of the current coding tree unit CTU (i.e., the top row of the current chroma block overlaps with the top row of the current CTU). Figure 19 is a schematic diagram showing an example of downsampling using a single row of adjacent luma samples for a luma block at the top boundary of the CTU. As shown in Figure 19, only A1 (containing one row of the reconstructed adjacent luma samples) is used to generate the top reconstructed adjacent luma samples downsampled at A.

[0312] The above explanation focuses on the reconstructed adjacent luma sample at the top of the CCLM_T mode, but a similar method can be used to reconstruct the left side of the CCLM_L mode. neighborhood It is important to understand that this can be applied to luma samples. For example, instead of using multiple columns of the reconstructed neighboring luma sample on the left for downsampling, if the current block is on the left boundary of the CTU (i.e., the left row of the current chroma block overlaps with the left row of the CTU), a single column of the reconstructed neighboring luma sample on the left is used for downsampling to generate a reference sample in the left template of the downsampled luma block.

[0313] Determining the availability of reference samples

[0314] In the example above, a luma template is used to determine the availability of reference samples for determining the maximum and minimum luma values. However, in some scenarios, available luma reference samples do not have corresponding chroma reference samples. For example, luma blocks and chroma blocks can be coded separately. Thus, if a reconstructed luma block is available, the corresponding chroma block may not yet be available. As a result, available luma reference samples do not have corresponding available chroma reference samples, which can lead to coding errors.

[0315] The technique presented herein addresses this problem by determining the availability of reference samples within a template after checking the availability of reference samples within the chroma template. In some cases, a chroma reference sample is available if it is not outside the current image, slice, or title, and the reference sample has been reconstructed. In some cases, a chroma reference sample is available if it is not outside the current image, slice, or title, the reference sample has been reconstructed, and the reference sample has not been omitted based on the encoding decision. Available reference samples in the current chroma block may be available reconstructed adjacent samples in the chroma block. Luma reference samples corresponding to available chroma reference samples are used to determine the maximum and minimum luma values. For example, if L reference samples are available in the chroma template, then the maximum and minimum luma values ​​are determined using the reference samples in the luma template corresponding to the L available chroma reference samples. The luma reference samples corresponding to chroma reference samples can be determined by identifying the chroma reference samples located within the chroma template and the luma reference samples located at the same position within the luma template (i.e., luma reference samples (position (x, y)) located in the same place, for example, the luma reference samples corresponding to a chroma reference sample may include adjacent luma reference samples (position (x-1, y)), luma reference samples (position (x, y)), and adjacent luma reference samples (position (x+1, y)). The luma reference samples with maximum and minimum luma values, and the chroma reference samples corresponding to the luma reference samples associated with the maximum and minimum luma values, are then used to determine the model coefficients as described above.

[0316] In one example, model coefficients, e.g., L=4, can be determined using L available chroma reference samples. In another example, model coefficients are determined using a portion of L available chroma reference samples. For example, a fixed number of available chroma reference samples are selected from L available chroma reference samples. Luma reference samples corresponding to the selected chroma reference samples are identified and used together with the selected chroma reference samples to determine the model coefficients. For example, if the selected chroma reference samples are 4 available chroma reference samples, then 24 reconstructed adjacent luma samples corresponding to the 4 available chroma reference samples are identified. The luma reference samples used to determine the model coefficients are obtained by downsampling the 24 reconstructed adjacent luma samples, where a 6-tap filter is used in the downsampling process.

[0317] In the example shown in Figure 16, the upper template sample range for CCLM_T mode is W2. Next, the availability of reference samples within the upper chroma template of the current chroma block is determined. If W3 chroma reference samples are available (W3 <= W2), then W3 or fewer corresponding chroma reference samples are acquired. These acquired chroma and chroma reference samples are used to derive the model coefficients for CCLM_T mode.

[0318] Similarly, in the example shown in Figure 16, the left template sample range is H2 in CCLM_L mode. The availability of reference samples in the left chroma template of the current chroma block is checked. If H3 chroma reference samples are available (H3 <= H2), then H3 or fewer corresponding chroma reference samples are obtained. These obtained chroma and chroma reference samples are used to derive the model coefficients in CCLM_L mode. The availability of reference samples in CCLM mode can be determined similarly, that is, the availability of reference samples in both the upper and left chroma templates is determined, and then the corresponding chroma reference samples can be found and the model coefficients determined as described above.

[0319] Details of the proposed method are described in Table 1 in the format of the INTRA_CCLM, INTRA_CCLM_L, or INTRA_CCLM_T intra-prediction mode specifications. Table 2 shows alternative embodiments of the method proposed herein. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 2]

[0320] MDLM mode binarization

[0321] To encode MDLM modes in a video signal bitstream, it is necessary to perform MDLM mode binarization so that the selected MDLM mode can be encoded in the bitstream, allowing the decoder to determine the mode selected for decoding. Existing binarization methods do not include the two chroma modes of MDLM, namely CCLM_L and CCLM_T. Here, we propose a novel chroma mode coding method.

[0322] Tables 3 and 4 provide details on the binarization of these two chroma modes. In Table 3, 77 represents the CCLM mode, and the intra_chroma_pred_mode index is 4. 78 represents the CCLM_L mode, and the intra_chroma_pred_mode index is 5. 79 represents the CCLM_T mode, and the intra_chroma_pred_mode index is 6. When the intra_chroma_pred_mode index is equal to 7, the selected mode is the DM mode. The remaining index values ​​0, 1, 2, and 3 represent the planar mode, vertical mode, horizontal mode, and DC mode, respectively. [Table 3] [Table 4]

[0323] Table 4 shows examples of bit strings or syntax elements used for each of the chromatic intra-prediction modes. As shown in Table 4, the syntax element for DM mode (index 7) is 0, for CCLM mode (index 4) it is 10, for CCLM_L mode (index 5) it is 1110, for CCLM_T mode (index 6) it is 1111, for planar mode (index 0) it is 11000, for vertical mode (index 1) it is 11001, for horizontal mode (index 2) it is 11010, and for DC mode (index 3) it is 11011. Table 5 shows another example of bit strings or syntax elements used for each of the chromatic intra-prediction modes. Depending on the coding mode selected by the encoder, the corresponding syntax elements are included in the bitstream of the encoded video. [Table 5]

[0324] When a video encoder, such as the video encoder 20 in Figure 1, performs intra-prediction of chroma blocks of a video signal based on an intra-chroma prediction mode, the video encoder generates a bitstream of the video signal by selecting an intra-chroma prediction mode and including a syntax element in the bitstream that indicates the selected intra-chroma prediction mode. The video encoder can select an intra-chroma prediction mode from a set of multiple modes. For example, the modes may include a first set of modes that includes derived mode (DM) or cross-component linear model (CCLM) prediction modes, or both. The modes may also include a second set of modes that includes at least one of the CCLM_L mode or CCLM_T mode. The modes may further include a third set of modes that may include at least one of the vertical mode, horizontal mode, DC mode, or planar mode.

[0325] In some examples, the number of bits in the syntax element of the intrachroma prediction mode when the intrachroma prediction mode is selected from the first mode set is smaller than the number of bits in the syntax element when the intrachroma prediction mode is selected from the second mode set. Furthermore, the number of bits in the syntax element of the intrachroma prediction mode when the intrachroma prediction mode is selected from the second mode set is smaller than the number of bits in the syntax element when the intrachroma prediction mode is selected from the third mode set. In some examples, the syntax elements of various intrachroma prediction modes are selected according to the examples shown in Table 4 or Table 5.

[0326] When a decoder receives an encoded bitstream of a video signal and decodes the video, the decoder, such as the video decoder 30 in Figure 1, analyzes syntax elements from the video signal bitstream and determines an intrachroma prediction mode to be used for the chroma block based on the syntax elements selected from the analyzed elements. Based on the determined intrachroma prediction mode, the decoder performs an intra-prediction of the current chroma block of the video signal.

[0327] Figure 20 is a flowchart of a method for performing intraprediction using a linear model according to several aspects of this disclosure. In block 2002, the luma block (e.g., luma block 811) corresponding to the current chroma block (e.g., chroma block 801) is determined.

[0328] In block 2004, the luma reference sample of the luma block is obtained based on determining the L available chroma reference samples of the current chroma block. The luma reference sample obtained in the luma block is a downsampled luma reference sample. In some examples, the luma reference sample obtained in the luma block is a downsampled luma reference sample obtained by downsampling adjacent luma samples selected based on the L available chroma reference samples (or based on some or all of the L available chroma reference samples). In other words, the luma reference sample obtained in the luma block is a downsampled luma reference sample obtained by downsampling adjacent luma samples corresponding to the available chroma reference samples. In some examples, the obtained luma reference sample corresponds to the L available chroma reference samples. In additional examples, the obtained luma reference sample corresponds to some of the L available chroma reference samples. It can be understood that the correspondence between the acquired luma reference sample (i.e., the downsampled luma reference sample) and L available chroma reference samples is not limited to a "one-to-one correspondence," and that the correspondence between the acquired luma reference sample (i.e., the downsampled luma reference sample) and L available chroma reference samples can be an "M to N correspondence." For example, M=4, N=4, or M=4, N>4.

[0329] In some examples, the chroma reference sample of the current chroma block includes the reconstructed neighboring sample of the current chroma block. L available chroma reference samples are determined from the reconstructed neighboring sample. Similarly, the neighboring sample of the luma block is also the reconstructed neighboring sample of the luma block. The acquired luma reference sample of the luma block is obtained by downsampling the reconstructed neighboring sample selected based on the L available chroma reference samples. Such an example is L=4.

[0330] In some cases, chroma reference samples are available when the chroma reference sample is not outside the current image, slice, or title, and the reference sample has been reconstructed. In some cases, chroma reference samples are available when the chroma reference sample is not outside the current image, slice, or title, the reference sample has been reconstructed, and the reference sample has not been omitted based on the encoding decision. Available reference samples for the current chroma block may be available reconstructed adjacent samples for the chroma block. A chroma reference sample corresponding to an available chroma reference sample is obtained.

[0331] In some examples, the L available chroma reference samples are determined by determining that L of the top adjacent chroma samples of the current chroma block are available, where 1 <= L <= W2, and L and W2 are positive integers. W2 represents the top reference sample range, and L of the top adjacent chroma samples are used as available chroma reference samples. In some examples, W2 is equal to either 2*W or W+H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0332] In other examples, the L available chroma reference samples are determined by determining the L available left-side neighboring chroma samples of the current chroma block, where 1 <= L <= H2, and L and H2 are positive integers. H2 represents the left-side reference sample range. The L left-side neighboring chroma samples are used as the available chroma reference samples. In some examples, H2 is equal to either 2*H or W+H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0333] In further examples, the L available chroma reference samples are determined by determining the L1 upper adjacent chroma samples and L2 left adjacent chroma samples available for the current chroma block, where 1 <= L1 <= W2 and 1 <= L2 <= H2. W2 represents the upper reference sample range, and H2 represents the left reference sample range. L1, L2, W2, and H2 are positive integers, and L1 + L2 = L. In these examples, L1 upper adjacent chroma samples and L2 left adjacent chroma samples are used as the available chroma reference samples.

[0334] In one example, the Luma reference sample is obtained by downsampling only neighboring samples that are located above the Luma block and selected based on L available chroma reference samples. In another example, the Luma reference sample is obtained by downsampling only neighboring samples that are located to the left of the Luma block and selected based on L available chroma reference samples.

[0335] In the example above, the downsampled luma block is obtained by downsampling the reconstructed luma block corresponding to the current chroma block. In some cases, such as when the luma reference sample is obtained based only on adjacent samples above the luma block, and when the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), the luma reference sample is obtained using only one row of the reconstructed adjacent luma sample from the reconstructed version of the luma block.

[0336] In block 2006, the linear model coefficients used for cross-component prediction are calculated based on the luma reference sample obtained in step 2004 and the chroma reference sample corresponding to the luma reference sample. In some examples, the chroma reference sample corresponding to the luma reference sample is a chroma reference sample located in the same location as the luma reference sample.

[0337] In Block 2008, the current chroma block prediction is generated based on the calculated linear model coefficients and the downsampled luma block values ​​obtained by downsampling luma blocks (such as luma block 811).

[0338] Figure 21 is a flowchart of a cross-component linear model (CCLM) prediction method according to another aspect of the present disclosure. In block 2102, the luma block (e.g., luma block 811) corresponding to the current chroma block (e.g., chroma block 801) is determined.

[0339] In block 2104, the luma reference sample of the luma block is obtained by downsampling the adjacent samples of the luma block. In some examples, the luma reference sample includes only the luma reference sample obtained based on the adjacent sample above the luma block. In other examples, the luma reference sample includes only the luma reference sample obtained based on the adjacent sample to the left of the luma block.

[0340] In block 2106, the maximum and minimum luma values ​​are determined based on the luma reference sample.

[0341] In block 2108, the first chroma value is obtained at least partially based on one or more locations of one or more luma reference samples associated with the maximum luma value. The second chroma value is also obtained at least partially based on one or more locations of one or more luma reference samples associated with the minimum luma value.

[0342] In block 2110, the linear model coefficients are calculated based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value.

[0343] In block 2112, the current chroma block prediction is generated based on the linear model coefficients and the downsampled luma block values.

[0344] Figure 22 is a block diagram showing an exemplary structure of equipment 2200 for performing intra-prediction using a linear model. Equipment 2200 may include a decision unit 2202 and an intra-prediction processing unit 2204. In one example, equipment 2200 may correspond to the intra-prediction unit 254 in Figure 2. In another example, equipment 2200 may correspond to the intra-prediction unit 354 in Figure 3.

[0345] The decision unit 2202 is configured to determine the luma block (block 811, etc.) corresponding to the current chroma block (chroma block 801, etc.). Based on determining L available chroma reference samples for the current chroma block, the decision unit 2202 is further configured to obtain luma reference samples for the luma block. The obtained luma reference samples for the luma block are downsampled luma reference samples.

[0346] In some examples, the chroma reference sample of the current chroma block includes the reconstructed neighboring sample of the current chroma block. L available chroma reference samples are determined from the reconstructed neighboring sample. Similarly, the neighboring sample of the luma block is also the reconstructed neighboring sample of the luma block. The acquired luma reference sample of the luma block is obtained by downsampling the reconstructed neighboring sample of the luma block. In some examples, the acquired luma reference sample of the luma block is a downsampled luma reference sample obtained by downsampling the reconstructed neighboring sample of the luma block selected based on L available chroma reference samples. In some examples, the acquired luma reference sample of the luma block is a downsampled luma reference sample obtained by downsampling the reconstructed neighboring sample corresponding to L available chroma reference samples.

[0347] In some cases, chroma reference samples are available when the chroma reference sample is not outside the current image, slice, or title, the reference sample has been reconstructed, and the reference sample is not omitted based on the encoding decision. Available reference samples for the current chroma block may be available reconstructed adjacent samples for the chroma block. A chroma reference sample corresponding to the available chroma reference sample is obtained.

[0348] In some examples, the L available chroma reference samples are determined by determining that L of the top adjacent chroma samples of the current chroma block are available, where 1 <= L <= W2, and L and W2 are positive integers. W2 represents the top reference sample range, and the L top adjacent chroma samples are used as available chroma reference samples. In some examples, W2 is equal to either 2*W or W+H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0349] In other examples, the L available chroma reference samples are determined by determining the L available left-side neighboring chroma samples of the current chroma block, where 1 <= L <= H2, and L and H2 are positive integers. H2 represents the left-side reference sample range. The L left-side neighboring chroma samples are used as the available chroma reference samples. In some examples, H2 is equal to either 2*H or W+H, where W represents the width of the current chroma block and H represents the height of the current chroma block.

[0350] In further examples, the L available chroma reference samples are determined by determining the L1 upper adjacent chroma samples and L2 left adjacent chroma samples available for the current chroma block, where 1 <= L1 <= W2 and 1 <= L2 <= H2. W2 represents the upper reference sample range, and H2 represents the left reference sample range. L1, L2, W2, and H2 are positive integers, and L1 + L2 = L. In these examples, L1 upper adjacent chroma samples and L2 left adjacent chroma samples are used as the available chroma reference samples.

[0351] In one example, the luma reference sample is obtained by downsampling only adjacent samples that are located above the luma block and selected based on L available chroma reference samples. In another example, the luma reference sample is obtained by downsampling only adjacent samples that are located to the left of the luma block and selected based on L available chroma reference samples.

[0352] In the example above, the downsampled luma block is obtained by downsampling the reconstructed luma block corresponding to the current chroma block. In some cases, such as when the luma reference sample is obtained based only on adjacent samples above the luma block, and when the top row of the current chroma block overlaps with the top row of the current coding tree unit (CTU), the luma reference sample is obtained using only one row of the reconstructed adjacent luma sample from the reconstructed version of the luma block.

[0353] The intra-prediction processing unit 2204 is configured to calculate linear model coefficients (α and β, etc.) based on the luma reference sample and the chroma reference sample corresponding to the luma reference sample. The intra-prediction processing unit 2204 is further configured to obtain a prediction of the current chroma block based on the linear model coefficients and the downsampled luma block values ​​of the luma block.

[0354] Figure 23 is a flowchart illustrating a method for coding a chromatic intracoding mode in a bitstream of a video signal according to several aspects of the present disclosure.

[0355] In block 2302, intra-prediction of the chroma block of the video signal is performed based on the intra-chroma prediction mode. The intra-chroma prediction mode can be selected from multiple modes. In some examples, the multiple modes include three sets: a first mode set including at least one of the derivation mode (DM) or cross-component linear model (CCLM) prediction modes; a second mode set including at least one of the CCLM_L mode or CCLM_T mode; or a third mode set including at least one of the vertical mode, horizontal mode, DC mode, or planar mode.

[0356] In block 2304, the video signal bitstream is generated by including a syntax element in the bitstream that indicates the intrachroma prediction mode. In some examples, the number of bits in the syntax element when the intrachroma prediction mode is selected from a first mode set is less than the number of bits in the syntax element when the intrachroma prediction mode is selected from a second mode set. The number of bits in the syntax element when the intrachroma prediction mode is selected from a second mode set is less than the number of bits in the syntax element when the intrachroma prediction mode is selected from a third mode set.

[0357] For example, the syntax element for DM mode is 0. The syntax element for CCLM mode is 10. The syntax element for CCLM_L mode is 1110. The syntax element for CCLM_T mode is 1111. The syntax element for planar mode is 11000. The syntax element for vertical mode is 11001. The syntax element for horizontal mode is 11010. The syntax element for DC mode is 11011.

[0358] In another example, the syntax element for DM mode is 00. The syntax element for CCLM mode is 10. The syntax element for CCLM_L mode is 110. The syntax element for CCLM_T mode is 111. The syntax element for planar mode is 0100. The syntax element for vertical mode is 0101. The syntax element for horizontal mode is 0110. The syntax element for DC mode is 0111.

[0359] Figure 24 is a flowchart of a method for decoding the chromatic intracoding mode in a bitstream of a video signal according to several aspects of the present disclosure.

[0360] In block 2402, multiple syntax elements are analyzed from the video signal bitstream. In block 2404, the intrachroma prediction mode is determined based on the syntax element indicating the intrachroma prediction mode from the multiple syntax elements. In some examples, the intrachroma prediction mode is determined from multiple modes. For example, the multiple modes include three sets: a first mode set containing at least one of the derivation mode (DM) or cross-component linear model (CCLM) prediction modes; a second mode set containing at least one of the CCLM_L mode or CCLM_T mode; or a third mode set containing at least one of the vertical mode, horizontal mode, DC mode, or planar mode. Of these intrachroma prediction mode sets, the number of bits in the syntax element when the intrachroma prediction mode is selected from the first mode set is less than the number of bits in the syntax element when the intrachroma prediction mode is selected from the second mode set. The number of bits in the syntax element when the intrachroma prediction mode is selected from the second mode set is smaller than the number of bits in the syntax element when the intrachroma prediction mode is selected from the third mode set.

[0361] In block 2406, intra-prediction of the current chroma block of the video signal is performed based on the intra-chroma prediction mode.

[0362] Figure 25 is a block diagram showing an exemplary structure of a device 2500 for generating a video bitstream. The device 2500 may include an intra-prediction processing unit 2502 and a binarization unit 2504. In one example, the intra-prediction processing unit 2502 may correspond to the intra-prediction unit 254 in Figure 2. In one example, the binarization unit 2504 may correspond to the entropy coding unit 270 in Figure 2.

[0363] The intra-prediction processing unit 2502 is configured to perform intra-prediction of chroma blocks of a video signal based on an intra-chroma prediction mode. The intra-chroma prediction mode is selected from a first mode set, a second mode set, or a third mode set. The first mode set includes at least one of the derivation mode (DM) or cross-component linear model (CCLM) prediction modes. The second mode set includes at least one of the CCLM_L mode or CCLM_T mode. The third mode set includes at least one of the vertical mode, horizontal mode, DC mode, or planar mode.

[0364] The binarization unit 2504 is configured to generate a bitstream of the video signal by including a syntax element indicating the intrachroma prediction mode. The number of bits in the syntax element when the intrachroma prediction mode is selected from a first mode set is less than the number of bits in the syntax element when the intrachroma prediction mode is selected from a second mode set, and the number of bits in the syntax element when the intrachroma prediction mode is selected from a second mode set is less than the number of bits in the syntax element when the intrachroma prediction mode is selected from a third mode set.

[0365] For example, the syntax element for DM mode is 0. The syntax element for CCLM mode is 10. The syntax element for CCLM_L mode is 1110. The syntax element for CCLM_T mode is 1111. The syntax element for planar mode is 11000. The syntax element for vertical mode is 11001. The syntax element for horizontal mode is 11010. The syntax element for DC mode is 11011.

[0366] In another example, the syntax element for DM mode is 00. The syntax element for CCLM mode is 10. The syntax element for CCLM_L mode is 110. The syntax element for CCLM_T mode is 111. The syntax element for planar mode is 0100. The syntax element for vertical mode is 0101. The syntax element for horizontal mode is 0110. The syntax element for DC mode is 0111.

[0367] Figure 26 is a block diagram showing an exemplary structure of a device 2600 for decoding a video bitstream. The device may include an analysis unit 2602, a decision unit 2604, and an intra-prediction processing unit 2606. In one example, the analysis unit 2602 may correspond to the entropy coding unit 304 in Figure 3. In one example, the decision unit 2604 and the intra-prediction processing unit 2606 may correspond to the intra-prediction unit 354 in Figure 3.

[0368] The analysis unit 2602 is configured to analyze syntax elements from the bitstream of the video signal. The decision unit 2604 is configured to determine an intrachroma prediction mode based on syntax elements from a plurality of syntax elements. The intrachroma prediction mode is determined from one of a first mode set, a second mode set, or a third mode set. The first mode set includes at least one of the derivation mode (DM) or cross-component linear model (CCLM) prediction modes. The second mode set includes at least one of the CCLM_L mode or CCLM_T mode. The third mode set includes at least one of the vertical mode, horizontal mode, DC mode, or planar mode.

[0369] The number of bits in the syntax element when the intrachroma prediction mode is selected from the first mode set is less than the number of bits in the syntax element when the intrachroma prediction mode is selected from the second mode set, and the number of bits in the syntax element when the intrachroma prediction mode is selected from the second mode set is less than the number of bits in the syntax element when the intrachroma prediction mode is selected from the third mode set.

[0370] For example, the syntax element for DM mode is 0. The syntax element for CCLM mode is 10. The syntax element for CCLM_L mode is 1110. The syntax element for CCLM_T mode is 1111. The syntax element for planar mode is 11000. The syntax element for vertical mode is 11001. The syntax element for horizontal mode is 11010. The syntax element for DC mode is 11011.

[0371] In another example, the syntax element for DM mode is 00. The syntax element for CCLM mode is 10. The syntax element for CCLM_L mode is 110. The syntax element for CCLM_T mode is 111. The syntax element for planar mode is 0100. The syntax element for vertical mode is 0101. The syntax element for horizontal mode is 0110. The syntax element for DC mode is 0111.

[0372] The intra-prediction processing unit 2606 is configured to perform intra-prediction of the current chroma block of the video signal based on the intra-chroma prediction mode.

[0373] The following references are incorporated by reference to better understand the current disclosure: JCTVC-H0544, Description of MDLM, JVET-G1001, Description of CCLM or LM, Section 2.2.4, and JVET-K0204, Description for Deriving Model Coefficients Using Maximum and Minimum Values.

[0374] The following describes applications of the encoding and decoding methods shown in the embodiments described above, as well as systems that use these applications.

[0375] Figure 27 is a block diagram showing a content supply system 3100 for realizing a content distribution service. This content supply system 3100 includes an input device 3102 and a terminal device 3106, and optionally includes a display 3126. The input device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, Wi-Fi, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination thereof.

[0376] The acquisition device 3102 can generate data and encode the data using an encoding method as shown in the above embodiment. Alternatively, the acquisition device 3102 can distribute the data to a streaming server (not shown in the figure), which encodes the data and transmits the encoded data to the terminal device 3106. The acquisition device 3102 includes, but is not limited to, a camera, smartphone or tablet, computer or laptop, video conferencing system, PDA, in-vehicle device, or any combination thereof. For example, the acquisition device 3102 may include a source device 12 as described above. If the data includes video, the video encoder 20 included in the acquisition device 3102 can actually perform video encoding. If the data includes sound (i.e., audio), the audio encoder included in the acquisition device 3102 can actually perform audio encoding. In some practical scenarios, the acquisition device 3102 distributes the encoded video and audio data by multiplexing them together (the video encoded and audio encoded data). In other practical scenarios, such as a video conferencing system, the encoded audio data and encoded video data are not multiplexed. The input device 3102 distributes the encoded audio data and encoded video data separately to the terminal device 3106.

[0377] In the content supply system 3100, the terminal device 310 receives and plays back the encoded data. The terminal device 3106 may be a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a television 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof, which can decode the encoded data described above. For example, the terminal device 3106 may include the destination device 14 as described above. If the encoded data includes video, the video decoder 30 included in the terminal device will prioritize video decoding. If the encoded data includes audio, the audio decoder included in the terminal device will prioritize audio decoding.

[0378] In the case of terminal devices including a display, such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a television 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can supply the decoded data to its display. In the case of terminal devices without a display, such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 contacts them to receive and display the decoded data.

[0379] When each device in this system performs encoding or decoding, an image encoding device or an image decoding device can be used, as shown in the embodiments described above.

[0380] Figure 28 shows the structure of an example terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol processing unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, Real-Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-Time Transport Protocol (RTP), Real-Time Messaging Protocol (RTMP), or any combination thereof.

[0381] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some practical scenarios, for example in a video conferencing system, the encoded audio data and encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and audio decoder 3208 without going through the demultiplexing unit 3204.

[0382] Through demultiplexing, a video elementary stream (ES), an audio ES, and optional subtitles are generated. A video decoder 3206, including a video decoder 30 as described in the above-described embodiment, decodes the video ES to generate video frames using the decoding method shown in the above-described embodiment and supplies this data to the synchronization unit 3212. An audio decoder 3208 decodes the audio ES to generate audio frames and supplies this data to the synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in Figure 28) before supplying their data to the synchronization unit 3212. Similarly, the audio frames may be stored in a buffer (not shown in Figure 28) before supplying their data to the synchronization unit 3212.

[0383] The synchronization unit 3212 synchronizes video frames and audio frames and supplies video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. The information can be coded using syntax with timestamps related to the presentation of coded audio and visual data and timestamps related to the delivery of the data stream itself.

[0384] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes the decoded data with the video and audio frames, and supplies the video / audio / subtitles to the video / audio / subtitle display 3216.

[0385] The present invention is not limited to the system described above, and either the image encoding device or the image decoding device in the above-described embodiment can be incorporated into other systems, such as an automotive system.

[0386] For example, but not limited to, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage devices, magnetic disk storage devices, or other magnetic storage devices, flash memory, or any other media that can be used to store desired program code in the form of instructions or data structures and are accessible by a computer. Connections are also appropriately referred to as computer-readable media. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other temporary media, but instead refer to non-temporary, tangible storage media. As used herein, the terms "disk" and "disc" include compact discs (CDs), laser discs, optical discs, digital-purpose discs (DVDs), floppy disks, and Blu-ray discs, where a "disk" typically reproduces data magnetically, while a "disc" reproduces data optically using a laser. Any combination of the above should also be included within the scope of computer-readable media.

[0387] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, as used herein, the term “processor” may refer to any of the aforementioned structures or any other structure suitable for embodiments of the technology described herein. Furthermore, in some embodiments, the functions described herein may be provided within dedicated hardware and / or software modules incorporated into codecs configured or combined for encoding and decoding. These technologies can also be fully implemented in one or more circuits or logic elements.

[0388] The technology of this disclosure can be implemented in a wide variety of devices or equipment, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). While various components, modules, or units are described in order to highlight the functional aspects of devices configured to perform the disclosed technology, these various components, modules, or units do not necessarily require implementation by different hardware units. Rather, as described above, these various units may be provided by a set of inter-operative hardware units, including one or more of the processors, combined with a codec hardware unit or appropriate software and / or firmware.

[0389] While several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of this disclosure. These embodiments should be considered illustrative and non-limiting, and the intent should not be limited to the details given herein. For example, different elements or components can be combined or integrated into other systems, and certain functions can be omitted or not implemented.

[0390] Furthermore, the technologies, systems, subsystems, and methods described and illustrated in various embodiments, either discretely or separately, can be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as being coupled or directly coupled or communicating with one another can be indirectly coupled or communicated through any interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of modifications, substitutions, and alterations are readily apparent to those skilled in the art and can be made without departing from the spirit and scope disclosed herein.

Claims

1. A method for encoding video data implemented by an encoding device, wherein the method is A step of performing an intra-prediction of the current chroma block of the video data based on a chroma intra-prediction mode, wherein the chroma intra-prediction mode is a selected CCLM_L mode, and the intra-prediction of the current chroma block is To determine the luma block corresponding to the current chroma block, The process involves determining L available chroma reference samples for the current chroma block by checking the availability of neighboring chroma samples to the left of the current chroma block, where L is a positive integer. A step of obtaining a luma reference sample of the luma block based on the L available chroma reference samples determined for the current chroma block, wherein the obtained luma reference sample of the luma block is a downsampled luma reference sample obtained by downsampling a neighboring luma sample relating to the luma block that corresponds to the available chroma reference sample. Calculating linear model coefficients based on the aforementioned Luma reference sample and the corresponding Chroma reference sample, and The steps include obtaining a prediction of the current chroma block based on the linear model coefficients and the sampled values ​​of the downsampled chroma block, The steps include generating a bitstream of the video data by including syntactic elements indicating the chromatic intra prediction mode, method.

2. Determining the L available chromatic reference samples is: The method according to claim 1, comprising determining that the L left neighbor chromatic samples of the current chromatic block are available by checking the availability of the left neighbor chromatic samples within the left reference sample range, wherein 1 < L < H2, where H2 represents the left reference sample range, L and H2 are positive integers, and the L left neighbor chromatic samples are used as the L available chromatic reference samples.

3. The method according to claim 2, wherein H2 is equal to either 2 * H or W + H, W represents the width of the current chroma block, and H represents the height of the current chroma block.

4. The method according to claim 1, wherein the luma reference sample is located to the left of the luma block and is obtained by downsampling only the neighboring samples selected based on the L available chroma reference samples.

5. The method according to claim 1, wherein the downsampled lumablock of the lumablock is obtained by downsampling the reconstructed lumablock of the lumablock corresponding to the current chromablock.

6. Calculating the linear model coefficients based on the Luma reference sample and the corresponding Chroma reference sample is: Based on the aforementioned luma reference sample, determine the maximum luma value and the minimum luma value. Obtaining a first chroma value based at least partially on the position of the chroma reference sample associated with the maximum chroma value, Obtaining a second chroma value based at least partially on the position of the chroma reference sample associated with the minimum chroma value, and The method according to claim 1, comprising calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value.

7. Obtaining the first chroma value based at least partially on the position of the chroma reference sample associated with the maximum chroma value is: This includes obtaining the first chroma value based at least partially on one or more locations of one or more chroma reference samples associated with the maximum chroma value, Obtaining the second chroma value based at least partially on the position of the chroma reference sample associated with the minimum chroma value is: The method according to claim 6, comprising obtaining the second chroma value based at least partially on one or more locations of one or more chroma reference samples associated with the minimum chroma value.

8. A method for decoding video data implemented by a decoding device, wherein the method is A step of parsing syntactic elements from a bitstream, wherein the syntactic elements represent chromatintra prediction modes, A step of determining the chromatintra prediction mode based on the syntactic elements, wherein the chromatintra prediction mode is CCLM_L mode, A step of performing an intra-prediction of the current chroma block of the video data based on the chroma intra-prediction mode, wherein the intra-prediction of the current chroma block is: To determine the luma block corresponding to the current chroma block, The process involves determining L available chroma reference samples for the current chroma block by checking the availability of neighboring chroma samples to the left of the current chroma block, where L is a positive integer. The process involves obtaining a luma reference sample of the luma block based on the L available chroma reference samples determined for the current chroma block, wherein the obtained luma reference sample of the luma block is a downsampled luma reference sample obtained by downsampling a neighboring luma sample relating to the luma block that corresponds to the available chroma reference sample, and The steps include: calculating linear model coefficients based on the aforementioned luma reference sample and the corresponding chroma reference sample; The step of obtaining a prediction of the current chromablock based on the linear model coefficients and the sampled values ​​of the downsampled chromablock. method.

9. Determining the L available chromatic reference samples is: The method of claim 8, comprising determining that the L left neighbor chromatic samples of the current chromatic block are available by checking the availability of the left neighbor chromatic samples within the left reference sample range, wherein 1 < L < H2, where H2 represents the left reference sample range, L and H2 are positive integers, and the L left neighbor chromatic samples are used as the available chromatic reference samples.

10. The method according to claim 9, wherein H2 is equal to either 2 * W or W + H, H represents the width of the current chroma block, and H represents the height of the current chroma block.

11. The method according to claim 8, wherein the luma reference sample is located to the left of the luma block and is obtained by downsampling only the neighboring samples selected based on the determined L available chroma reference samples.

12. The method according to claim 8, wherein the downsampled lumablock of the lumablock is obtained by downsampling the reconstructed lumablock of the lumablock corresponding to the current chromablock.

13. Calculating the linear model coefficients based on the Luma reference sample and the corresponding Chroma reference sample is: Based on the aforementioned luma reference sample, determine the maximum luma value and the minimum luma value. Obtaining a first chroma value based at least partially on the position of the chroma reference sample associated with the maximum chroma value, Obtaining a second chroma value based at least partially on the position of the chroma reference sample associated with the minimum chroma value, and The method according to claim 8, comprising calculating linear model coefficients based on the first chroma value, the second chroma value, the maximum luma value, and the minimum luma value.

14. Obtaining the first chroma value based at least partially on the position of the chroma reference sample associated with the maximum chroma value is: This includes obtaining a first chroma value based at least partially on one or more locations of one or more chroma reference samples associated with the maximum chroma value, Obtaining the second chroma value based at least partially on the position of the chroma reference sample associated with the minimum chroma value is: The method according to claim 13, comprising obtaining a second chroma value based at least partially on one or more locations of one or more chroma reference samples associated with the minimum chroma value.

15. An encoding device, said encoding device, At least one processor, The at least one processor is coupled to one or more memories for storing programming instructions, When the programming instruction is executed by the at least one processor, the encoding device will: A step of performing an intra-prediction of the current chroma block of video data based on a chroma intra-prediction mode, wherein the chroma intra-prediction mode is a selected CCLM_L mode, and the intra-prediction of the current chroma block is To determine the luma block corresponding to the current chroma block, The process involves determining L available chroma reference samples for the current chroma block by checking the availability of neighboring chroma samples to the left of the current chroma block, where L is a positive integer. The process involves obtaining a luma reference sample of the luma block based on the L available chromatic reference samples determined for the current chromatic block, wherein the obtained luma reference sample of the luma block is a downsampled luma reference sample obtained by downsampling a neighboring luma sample relating to the luma block that corresponds to the available chromatic reference sample. Calculating linear model coefficients based on the aforementioned Luma reference sample and the corresponding Chroma reference sample, and The steps include obtaining a prediction of the current chroma block based on the linear model coefficients and the sampled values ​​of the downsampled chroma block, The steps include: generating a bitstream of the video data by including a syntactic element indicating the chromatic intra prediction mode; Encoding device.

16. The encoding apparatus according to claim 15, wherein the luma reference sample is located to the left of the luma block and is obtained by downsampling only neighboring samples selected based on the L available chroma reference samples.

17. A decoding device, said decoding device is At least one processor, The at least one processor is coupled to one or more memories for storing programming instructions, When the programming instruction is executed by the at least one processor, the decoding device will A step of parsing syntactic elements from a bitstream, wherein the syntactic elements represent chromatintra prediction modes, A step of determining the chromatintra prediction mode based on the syntactic elements, wherein the chromatintra prediction mode is CCLM_L mode, A step of performing an intra-prediction of the current chroma block of video data based on the chroma intra-prediction mode, wherein the intra-prediction of the current chroma block is: To determine the luma block corresponding to the current chroma block, The process involves determining L available chroma reference samples for the current chroma block by checking the availability of neighboring chroma samples to the left of the current chroma block, where L is a positive integer. The process involves obtaining a luma reference sample of the luma block based on the L available chroma reference samples determined for the current chroma block, wherein the obtained luma reference sample of the luma block is a downsampled luma reference sample obtained by downsampling a neighboring luma sample relating to the luma block that corresponds to the available chroma reference sample, and The steps include: calculating linear model coefficients based on the aforementioned luma reference sample and the corresponding chroma reference sample; The process involves performing the steps of obtaining a prediction of the current chroma block based on the linear model coefficients and the sampled values ​​of the downsampled chroma block. Decryption device.

18. The decoding apparatus according to claim 17, wherein the luma reference sample is located to the left of the luma block and is obtained by downsampling only neighboring samples selected based on the L available chroma reference samples.

19. A non-temporary computer-readable medium that holds program instructions, wherein when the program instructions are executed by a computer device or one or more processors, the computer device or one or more processors, A step of parsing syntactic elements from a bitstream, wherein the syntactic elements represent chromatintra prediction modes, A step of determining the chromatintra prediction mode based on the syntactic elements, wherein the chromatintra prediction mode is CCLM_L mode, A step of performing an intra-prediction of the current chroma block of video data based on the chroma intra-prediction mode, wherein the intra-prediction of the current chroma block is: To determine the luma block corresponding to the current chroma block, The process involves determining L available chroma reference samples for the current chroma block by checking the availability of neighboring chroma samples to the left of the current chroma block, where L is a positive integer. The process involves obtaining a luma reference sample of the luma block based on the L available chromatic reference samples determined for the current chromatic block, wherein the obtained luma reference sample of the luma block is a downsampled luma reference sample obtained by downsampling a neighboring luma sample relating to the luma block that corresponds to the available chromatic reference sample. The steps include: calculating linear model coefficients based on the aforementioned luma reference sample and the corresponding chroma reference sample; The procedure involves performing the steps of obtaining a prediction of the current chroma block based on the linear model coefficients and the sample values ​​of the downsampled chroma block. A non-temporary computer-readable medium.

20. A non-temporary computer-readable medium that holds program instructions, wherein when the program instructions are executed by a computer device or processor, the computer device or processor... A step of performing an intra-prediction of the current chroma block of video data based on a chroma intra-prediction mode, wherein the chroma intra-prediction mode is a selected CCLM_L mode, and the intra-prediction of the current chroma block is To determine the luma block corresponding to the current chroma block, The process involves determining L available chroma reference samples for the current chroma block by checking the availability of neighboring chroma samples to the left of the current chroma block, where L is a positive integer. The process involves obtaining a luma reference sample of the luma block based on the L available chromatic reference samples determined for the current chromatic block, wherein the obtained luma reference sample of the luma block is a downsampled luma reference sample obtained by downsampling a neighboring luma sample relating to the luma block that corresponds to the available chromatic reference sample. Calculating linear model coefficients based on the aforementioned Luma reference sample and the corresponding Chroma reference sample, and The steps include obtaining a prediction of the current chroma block based on the linear model coefficients and the sampled values ​​of the downsampled chroma block, The steps include generating a bitstream of the video data by including syntactic elements that indicate the chromatic intra prediction mode, and causing the following to be performed. A non-temporary computer-readable medium.