Video encoder, video decoder, and corresponding encoding and decoding method

Intra prediction using cross-component linear models addresses the challenge of balancing video quality and bandwidth by predicting chroma samples from luma samples, enhancing encoding and decoding efficiency in high-resolution video transmission.

JP2025160398APending Publication Date: 2025-10-22HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025127737
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-07-17
Filing Date
2025-07-30
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in balancing video quality and bandwidth requirements, particularly in high-resolution video transmission, where compression often compromises quality.

Method used

Intra prediction using cross-component linear models (CCLMs) is employed to predict chroma samples from luma samples, allowing for flexible video coding by determining linear model parameters based on downsampled reconstructed neighboring luma samples and chroma values to enhance prediction accuracy.

Benefits of technology

This approach improves video quality while reducing bandwidth requirements by leveraging cross-component linear models for more efficient video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025160398000001_ABST
    Figure 2025160398000001_ABST
Patent Text Reader

Abstract

To provide techniques for linear model prediction modes.SOLUTION: Two pairs of luma values and chroma values are determined according to N reconstructed neighboring luma samples, N reconstructed neighboring chroma samples corresponding to the N reconstructed neighboring luma samples, M reconstructed neighboring luma samples, and M reconstructed neighboring chroma samples corresponding to the M reconstructed neighboring luma samples. The minimum value of the N reconstructed neighboring luma samples is not less than the luma value of the remaining reconstructed neighboring luma samples of a set of the reconstructed neighboring luma samples. The maximum value of the M reconstructed neighboring luma samples is not larger than the luma value of the remaining reconstructed neighboring luma samples of a set of the reconstructed neighboring luma samples. M and N are positive integers and are greater than 1. One or more linear model parameters are determined based on the two pairs of the luma value and the chroma value, and a predictive block is determined based on the one or more linear model parameters.SELECTED DRAWING: Figure 23
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] TECHNICAL FIELD Embodiments of the present disclosure relate generally to video data encoding and decoding techniques, and more particularly to techniques for intra prediction using cross-component linear models (CCLMs). [Background technology]

[0002] Digital video has been widely used since the introduction of the Digital Versatile Disc (DVD). In addition to distributing video programs using DVDs, video programs can now be transmitted using wired computer networks (e.g., the Internet) or wireless communication networks. Before transmitting video data over a transmission medium, the video is encoded. A viewer receives the encoded video and decodes and displays the video using a viewing device. Over the years, video quality has improved, for example, due to higher resolutions, color depths, and frame rates. The improved quality of transmitted video data has led to larger data streams, which are now commonly carried over the Internet and mobile communication networks.

[0003] Higher resolution videos typically require more bandwidth because they carry more information. To reduce bandwidth requirements, video coding schemes have been introduced that involve compressing the video. When video is coded, the bandwidth requirements (or corresponding memory requirements in the case of storage) are reduced compared to uncoded video. Often, this reduction comes at the expense of quality. Thus, video coding standards strive to find a balance between bandwidth requirements and quality.

[0004] Because there is a continuing need to improve quality and reduce bandwidth requirements, solutions are constantly being sought that maintain quality while reducing bandwidth requirements, or improve quality while maintaining bandwidth requirements. Sometimes a compromise between the two is acceptable. For example, if improved quality is important, then increasing bandwidth requirements may be acceptable.

[0005] High Efficiency Video Coding (HEVC) is a widely known video coding scheme. In HEVC, a coding unit (CU) is divided into multiple prediction units (PUs) or transform units (TUs). The Versatile Video Coding (VVC) standard, a next-generation video coding standard, is a recent joint video coding project between the Video Coding Experts Group (VCEG) of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC). The two standardization organizations collaborate in a partnership known as the Joint Video Exploration Team (JVET). The VVC standard is also known as the ITU-T H.266 standard or the Next Generation Video Coding (NGVC) standard. In the VVC standard, the concept of multiple partition types, i.e., the partitioning of CU, PU, ​​and TU concepts, is removed except when necessary for CUs whose size is too large for the maximum transform length, and VVC supports greater flexibility in CU partitioning shapes. Summary of the Invention

[0006]

[0006] Embodiments of the present application provide an apparatus and method for encoding and decoding video data. In particular, by using luma samples to predict chroma samples by intra prediction as part of a video coding mechanism, intra predictive coding using a cross-component linear model can be realized in a flexible manner.

[0007] Particular embodiments are set out in the accompanying independent claims, and other embodiments are set out in the dependent claims.

[0008] According to a first aspect, the present disclosure relates to a method for decoding video data, the method comprising: determining a luma block corresponding to a chroma block; determining a set of downsampled reconstructed neighboring luma samples, the set of downsampled reconstructed neighboring luma samples comprising a plurality of downsampled reconstructed luma samples above the luma block and / or a plurality of downsampled reconstructed luma samples to the left of the luma block; determining two pairs of luma and chroma values ​​according to the N reconstructed chroma samples corresponding to the N downsampled reconstructed luma samples with the largest values ​​and / or the M reconstructed chroma samples corresponding to the M downsampled reconstructed luma samples with the smallest values, where M and N are positive integers greater than 1, when the N downsampled reconstructed luma samples with the largest values ​​and / or the M downsampled reconstructed luma samples with the smallest values ​​are included in the set of downsampled reconstructed neighboring luma samples; determining one or more linear model parameters based on the determined two pairs of luma and chroma values; determining a prediction block based on one or more linear model parameters; Linear model (LM) predictive decoding of chroma blocks based on predicted blocks; and Linear model (LM) prediction includes cross-component linear model prediction, multi-directional linear model (MDLM) and MMLM.

[0009] As such, in a possible embodiment of the method according to the first aspect, the set of downsampled reconstructed neighboring luma samples may be an upper right neighboring luma sample outside the luma block and a luma sample to the right of the upper right neighboring luma sample outside the luma block; The luma sample below the lower-left neighboring luma sample outside the luma block and the luma sample below the lower-left neighboring luma sample outside the luma block. It further has:

[0010] As such, in possible embodiments of the method according to the first aspect or any of the above implementations of the first aspect, the plurality of reconstructed luma samples above the luma block are reconstructed neighboring luma samples proximate their respective top boundaries, and the plurality of reconstructed luma samples to the left of the luma block are reconstructed neighboring luma samples proximate their respective left boundaries.

[0011] As such, in possible embodiments of the method according to the first aspect or any of the above implementations of the first aspect, the set of (downsampled) reconstructed neighboring luma samples excludes luma samples that are above the upper-left neighboring luma sample outside the luma block and / or luma samples that are to the left of the upper-left neighboring luma sample outside the luma block.

[0012] As such, in a possible embodiment of a method according to the first aspect or any of the above implementations of the first aspect, the coordinates of the top-left luma sample of a luma block are (x0, y0), and the set of (downsampled) reconstructed neighboring luma samples excludes luma samples with x coordinates smaller than x0 and y coordinates smaller than y0.

[0013] As such, in a possible embodiment of the method according to the first aspect or any of the above implementations of the first aspect, the step of determining two pairs of luma and chroma values ​​when the N downsampled reconstructed luma samples with the largest values ​​and / or the M downsampled reconstructed luma samples with the smallest values ​​are included in the set of downsampled reconstructed neighboring luma samples comprises: determining (or selecting) two pairs of luma and chroma values ​​based on a chroma value difference between each chroma value of a first plurality of pairs of luma and chroma values ​​and each chroma value of a second plurality of pairs of luma and chroma values; Each of the first plurality of pairs of luma and chroma values ​​has one of the N downsampled reconstructed luma samples having a maximum value and a corresponding reconstructed adjacent chroma sample, and each of the second plurality of pairs of luma and chroma values ​​has one of the M downsampled reconstructed luma samples having a minimum value and a corresponding reconstructed adjacent chroma sample.

[0014] As such, in a possible embodiment of the method according to the first aspect or any of the above implementations of the first aspect, there is a minimum chroma value difference between a chroma value of the first pair of luma and chroma values ​​and a chroma value of the second pair of luma and chroma values, and the first pair of luma and chroma values ​​and the second pair of luma and chroma values ​​having the minimum chroma value difference are selected as the two pairs of luma and chroma values; or There is a maximum chroma value difference between the chroma value of the third pair of luma and chroma values ​​and the chroma value of the fourth pair of luma and chroma values, and the third pair of luma and chroma values ​​and the fourth pair of luma and chroma values ​​having the maximum chroma value difference are selected as two pairs of luma and chroma values. For example, the first pair of luma and chroma values ​​is included in the first plurality of pairs of luma and chroma values. For example, the second pair of luma and chroma values ​​is included in the second plurality of pairs of luma and chroma values. For example, the third pair of luma and chroma values ​​is included in the first plurality of pairs of luma and chroma values. For example, the fourth pair of luma and chroma values ​​is included in the second plurality of pairs of luma and chroma values.

[0015] As such, in a possible embodiment of the method according to the first aspect or any of the above implementations of the first aspect, the step of determining two pairs of luma and chroma values ​​when the N downsampled reconstructed luma samples with the largest values ​​and / or the M downsampled reconstructed luma samples with the smallest values ​​are included in the set of downsampled reconstructed neighboring luma samples comprises: determining a fifth pair of luma and chroma values ​​and a sixth pair of luma and chroma values ​​as two pairs of luma and chroma values; The corresponding chroma values ​​in the fifth pair of luma and chroma values ​​are the average chroma values ​​of the N reconstructed chroma samples corresponding to the N downsampled reconstructed luma samples with the largest values, and the corresponding chroma values ​​in the sixth pair of luma and chroma values ​​are the average chroma values ​​of the M reconstructed chroma samples corresponding to the M downsampled reconstructed luma samples with the smallest values. The luma values ​​in the fifth pair of luma and chroma values ​​are the luma values ​​of each of the N downsampled reconstructed luma samples with the largest values. The luma values ​​in the sixth pair of luma and chroma values ​​are the luma values ​​of each of the M downsampled reconstructed luma samples with the smallest values.

[0016] As such, in a possible embodiment of a method according to the first aspect or any of the above implementations of the first aspect, the set of downsampled reconstructed neighboring luma samples comprises a first set of downsampled reconstructed neighboring luma samples and a second set of downsampled reconstructed neighboring luma samples, the first set of downsampled reconstructed neighboring luma samples comprising downsampled reconstructed neighboring luma samples having luma values ​​that are less than or equal to a threshold, and the second set of downsampled reconstructed neighboring luma samples comprising downsampled reconstructed neighboring luma samples having luma values ​​that are greater than the threshold.

[0017] As such, in a possible embodiment of a method according to the first aspect or any of the above implementations of the first aspect, LM predictive decoding of a chroma block based on a predictive block comprises adding the predictive block to a residual block to reconstruct the chroma block.

[0018] As such, in a possible embodiment of the method according to the first aspect or any of the above implementations of the first aspect, the method comprises: decoding a flag of a current block including a luma block and a chroma block; The flag indicates that LM predictive coding is enabled for the chroma block, and decoding the flag includes decoding the flag based on a context having one or more flags that indicate whether LM predictive coding is enabled for neighboring blocks.

[0019] According to a second aspect, the present invention relates to a method for decoding video data, the method comprising the steps of: determining a luma block corresponding to a chroma block; determining a set of (downsampled) reconstructed neighboring luma samples, the set of (downsampled) reconstructed neighboring luma samples comprising a plurality of (downsampled) reconstructed luma samples above the luma block and / or a plurality of (downsampled) reconstructed luma samples to the left of the luma block; determining two pairs of luma values ​​and chroma values ​​according to the N (downsampled) reconstructed luma samples and N reconstructed chroma samples corresponding to the N (downsampled) reconstructed luma samples, and / or M (downsampled) reconstructed luma samples and M reconstructed chroma samples corresponding to the M (downsampled) reconstructed luma samples, wherein the minimum value of the N (downsampled) reconstructed luma samples is the minimum value of the remaining (downsampled) reconstructed luma values ​​of the set of (downsampled) reconstructed adjacent luma samples; the luma value of any one of the M (downsampled) reconstructed luma samples is not smaller than the luma value of any one of the remaining (downsampled) reconstructed luma samples of the set of (downsampled) reconstructed luma samples, and the maximum value of the M (downsampled) reconstructed luma samples is not greater than the luma values ​​of the remaining (downsampled) reconstructed luma samples of the set of (downsampled) reconstructed luma samples, M, N are positive integers greater than 1, i.e., the luma value of any one of the N downsampled reconstructed luma samples is greater than the luma value of any one of the M downsampled reconstructed luma samples, and the sum of N and M is less than or equal to the number of sets of downsampled reconstructed neighboring luma samples; determining one or more linear model parameters based on the determined two pairs of luma and chroma values; determining a prediction block based on one or more linear model parameters; Linear model (LM) predictive decoding of chroma blocks based on predicted blocks; It has.

[0020] As such, in a possible embodiment of the method according to the second aspect, the step of determining two pairs of luma and chroma values ​​comprises: determining a seventh pair of luma and chroma values ​​and an eighth pair of luma and chroma values ​​as two pairs of luma and chroma values; The luma value of the seventh pair of luma and chroma values ​​is the average luma value of the N (downsampled) reconstructed luma samples, the chroma value of the seventh pair of luma and chroma values ​​is the average chroma value of the N reconstructed chroma samples corresponding to the N (downsampled) reconstructed luma samples, the luma value of the eighth pair of luma and chroma values ​​is the average luma value of the M (downsampled) reconstructed luma samples, and the chroma value of the eighth pair of luma and chroma values ​​is the average chroma value of the M reconstructed chroma samples corresponding to the M (downsampled) reconstructed luma samples.

[0021] As such, in a possible embodiment of the method according to the second aspect or any of the above implementations of the second aspect, the step of determining two pairs of luma and chroma values ​​comprises: determining a ninth pair of luma and chroma values ​​and a tenth pair of luma and chroma values ​​as two pairs of luma and chroma values; the luma value of the ninth pair of luma and chroma values ​​is the average luma value of the N (downsampled) reconstructed luma samples in the first luma value range, and the chroma value of the ninth pair of luma and chroma values ​​is the average chroma value of the N reconstructed chroma samples corresponding to the N (downsampled) reconstructed luma samples in the first luma value range; The luma value of the tenth pair of luma and chroma values ​​is the average luma value of the M (downsampled) reconstructed luma samples in the second luma value range, and the chroma value of the tenth pair of luma and chroma values ​​is the average chroma value of the M reconstructed chroma samples corresponding to the M (downsampled) reconstructed luma samples in the second luma value range, and no value in the first luma value range is greater than any one of the second luma value range.

[0022] As such, in a possible embodiment of the method according to the second aspect or any of the above implementations of the second aspect, the first luma value is in the range [MaxlumaValue-T1, MaxlumaValue] and / or the second luma value is in the range [MinlumaValue, MinlumaValue+T2]; MaxlumaValue and MinlumaValue are the maximum and minimum luma values, respectively, among the set of downsampled reconstructed neighboring luma samples, and T1, T2 are predefined thresholds.

[0023] As such, in possible embodiments of the method according to the second aspect or any of the above implementations of the second aspect, M and N are equal or unequal.

[0024] As such, in a possible embodiment of the method according to the second aspect or any of the above implementations of the second aspect, M and N are defined based on the block size of the luma blocks.

[0025] As such, in a possible embodiment of a method according to the second aspect or any of the above implementations of the second aspect, M=(W+H)>>t, N=(W+H)>>r, where t and r are the number of right shift bits.

[0026] As such, in a possible embodiment of a method according to the second aspect or any of the above implementations of the second aspect, the set of downsampled reconstructed neighboring luma samples may be an upper right neighboring luma sample outside the luma block and a luma sample to the right of the upper right neighboring luma sample outside the luma block; The bottom-left neighboring luma sample outside the luma block and the luma sample below the bottom-left neighboring luma sample outside the luma block It further has:

[0027] As such, in possible embodiments of the method according to the second aspect or any of the above implementations of the second aspect, the plurality of reconstructed luma samples above the luma block are reconstructed neighboring luma samples proximate their respective top boundaries, and the plurality of reconstructed luma samples to the left of the luma block are reconstructed neighboring luma samples proximate their respective left boundaries.

[0028] As such, in possible embodiments of the method according to the second aspect or any of the above implementations of the second aspect, the set of (downsampled) reconstructed neighboring luma samples excludes luma samples that are above the upper-left neighboring luma sample outside the luma block and / or luma samples that are to the left of the upper-left neighboring luma sample outside the luma block.

[0029] As such, in a possible embodiment of a method according to the second aspect or any of the above implementations of the second aspect, the coordinates of the top-left luma sample of a luma block are (x0, y0), and the set of (downsampled) reconstructed neighboring luma samples excludes luma samples with x coordinates smaller than x0 and y coordinates smaller than y0.

[0030] As such, in a possible embodiment of a method according to the second aspect or any of the above implementations of the second aspect, the set of downsampled reconstructed neighboring luma samples comprises a first set of downsampled reconstructed neighboring luma samples and a second set of downsampled reconstructed neighboring luma samples, the first set of downsampled reconstructed neighboring luma samples comprising downsampled reconstructed neighboring luma samples having luma values ​​that are less than or equal to a threshold, and the second set of downsampled reconstructed neighboring luma samples comprising downsampled reconstructed neighboring luma samples having luma values ​​that are greater than the threshold.

[0031] According to a third aspect, the present invention relates to a device for decoding video data, the device comprising: a video data memory; a video decoder, the video decoder comprising: determining a luma block corresponding to the chroma block; determining a set of (downsampled) reconstructed neighboring luma samples, the set of (downsampled) reconstructed neighboring luma samples comprising a plurality of (downsampled) reconstructed luma samples above the luma block and / or a plurality of (downsampled) reconstructed luma samples to the left of the luma block; determining two pairs of luma and chroma values ​​according to the N reconstructed chroma samples corresponding to the N downsampled reconstructed luma samples with the largest values ​​and the N downsampled reconstructed luma samples with the largest values ​​and / or the M reconstructed chroma samples corresponding to the M downsampled reconstructed luma samples with the smallest values, where M, N are positive integers greater than 1, when the N downsampled reconstructed luma samples with the largest values ​​and / or the M downsampled reconstructed luma samples with the smallest values ​​are included in the set of downsampled reconstructed neighboring luma samples; determining one or more linear model parameters based on the determined two pairs of luma and chroma values; determining a prediction block based on one or more linear model parameters; Linear model (LM) predictive decoding of chroma blocks based on predicted blocks It is configured as follows.

[0032] As such, in a possible embodiment of the device according to the third aspect, the set of downsampled reconstructed adjacent luma samples may be an upper right neighboring luma sample outside the luma block and a luma sample to the right of the upper right neighboring luma sample outside the luma block; The luma sample below the lower-left neighboring luma sample outside the luma block and the luma sample below the lower-left neighboring luma sample outside the luma block. It further has:

[0033] As such, in possible embodiments of the method according to the third aspect or any of the above implementations of the third aspect, the plurality of reconstructed luma samples above the luma block are reconstructed neighboring luma samples proximate their respective top boundaries, and the plurality of reconstructed luma samples to the left of the luma block are reconstructed neighboring luma samples proximate their respective left boundaries.

[0034] As such, in possible embodiments of the method according to the third aspect or any of the above implementations of the third aspect, the set of (downsampled) reconstructed neighboring luma samples excludes luma samples that are above the upper-left neighboring luma sample outside the luma block and / or luma samples that are to the left of the upper-left neighboring luma sample outside the luma block.

[0035] As such, in a possible embodiment of a method according to the third aspect or any of the above implementations of the third aspect, the coordinates of the top-left luma sample of a luma block are (x0, y0), and the set of (downsampled) reconstructed neighboring luma samples excludes luma samples with x coordinates smaller than x0 and y coordinates smaller than y0.

[0036] As such, in a possible embodiment of the method according to the third aspect or any of the above implementations of the third aspect, when the N downsampled reconstructed luma samples having the largest values ​​and / or the M downsampled reconstructed luma samples having the smallest values ​​are included in the set of downsampled reconstructed neighboring luma samples, to determine two pairs of luma and chroma values, the video decoder: configured to determine (or select) two pairs of luma and chroma values ​​based on a chroma value difference between a chroma value of each of the first plurality of pairs of luma and chroma values ​​and a chroma value of each of the second plurality of pairs of luma and chroma values; Each of the first plurality of pairs of luma and chroma values ​​has one of the N downsampled reconstructed luma samples having a maximum value and a corresponding reconstructed adjacent chroma sample, and each of the second plurality of pairs of luma and chroma values ​​has one of the M downsampled reconstructed luma samples having a minimum value and a corresponding reconstructed adjacent chroma sample.

[0037] As such, in a possible embodiment of the method according to the third aspect or any of the above implementations of the third aspect, there is a minimum chroma value difference between a chroma value of the first pair of luma and chroma values ​​and a chroma value of the second pair of luma and chroma values, and the first pair of luma and chroma values ​​and the second pair of luma and chroma values ​​having the minimum chroma value difference are selected as the two pairs of luma and chroma values; or There is a maximum chroma value difference between the chroma value of the third pair of luma and chroma values ​​and the chroma value of the fourth pair of luma and chroma values, and the third pair of luma and chroma values ​​and the fourth pair of luma and chroma values ​​having the maximum chroma value difference are selected as two pairs of luma and chroma values. For example, the first pair of luma and chroma values ​​is included in the first plurality of pairs of luma and chroma values. For example, the second pair of luma and chroma values ​​is included in the second plurality of pairs of luma and chroma values. For example, the third pair of luma and chroma values ​​is included in the first plurality of pairs of luma and chroma values. For example, the fourth pair of luma and chroma values ​​is included in the second plurality of pairs of luma and chroma values.

[0038] As such, in a possible embodiment of the method according to the third aspect or any of the above implementations of the third aspect, when the N downsampled reconstructed luma samples having the largest values ​​and / or the M downsampled reconstructed luma samples having the smallest values ​​are included in the set of downsampled reconstructed neighboring luma samples, to determine two pairs of luma and chroma values, the video decoder: configured to determine a fifth pair of luma and chroma values ​​and a sixth pair of luma and chroma values ​​as two pairs of luma and chroma values; The corresponding chroma values ​​in the fifth pair of luma and chroma values ​​are the average chroma values ​​of the N reconstructed chroma samples corresponding to the N downsampled reconstructed luma samples with the largest values, and the corresponding chroma values ​​in the sixth pair of luma and chroma values ​​are the average chroma values ​​of the M reconstructed chroma samples corresponding to the M downsampled reconstructed luma samples with the smallest values. The luma values ​​in the fifth pair of luma and chroma values ​​are the luma values ​​of each of the N downsampled reconstructed luma samples with the largest values. The luma values ​​in the sixth pair of luma and chroma values ​​are the luma values ​​of each of the M downsampled reconstructed luma samples with the smallest values.

[0039] As such, in possible embodiments of a method according to the third aspect or any of the above implementations of the third aspect, the set of downsampled reconstructed neighboring luma samples comprises a first set of downsampled reconstructed neighboring luma samples and a second set of downsampled reconstructed neighboring luma samples, the first set of downsampled reconstructed neighboring luma samples comprising downsampled reconstructed neighboring luma samples having luma values ​​that are less than or equal to a threshold, and the second set of downsampled reconstructed neighboring luma samples comprising downsampled reconstructed neighboring luma samples having luma values ​​that are greater than the threshold.

[0040] As such, in a possible embodiment of a method according to the third aspect or any of the above implementations of the third aspect, in order to LM predictively decode a chroma block based on the predictive block, the video decoder is configured to add the predictive block to the residual block to reconstruct the chroma block.

[0041] As such, in a possible embodiment of a method according to the third aspect or any of the above implementations of the third aspect, the video decoder configured to decode a flag of a current block including a luma block and a chroma block; The flag indicates that LM predictive coding is enabled for the chroma block, and decoding the flag includes decoding the flag based on a context having one or more flags that indicate whether LM predictive coding is enabled for neighboring blocks.

[0042] According to a fourth aspect, the present invention relates to a device for decoding video data, the device comprising: a video data memory and a video decoder; The video decoder determining a luma block corresponding to the chroma block; determining a set of (downsampled) reconstructed neighboring luma samples, the set of (downsampled) reconstructed neighboring luma samples comprising a plurality of (downsampled) reconstructed luma samples above the luma block and / or a plurality of (downsampled) reconstructed luma samples to the left of the luma block; Determine two pairs of luma values ​​and chroma values ​​according to the N (downsampled) reconstructed luma samples and N reconstructed chroma samples corresponding to the N (downsampled) reconstructed luma samples, and / or the M (downsampled) reconstructed luma samples and M reconstructed chroma samples corresponding to the M (downsampled) reconstructed luma samples, and the minimum value of the N (downsampled) reconstructed luma samples is determined based on the remaining (downsampled) reconstructed luma samples of the set of (downsampled) reconstructed adjacent luma samples. the luma value of any one of the M (downsampled) reconstructed luma samples is not smaller than the luma value of any one of the M (downsampled) reconstructed luma samples in the set of adjacent (downsampled) reconstructed luma samples, and M and N are positive integers greater than 1, i.e., the luma value of any one of the N downsampled reconstructed luma samples is greater than the luma value of any one of the M downsampled reconstructed luma samples, and the sum of N and M is less than or equal to the number of sets of adjacent downsampled reconstructed luma samples; determining one or more linear model parameters based on the determined two pairs of luma and chroma values; determining a prediction block based on one or more linear model parameters; Linear model (LM) predictive decoding of chroma blocks based on predicted blocks It is configured as follows.

[0043] As such, in a possible embodiment of the method according to the fourth aspect, in order to determine the two pairs of luma and chroma values, the video decoder configured to determine a seventh pair of luma and chroma values ​​and an eighth pair of luma and chroma values ​​as two pairs of luma and chroma values; the luma value of the seventh pair of luma and chroma values ​​is the average luma value of the N (downsampled) reconstructed luma samples, and the chroma value of the seventh pair of luma and chroma values ​​is the average chroma value of the N reconstructed chroma samples corresponding to the N (downsampled) reconstructed luma samples; The luma value of the eighth pair of luma and chroma values ​​is the average luma value of the M (downsampled) reconstructed luma samples, and the chroma value of the eighth pair of luma and chroma values ​​is the average chroma value of the M reconstructed chroma samples corresponding to the M (downsampled) reconstructed luma samples.

[0044] As such, in a possible embodiment of a method according to the fourth aspect or any of the above implementations of the fourth aspect, to determine the two pairs of luma and chroma values, the video decoder configured to determine a ninth pair of luma and chroma values ​​and a tenth pair of luma and chroma values ​​as two pairs of luma and chroma values; the luma value of the ninth pair of luma and chroma values ​​is the average luma value of the N (downsampled) reconstructed luma samples in the first luma value range, and the chroma value of the ninth pair of luma and chroma values ​​is the average chroma value of the N reconstructed chroma samples corresponding to the N (downsampled) reconstructed luma samples in the first luma value range; The luma value of the tenth pair of luma and chroma values ​​is the average luma value of the M (downsampled) reconstructed luma samples in the second luma value range, and the chroma value of the tenth pair of luma and chroma values ​​is the average chroma value of the M reconstructed chroma samples corresponding to the M (downsampled) reconstructed luma samples in the second luma value range, and no value in the first luma value range is greater than any one of the second luma value range.

[0045] As such, in possible embodiments of the method according to the fourth aspect or any of the above implementations of the fourth aspect, the first luma value is in the range [MaxlumaValue-T1, MaxlumaValue] and / or the second luma value is in the range [MinlumaValue, MinlumaValue+T2], where MaxlumaValue and MinlumaValue are the maximum and minimum luma values, respectively, among the set of downsampled reconstructed neighboring luma samples, and T1, T2 are predefined thresholds.

[0046] As such, in possible embodiments of the method according to the fourth aspect or any of the above implementations of the fourth aspect, M and N are equal or unequal.

[0047] As such, in a possible embodiment of the method according to the fourth aspect or any of the above implementations of the fourth aspect, M and N are defined based on the block size of the luma blocks.

[0048] As such, in a possible embodiment of a method according to the fourth aspect or any of the above implementations of the fourth aspect, M=(W+H)>>t, N=(W+H)>>r, where t and r are the number of right shift bits.

[0049] As such, in a possible embodiment of a method according to the fourth aspect or any of the above implementations of the fourth aspect, the set of downsampled reconstructed neighboring luma samples may be an upper right neighboring luma sample outside the luma block and a luma sample to the right of the upper right neighboring luma sample outside the luma block; The bottom-left neighboring luma sample outside the luma block and the luma sample below the bottom-left neighboring luma sample outside the luma block It further has:

[0050] As such, in possible embodiments of a method according to the fourth aspect or any of the above implementations of the fourth aspect, the plurality of reconstructed luma samples above the luma block are reconstructed neighboring luma samples proximate their respective top boundaries, and the plurality of reconstructed luma samples to the left of the luma block are reconstructed neighboring luma samples proximate their respective left boundaries.

[0051] As such, in possible embodiments of the method according to the fourth aspect or any of the above implementations of the fourth aspect, the set of (downsampled) reconstructed neighboring luma samples excludes luma samples that are above the upper-left neighboring luma sample outside the luma block and / or luma samples that are to the left of the upper-left neighboring luma sample outside the luma block.

[0052] As such, in a possible embodiment of a method according to the fourth aspect or any of the above implementations of the fourth aspect, the coordinates of the top-left luma sample of a luma block are (x0, y0), and the set of (downsampled) reconstructed neighboring luma samples excludes luma samples with x coordinates smaller than x0 and y coordinates smaller than y0.

[0053] As such, in possible embodiments of a method according to the fourth aspect or any of the above implementations of the fourth aspect, the set of downsampled reconstructed neighboring luma samples comprises a first set of downsampled reconstructed neighboring luma samples and a second set of downsampled reconstructed neighboring luma samples, the first set of downsampled reconstructed neighboring luma samples comprising downsampled reconstructed neighboring luma samples having luma values ​​that are less than or equal to a threshold, and the second set of downsampled reconstructed neighboring luma samples comprising downsampled reconstructed neighboring luma samples having luma values ​​that are greater than the threshold.

[0054] As such, in a possible embodiment of a method according to the fourth aspect or any of the above implementations of the fourth aspect, to LM predictively decode a chroma block based on the predictive block, the video decoder is configured to add the predictive block to the residual block to reconstruct the chroma block.

[0055] As such, in a possible embodiment of a method according to the fourth aspect or any of the above implementations of the fourth aspect, the video decoder configured to decode a flag of a current block including a luma block and a chroma block; The flag indicates that LM predictive coding is enabled for the chroma block, and decoding the flag includes decoding the flag based on a context having one or more flags that indicate whether LM predictive coding is enabled for neighboring blocks.

[0056] The method according to the first aspect of the invention can be performed by a device according to the third aspect of the invention. Further features and embodiments of the device according to the third aspect of the invention correspond to the features and embodiments of the method according to the first aspect of the invention.

[0057] The method according to the second aspect of the invention can be performed by a device according to the fourth aspect of the invention. Further features and embodiments of the device according to the fourth aspect of the invention correspond to the features and embodiments of the method according to the second aspect of the invention.

[0058] According to a fifth aspect, the present invention relates to a method for encoding video data, the method comprising the steps of: determining a luma block corresponding to a chroma block; determining a set of downsampled reconstructed neighboring luma samples, the set of downsampled reconstructed neighboring luma samples comprising a plurality of downsampled reconstructed luma samples above the luma block and / or a plurality of downsampled reconstructed luma samples to the left of the luma block; determining two pairs of luma and chroma values ​​according to the N reconstructed chroma samples corresponding to the N downsampled reconstructed luma samples with the largest values ​​and / or the M reconstructed chroma samples corresponding to the M downsampled reconstructed luma samples with the smallest values, where M and N are positive integers greater than 1, when the N downsampled reconstructed luma samples with the largest values ​​and / or the M downsampled reconstructed luma samples with the smallest values ​​are included in the set of downsampled reconstructed neighboring luma samples; determining one or more linear model parameters based on the determined two pairs of luma and chroma values; determining a prediction block based on one or more linear model parameters; Linear model (LM) predictive coding of chroma blocks based on the predicted blocks; It has.

[0059] According to a sixth aspect, the present invention relates to a method for encoding video data, the method comprising the steps of: determining a luma block corresponding to a chroma block; determining a set of (downsampled) reconstructed neighboring luma samples, the set of (downsampled) reconstructed neighboring luma samples comprising a plurality of (downsampled) reconstructed luma samples above the luma block and / or a plurality of (downsampled) reconstructed luma samples to the left of the luma block; determining two pairs of luma values ​​and chroma values ​​according to the N (downsampled) reconstructed luma samples and N reconstructed chroma samples corresponding to the N (downsampled) reconstructed luma samples, and / or the M (downsampled) reconstructed luma samples and M reconstructed chroma samples corresponding to the M (downsampled) luma samples, wherein the minimum value of the N (downsampled) reconstructed luma samples is the minimum value of the remaining (downsampled) reconstructed luma samples among the set of (downsampled) reconstructed adjacent luma samples; the luma value of any one of the M (downsampled) reconstructed luma samples is not smaller than the luma value of any one of the M (downsampled) reconstructed luma samples, the maximum value of the M (downsampled) reconstructed luma samples is not greater than the luma values ​​of the remaining (downsampled) reconstructed luma samples of the set of (downsampled) reconstructed neighboring luma samples, M, N are positive integers greater than 1, i.e., the luma value of any one of the N downsampled reconstructed luma samples is greater than the luma value of any one of the M downsampled reconstructed luma samples, and the sum of N and M is less than or equal to the number of sets of downsampled reconstructed neighboring luma samples; determining one or more linear model parameters based on the determined two pairs of luma and chroma values; determining a prediction block based on one or more linear model parameters; Linear model (LM) predictive coding of chroma blocks based on the predicted blocks; It has.

[0060] According to a seventh aspect, the present invention relates to a device for encoding video data, the device comprising: a video data memory; a video encoder, the video encoder comprising: determining a luma block corresponding to the chroma block; determining a set of downsampled reconstructed neighboring luma samples, the set of downsampled reconstructed neighboring luma samples comprising a plurality of downsampled reconstructed luma samples above the luma block and / or a plurality of downsampled reconstructed luma samples to the left of the luma block; determining two pairs of luma and chroma values ​​according to the N reconstructed chroma samples corresponding to the N downsampled reconstructed luma samples with the largest values ​​and the N downsampled reconstructed luma samples with the largest values ​​and / or the M reconstructed chroma samples corresponding to the M downsampled reconstructed luma samples with the smallest values, where M, N are positive integers greater than 1, when the N downsampled reconstructed luma samples with the largest values ​​and / or the M downsampled reconstructed luma samples with the smallest values ​​are included in the set of downsampled reconstructed neighboring luma samples; determining one or more linear model parameters based on the determined two pairs of luma and chroma values; determining a prediction block based on one or more linear model parameters; Linear model (LM) predictive coding of chroma blocks based on predicted blocks It is configured as follows.

[0061] According to an eighth aspect, the present invention relates to a device for encoding video data, the device comprising: a video data memory; a video encoder, the video encoder comprising: determining a luma block corresponding to the chroma block; determining a set of (downsampled) reconstructed neighboring luma samples, the set of (downsampled) reconstructed neighboring luma samples comprising a plurality of (downsampled) reconstructed luma samples above the luma block and / or a plurality of (downsampled) reconstructed luma samples to the left of the luma block; determine two pairs of luma values ​​and chroma values ​​according to the N (downsampled) reconstructed luma samples and N reconstructed chroma samples corresponding to the N (downsampled) reconstructed luma samples, and / or M reconstructed chroma samples corresponding to the M (downsampled) reconstructed luma samples and the M (downsampled) reconstructed luma samples, wherein a minimum value of the N (downsampled) reconstructed luma samples is not smaller than a luma value of a remaining (downsampled) reconstructed luma sample in the set of adjacent (downsampled) reconstructed luma samples, and a maximum value of the M (downsampled) reconstructed luma samples is not greater than a luma value of a remaining (downsampled) reconstructed luma sample in the set of adjacent (downsampled) reconstructed luma samples, and M and N are positive integers greater than 1; determining one or more linear model parameters based on the determined two pairs of luma and chroma values; determining a prediction block based on one or more linear model parameters; Linear model (LM) predictive coding of chroma blocks based on predicted blocks It is configured as follows.

[0062] An encoding device according to any of the above aspects can be further extended by features of an encoding method according to a corresponding above aspect or implementation thereof to derive further implementations of an encoding device according to any of the above aspects.

[0063] An encoding method according to any of the above aspects can be further extended by features of a decoding method according to a corresponding above aspect or implementation thereof to derive further implementations of an encoding method according to any of the above aspects.

[0064] A computer-readable medium is provided that stores instructions that, when executed on a processor, cause the processor to perform any of the above methods according to any of the above aspects or any of the above implementations of any of the above aspects as such.

[0065] A decoding apparatus is provided, as such having modules / units / components / circuits for performing at least some of the steps of the above method according to any above aspect or any above implementation of any of the above aspects.

[0066] A decoding apparatus is provided, the decoding apparatus having a memory storing instructions and a processor coupled to the memory, the processor configured to execute the instructions stored in the memory so as to cause the processor to perform the above method according to any above aspect or any above implementation of any of the above aspects as such.

[0067] A computer-readable storage medium is provided, having a program recorded on the computer-readable storage medium, the program causing a computer to perform a method according to any above aspect or any above implementation of any above aspect as such.

[0068] A computer program is provided, the computer program being configured to cause a computer to carry out a method according to any above aspect or any above implementation of any above aspect as such.

[0069] For clarity, any one of the above embodiments may be combined with any one or more of the other above embodiments to create new embodiments within the scope of the present disclosure.

[0070] These and other features will be more clearly understood from the following detailed description considered in conjunction with the accompanying drawings and claims.

[0071] The following is a brief description of the accompanying drawings used in describing embodiments of the present application. [Brief explanation of the drawings]

[0072] [Figure 1A] 1 is a block diagram of a video data coding system in which embodiments of the present disclosure may be implemented. [Figure 1B] FIG. 10 is a block diagram of another video data coding system in which embodiments of the present disclosure may be implemented. [Figure 2] 1 is a block diagram of a video data encoder in which embodiments of the present disclosure may be implemented. [Figure 3] FIG. 1 is a block diagram of a video data decoder in which embodiments of the present disclosure may be implemented. [Figure 4] 1 is a schematic diagram of a video coding device according to an embodiment of the present disclosure. [Figure 5] 1 is a simplified block diagram of a video data coding device in which various embodiments of the present disclosure may be implemented. [Figure 6] This is an example of an intra prediction mode in H.265 / HEVC. [Figure 7] 1 is an example of a reference sample. [Figure 8] FIG. 1 is a conceptual diagram illustrating the nominal relative vertical and horizontal positions of luma and chroma samples. [Figure 9]9A and 9B are schematic diagrams illustrating an example of a mechanism for performing cross-component linear model (CCLM) intra prediction. FIG. 9A illustrates an example of neighboring reconstructed pixels of a co-located luma block. FIG. 9B illustrates an example of neighboring reconstructed pixels of a chroma block. [Figure 10] FIG. 1 is a conceptual diagram illustrating an example of luma and chroma positions of downsampled samples of a luma block to generate a predictive block. [Figure 11] FIG. 10 is a conceptual diagram illustrating another example of luma and chroma positions of downsampled samples of a luma block for generating a prediction block. [Figure 12] FIG. 1 is a schematic diagram illustrating an example of a downsampling mechanism for supporting cross-component intra prediction. [Figure 13] FIG. 1 is a schematic diagram illustrating an example of a downsampling mechanism for supporting cross-component intra prediction. [Figure 14] FIG. 1 is a schematic diagram illustrating an example of a downsampling mechanism for supporting cross-component intra prediction. [Figure 15] FIG. 1 is a schematic diagram illustrating an example of a downsampling mechanism for supporting cross-component intra prediction. [Figure 16] 1 is an illustration of a straight line between the minimum and maximum luma values. [Figure 17] 1 is an example of a Cross Component Intra Prediction_A (CCIP_A) mode. [Figure 18] 10 is an example of a cross-component intra-prediction_L (CCIP_L) mode. [Figure 19] 1 is a graph illustrating an example of a mechanism for determining linear model parameters to support multi-model CCLM (MMLM) intra prediction. [Figure 20] 1 is a schematic diagram illustrating an example of a mechanism for using neighboring samples above and to the left to support cross-component intra prediction. [Figure 21]1 is a schematic diagram illustrating an example of a mechanism for using extended samples to support cross-component intra prediction. [Figure 22] 1 is a flowchart of a method of cross-component linear model (CCLM) prediction according to some embodiments of the present disclosure. [Figure 23] 1 is a flowchart of a method for decoding video data using cross-component linear model (CCLM) prediction according to an embodiment of the present disclosure. [Figure 24] 1 is a flowchart of a method for encoding video data using cross-component linear model (CCLM) prediction according to an embodiment of the present disclosure. [Figure 25] 10 is a flowchart of a method for decoding video data using cross-component linear model (CCLM) prediction according to another embodiment of the present disclosure. [Figure 26] 10 is an exemplary flowchart of a method for encoding video data using cross-component linear model (CCLM) prediction according to another embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0073] In the various figures, the same reference numbers will be used for identical or functionally equivalent features.

[0074] While illustrative implementations of one or more embodiments are provided below, it should be understood at the outset that the disclosed systems and / or methods may be implemented using any number of technologies, whether currently known or in existence. The present disclosure should in no way be limited to the illustrative implementations, drawings, and technologies described below, including the example designs and implementations shown and described herein, but may vary within the full scope of the appended claims and their equivalents.

[0075] 1A is a block diagram of a video data coding system in which embodiments of the present disclosure may be implemented. As shown in FIG. 1A, coding system 10 includes a source device 12 that provides encoded video data and a destination device 14 that decodes the encoded video data provided by source device 12. In particular, source device 12 may provide the video data to destination device 14 via a transport medium 16. Source device 12 and destination device 14 may be any of a wide range of electronic devices, such as desktop computers, notebook computers (i.e., laptop computers), tablet computers, set-top boxes, mobile phone handsets (i.e., "smartphone" phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 12 and destination device 14 may be equipped for wireless communication.

[0076] The destination device 14 may receive the encoded video data via a transport medium 16. The transport medium 16 may be any type of medium or device capable of carrying the encoded video data from the source device 12 to the destination device 14. In one example, the transport medium 16 may be a communications medium that enables the source device 12 to transmit the encoded video data directly to the destination device 14 in real time. The encoded video data may be modulated according to a communications standard, such as a wireless communications protocol, and the modulated encoded video data is transmitted to the destination device 14. The communications medium may be any wireless or wired communications medium, such as radio frequency (RF) spectrum waves or one or more physical transmission paths. The communications medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communications medium may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from the source device 12 to the destination device 14.

[0077] At the source device 12, the encoded data may be output from the output interface 22 to a storage device (not shown in FIG. 1A). Similarly, the encoded data may be accessed from the storage device by the input interface 28 of the destination device 14. The storage device may be a hard drive, a Blu-ray TM It may include any of a variety of distributed or locally accessed data storage media, such as a disk, a digital video disk (DVD), a compact disk read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium that stores encoded video data.

[0078] In a further example, the storage device may correspond to a file server or other intermediate storage device that can store the encoded video generated by the source device 12. The destination device 14 may access the stored video data from the storage device by streaming or downloading. The file server may be any type of server capable of storing and transmitting encoded video data to the destination device 14. Exemplary file servers include a web server (e.g., for a website), a file transfer protocol (FTP) server, a network-attached storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data through any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on a file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.

[0079] The techniques of this disclosure are not necessarily limited to wireless applications or settings. The techniques may be applied to video coding in support of any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of video data stored on a data storage medium, or other uses. In some examples, coding system 10 may be configured to support unidirectional or bidirectional video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0080] In the example of FIG. 1A , source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. Consistent with this disclosure, video encoder 20 of source device 12 and / or video decoder 30 of destination device 14 may be configured to apply techniques for bidirectional prediction. In other examples, source device 12 and destination device 14 may include other components or arrangements. For example, source device 12 may receive video data from an external video source, such as an external camera. Similarly, destination device 14 may not include a built-in display device but may interface with an external display device.

[0081] The depicted coding system 10 of FIG. 1A is merely an example. The bidirectional prediction method may be performed by any digital video encoding or decoding device. While the techniques of this disclosure are generally used by video coding devices, the techniques may also be used by video encoders / decoders, commonly referred to as "codecs." Furthermore, the techniques of this disclosure may also be used by video preprocessors. The video encoder and / or video decoder may be a graphics processing unit (GPU) or similar device.

[0082] Source device 12 and destination device 14 are merely examples of encoding / decoding devices in a video data coding system in which source device 12 generates encoded video data for transmission to destination device 14. In some examples, source device 12 and destination device 14 may operate in a substantially symmetric manner such that source device 12 and destination device 14 each include video encoding and decoding components. Thus, coding system 10 may support unidirectional or bidirectional video transmission between video devices 12 and 14, e.g., for video streaming, video playback, video broadcasting, or video conferencing.

[0083] Video source 18 of originating device 12 may include a video capture device such as a video camera, a video archive containing previously captured video, and / or a video feed interface that receives video from a video content provider. As a further alternative, video source 18 may generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video.

[0084] In some cases, when video source 18 is a video camera, source device 12 and destination device 14 may form a so-called camera phone or video phone. As noted above, however, the techniques described in this disclosure may be applicable to video coding generally and may be applied to wireless and / or wired applications. In each case, captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video information may then be output by output interface 22 onto transport medium 16.

[0085] The transport medium 16 may be a transitory medium such as a wireless broadcast or wired network transmission, or may be a hard disk, flash drive, compact disc, digital video disc, Blue-ray TM Transport medium 16 may include a storage medium (i.e., a non-transitory storage medium), such as a disk or other computer-readable medium. In some examples, a network server (not shown) may receive the encoded video data from source device 12 and provide the encoded video data to destination device 14, e.g., via network transmission. Similarly, a computing device at a media production facility, such as a disk stamping facility, may receive the encoded video data from source device 12 and produce disks containing the encoded video data. Thus, transport medium 16 may be understood to include one or more computer-readable media of various forms, in various examples.

[0086] An input interface 28 of destination device 14 receives information from transport medium 16. The information on transport medium 16 may include syntax information defined by video encoder 20, which is also used by video decoder 30, including syntax elements that describe the characteristics and / or processing of blocks and other coded units, e.g., groups of pictures (GOPs). Display device 32 displays the decoded video data to a user and may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0087] Video encoder 20 and video decoder 30 may operate in accordance with a video coding standard, such as the High Efficiency Video Coding (HEVC) standard currently under development, and may conform to the HEVC Test Model (HM). Alternatively, video encoder 20 and video decoder 30 may operate in accordance with other proprietary or industry standards, such as the International Telecommunications Union Telecommunication Standardization Sector (ITU-T) H.264 standard, H.265 / HEVC, alternatively referred to as Motion Picture Expert Group (MPEG)-4, Part 10, Advanced Video Coding (AVC), or extensions of such standards. The techniques provided in this disclosure, however, are not limited to any particular coding standard. Other examples of video coding standards include MPEG-2 and ITU-T H.263. 1A, in some aspects, video encoder 20 and video decoder 30 may be integrated with an audio encoder and decoder, respectively, and may include appropriate multiplexer-demultiplexer (MUX-DEMUX) units or other hardware and software to handle the encoding of both audio and video in a common data stream or separate data streams. Where applicable, the MUX-DEMUX units may conform to the ITU H.223 multiplexer protocol or other protocols such as the User Datagram Protocol (UDP).

[0088] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If the techniques are implemented partially in software, a device may store software instructions on a suitable, non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, either of which may be incorporated as part of a combined encoder / decoder (CODEC) in the respective device. The device including video encoder 20 and / or video decoder 30 may be an integrated circuit, a microprocessor, and / or a wireless communication device such as a mobile phone.

[0089] 1B is a block diagram of an example video coding system 40, including video encoder 20 and / or decoder 30. As shown in FIG. 1B, video coding system 40 may include one or more imaging devices 41, video encoder 20, video decoder 30, antenna 42, one or more processors 43, one or more memory stores 44, and may further include a display device 45.

[0090] As shown, imaging device 41, antenna 42, processing circuitry 46, video encoder 20, video decoder 30, processor 43, memory store 44, and display device 45 may be in communication with one another. Although both video encoder 20 and video decoder 30 are shown, video coding system 40 may include only video encoder 20 or only video decoder 30 in various examples.

[0091] In some examples, antenna 42 of video coding system 40 may be configured to transmit or receive an encoded bitstream of video data. Furthermore, in some examples, display device 45 of video coding system 40 may be configured to present the video data. In some examples, processing circuitry 46 of video coding system 40 may be implemented by a processing unit. The processing unit may include application specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. Video coding system 40 may also include an optional processor 43, which may also include application specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. In some examples, processing circuitry 46 may be implemented by hardware, dedicated video coding hardware, etc. Furthermore, memory store 44 may be any type of memory, such as volatile memory (static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In one example, memory store 44 may be implemented by a cache memory. In some examples, processing circuitry 46 may access memory store 44 (e.g., for implementing an image buffer). In other examples, processing circuitry 46 may include a memory store (e.g., a cache, etc.) for implementing an image buffer, etc.

[0092] In some examples, video encoder 20 implemented by a processing circuit may embody various modules discussed in connection with FIG. 2 and / or any other encoder system or subsystem described herein. The processing circuit may be configured to perform various operations discussed herein.

[0093] Video decoder 30 may be implemented in a similar manner as implemented by processing circuitry 46 to embody the various modules discussed in connection with decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein.

[0094] In some examples, antenna 42 of video coding system 40 may be configured to receive an encoded bitstream of video data. The encoded bitstream may include data, indicators, etc. related to encoding the video frames. Video coding system 40 may also include video decoder 30 coupled to antenna 42 and configured to decode the encoded bitstream. Display device 45 is configured to present the video frames.

[0095] 2 is a block diagram illustrating an example of a video encoder 20 that may implement the techniques of the present application. The video encoder 20 may perform intra- and inter-coding of video blocks within video slices. Intra-coding relies on spatial prediction to reduce or remove spatial redundancy in video within a given video frame or picture. Inter-coding relies on temporal prediction to reduce or remove temporal redundancy in video within adjacent frames or pictures of a video sequence. Intra-mode (I-mode) may refer to any of several spatial-based coding modes. Inter-mode, such as unidirectional prediction (P-mode) or bidirectional prediction (B-mode), may refer to any of several temporal-based coding modes.

[0096] As shown in FIG. 2, video encoder 20 receives a current video block in a video frame to be encoded. In the example of FIG. 2, video encoder 20 includes a mode select unit 40, a reference frame memory 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Mode select unit 40, in turn, includes a motion compensation unit 44, a motion estimation unit 42, an intra prediction unit 46, and a partitioning unit 48. For video block reconstruction, video encoder 20 also includes an inverse quantization unit 58, an inverse transform unit 60, and an adder 62. A deblocking filter (not shown in FIG. 2) may also be included to filter block boundaries to remove block artifacts from the reconstructed video. If desired, the deblocking filter will typically filter the output of adder 62. Additional filters (in-loop or post-loop) may also be used in addition to the deblocking filter. Such a filter is not shown for simplicity; if desired, the output of summer 50 may be filtered (as an in-loop filter).

[0097] During the encoding process, video encoder 20 receives a video frame or slice to be encoded. The frame or slice may be divided into multiple video blocks. Motion estimation unit 42 and motion compensation unit 44 perform inter-predictive coding of the received video block relative to one or more blocks in one or more reference frames to provide temporal prediction. Intra-prediction unit 46 may alternatively perform intra-predictive coding of the received video block relative to one or more neighboring blocks in the same frame or slice as the block to be encoded to provide spatial prediction. Video encoder 20 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0098] Further, partitioning unit 48 may partition blocks of video data into sub-blocks based on an evaluation of a previous partitioning scheme in a previous coding pass. For example, partitioning unit 48 may first partition a frame or slice into largest coding units (LCUs) and then partition each of the LCUs into sub-coding units (sub-CUs) based on a rate-distortion analysis (e.g., rate-distortion optimization). Mode selection unit 40 may further generate a quadtree data structure that indicates the partitioning of the LCUs into sub-CUs. A leaf-node CU of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs).

[0099] In this disclosure, the term "block" is used to refer to either a coding unit (CU), a prediction unit (CU), or a transform unit (TU) in the context of HEVC, or to similar data structures in the context of other standards (e.g., a macroblock and its sub-blocks in H.264 / AVC). A CU includes a coding node, PUs, and TUs associated with the coding node. The size of a CU corresponds to the size of the coding node and is rectangular in shape. The size of a CU may range from 8x8 pixels to the size of a treeblock, which may be up to 64x64 pixels or larger. Each CU may include one or more PUs and one or more TUs. Syntax data associated with a CU may, for example, describe the partitioning of the CU into one or more PUs. The partitioning mode may differ depending on whether the CU is coded in skip or direct mode, intra-prediction mode, or inter-prediction mode. A PU may be partitioned to be non-square in shape. Syntax data associated with a CU may also describe the partitioning of the CU into one or more TUs, for example, according to a quadtree. In embodiments, a CU, PU, ​​or TU can be square or non-square (eg, rectangular) in shape.

[0100] Mode select unit 40 may select one of the coding modes, i.e., intra or inter, based on, for example, the error result, and provides the resulting intra- or inter-coded block to summer 50 for generating residual block data and to summer 62 for reconstructing the coded block to be used as a reference frame. Mode select unit 40 also provides syntax elements, such as motion vectors, intra-mode indicators, partition information, and other such syntax information, to entropy coding unit 56.

[0101] Motion estimation unit 42 and motion compensation unit 44 may be highly integrated but are represented separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. A motion vector may indicate, for example, the displacement of a PU of a video block in a current video frame (or picture) relative to a predictive block in a reference frame (or other coded unit), or may indicate the displacement of a PU of a video block in a current video (or picture) relative to a coded block in the current frame (or other coded unit). A predictive block is a block that is found to be a good match for a block to be coded in terms of pixel difference, which may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some examples, video encoder 20 may calculate values ​​for sub-integer pixel locations of a reference picture stored in reference frame memory 64. For example, video encoder 20 may interpolate values ​​at quarter-pixel positions, eighth-pixel positions, or other fractional-pixel positions of a reference picture. Accordingly, motion estimation unit 42 may perform motion searches for whole-pixel and fractional-pixel positions and output motion vectors that include fractional-pixel predictions.

[0102] Motion estimation unit 42 calculates a motion vector for a PU of a video block in an inter-coded slice by comparing the position of the PU with the position of a predictive block of a reference picture. The reference picture may be selected from a first reference picture list (e.g., list 0) or a second reference picture list (e.g., list 1), each of which identifies one or more reference pictures stored in reference frame memory 64. Motion estimation unit 42 sends the calculated motion vector to entropy encoding unit 56 and motion compensation unit 44.

[0103] The motion compensation performed by motion compensation unit 44 may include fetching or generating a predictive block based on the motion vector determined by motion estimation unit 42. Again, motion estimation unit 42 and motion compensation unit 44 may be functionally integrated in some examples. Upon receiving the motion vector for the PU of the current video block, motion compensation unit 44 may find the predictive block to which the motion vector points in one of the reference picture lists. Adder 50 forms a residual video block by subtracting pixel values ​​of the predictive block from pixel values ​​of the current video block being coded to form pixel difference values, as discussed below. Generally, motion estimation unit 42 performs motion estimation on the luma component, and motion compensation unit 44 uses the motion vector calculated based on the luma component for both the chroma and luma components. Mode select unit 40 may also generate syntax elements associated with the video blocks and video slices for use by video decoder 30 in decoding the video blocks of the video slices.

[0104] Intra prediction unit 46 may intra predict the current block as an alternative to the inter prediction performed by motion estimation unit 42 and motion compensation unit 44, as described above. In particular, intra prediction unit 46 may determine an intra prediction mode to use to encode the current block. In some examples, intra prediction unit 46 may encode the current block using various intra prediction modes, e.g., during separate encoding passes, and intra prediction unit 46 (or, in some examples, mode selection unit 40) may select an appropriate intra prediction mode to use from the tested modes.

[0105] For example, intra prediction unit 46 may calculate rate-distortion values ​​for various tested intra prediction modes using rate-distortion analysis and select the intra prediction mode having the best rate-distortion characteristics among the tested modes. The rate-distortion analysis generally determines the bitrate (i.e., number of bits) used to generate the coded block, along with the amount of distortion (or error) between the coded block and the original uncoded block that was coded to generate the coded block. Intra prediction unit 46 may calculate a ratio from the distortion and rate of the various coded blocks to determine which intra prediction mode exhibits the best rate-distortion value for that block.

[0106] Moreover, intra prediction unit 46 may be configured to code depth blocks of the depth map using a depth modeling mode (DMM). Mode selection unit 40 may determine whether an available DMM mode produces better coding results than intra prediction modes and other DMM modes, for example, using rate-distortion optimization (RDO). Data of texture images corresponding to the depth map may be stored in reference frame memory 64. Motion estimation unit 42 and motion compensation unit 44 may also be configured to inter predict depth blocks of the depth map.

[0107] After selecting an intra-prediction mode for the block (e.g., one of a conventional intra-prediction mode or a DMM mode), intra-prediction unit 46 may provide information indicating the selected intra-prediction mode for the block to entropy encoding unit 56. Entropy encoding unit 56 may encode the information indicating the selected intra-prediction mode. Video encoder 20 may include definitions of coding contexts for various blocks and indications of the most probable intra-prediction mode, intra-prediction mode index table, and modified intra-prediction mode index table to use for each of the contexts in transmitted bitstream configuration data, which may include multiple intra-prediction mode index tables and multiple modified intra-prediction mode index tables (also referred to as codeword mapping tables).

[0108] Video encoder 20 forms a residual video block by subtracting the prediction data from mode select unit 40 from the original video block being coded. Summer 50 represents the component or components that perform this subtraction operation.

[0109] Transform processing unit 52 applies a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform, to the residual block to produce a video block that includes residual transform coefficient values. Transform processing unit 52 may also perform other transforms that are conceptually similar to the DCT. Wavelet transforms, integer transforms, subband transforms, or other types of transforms may also be used.

[0110] Transform processing unit 52 applies a transform to the residual block, generating a block of residual transform coefficients. The transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be changed by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan of a matrix containing the quantized transform coefficients. Alternatively, entropy encoding unit 56 may perform the scan.

[0111] Following quantization, entropy coding unit 56 entropy codes the quantized transform coefficients. For example, entropy coding unit 56 may perform context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. In the case of context-based entropy coding, the context may be based on neighboring blocks. Following entropy coding by entropy coding unit 56, the coded bitstream may be transmitted to another device (e.g., video decoder 30) or archived for later transmission or retrieval.

[0112] Inverse quantization unit 58 and inverse transform unit 60 apply inverse quantization and inverse transformation, respectively, to reconstruct the residual block in the pixel domain, e.g., for later use as a reference block. Motion compensation unit 44 may calculate a reference block by adding the residual block to a predictive block of one of the frames in reference frame memory 64. Motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values ​​for use in motion estimation. Adder 62 adds the reconstructed residual block to the motion-compensated predictive block generated by motion compensation unit 44 to generate a reconstructed video block for storage in reference frame memory 64. The reconstructed video block may be used by motion estimation unit 42 and motion compensation unit 44 as a reference block for inter-coding blocks in subsequent video frames.

[0113] Other structural variations of video encoder 20 can be used to encode the video stream. For example, a non-transform-based encoder 20 can quantize the residual signal directly for a particular block or frame without relying on transform processing unit 52. In other implementations, encoder 20 can have quantization unit 54 and inverse quantization unit 58 combined into a single unit.

[0114] 3 is a block diagram illustrating an example of a video decoder 30 that may implement the techniques of the present application. In the example of FIG. 3, video decoder 30 includes an entropy decoding unit 70, a motion compensation unit 72, an intra prediction unit 74, an inverse quantization unit 76, an inverse transform unit 78, a reference frame memory 82, and an adder 80. Video decoder 30 may, in some examples, perform a decoding path that generally correlates with the encoding path described with respect to video encoder 20 (shown in FIG. 2). Motion compensation unit 72 may generate prediction data based on the motion vector received from entropy decoding unit 70, while intra prediction unit 74 may generate prediction data based on the intra prediction mode indicator received from entropy decoding unit 70.

[0115] During the decoding process, video decoder 30 receives from video encoder 20 an encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements. Entropy decoding unit 70 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. Entropy decoding unit 70 forwards the motion vectors and other syntax elements to motion compensation unit 72. Video decoder 30 may receive syntax elements at the video slice level and / or the video block level.

[0116] If a video slice is encoded as an intra-coded (I) slice, intra prediction unit 74 may generate prediction data for video blocks of the current video slice based on data from previously decoded blocks of the current frame or picture and the signaled intra-prediction mode. If a video frame is encoded as an inter-coded (i.e., B, P, or GPB) slice, motion compensation unit 72 generates prediction blocks for video blocks of the current video slice based on motion vectors and other syntax elements received from entropy decoding unit 70. The prediction blocks may be generated from one of the reference pictures in one of the reference picture lists. Video decoder 30 may construct reference frame lists, e.g., List 0 and List 1, using a default construction technique based on reference pictures stored in reference frame memory 82.

[0117] Motion compensation unit 72 determines prediction information for video blocks of the current video slice by parsing the motion vectors and other syntax elements, and uses the prediction information to generate a prediction block for the current video block being decoded. For example, motion compensation unit 72 uses some of the received syntax elements to determine the prediction mode (e.g., intra- or inter-prediction) used to code the video blocks of the video slice, the inter-prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more reference picture lists for the slice, the motion vector for each inter-coded video block of the slice, the inter-prediction status for each inter-coded video block of the slice, and other information for decoding video blocks in the current video slice.

[0118] Motion compensation unit 72 may also perform interpolation based on an interpolation filter. Motion compensation unit 72 may use the interpolation filter used by video encoder 20 during encoding of the video block to calculate interpolated values ​​for sub-integer pixels of the reference block. In this case, motion compensation unit 72 may determine the interpolation filter used by video encoder 20 from the received syntax element and use the interpolation filter to generate the predictive block.

[0119] Data for texture images corresponding to the depth maps may be stored in reference frame memory 82. Motion compensation unit 72 may also be configured to inter-predict depth blocks of the depth maps.

[0120] As will be appreciated by those skilled in the art, coding system 10 of Figure 1A is suitable for implementing various video coding or compression techniques. Some video compression techniques, such as inter-prediction, intra-prediction, and / or loop filter, are discussed below. Accordingly, video compression techniques are employed in various video coding standards, such as H.264 / AVC and H.265 / HEVC.

[0121] Various coding tools, such as adaptive motion vector prediction (AMVP) and merge mode (MERGE), are used to predict motion vectors (MVs) and improve inter-prediction efficiency and therefore overall video compression efficiency.

[0122] Other variations of video decoder 30 can be used to decode the compressed bitstream. For example, decoder 30 may generate an output video stream without a loop filtering unit. For example, a non-transform-based decoder 30 may inverse quantize the residual signal directly for a particular block or frame without an inverse transform processing unit 78. In other implementations, video decoder 30 may have inverse quantization unit 76 and inverse transform processing unit 78 combined into a single unit.

[0123] 4 is a schematic diagram of a video coding device according to an embodiment of the present disclosure. Video coding device 400 is suitable for implementing the disclosed embodiments described herein. In an embodiment, video coding device 400 may be a decoder such as video decoder 30 of FIG. 1A or an encoder such as video encoder 20 of FIG. 1A. In an embodiment, video coding device 400 may be one or more components of video decoder 30 of FIG. 1A or video encoder 20 of FIG. 1A, as described above.

[0124] Video coding device 400 includes an ingress port 410 and a receiver unit (Rx) 420 for receiving data, a processor 430 (which may be a logic unit or a central processing unit (CPU)) for processing the data, a transmitter unit (Tx) 440 and an egress port 450 for transmitting the data, and a memory 460 for storing the data. Video coding device 400 may also include optical-electrical (OE) and electro-optical (EO) components coupled to ingress port 410, receiver unit 420, transmitter unit 440, and egress port 450 for the entry and exit of optical or electrical signals.

[0125] The processor 430 is implemented in hardware and / or software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGA, ASIC, and DSP. The processor 430 communicates with the ingress port 410, the receiver unit 420, the transmitter unit 440, the egress port 450, and the memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the disclosed embodiments described herein. For example, the coding module 470 implements, processes, prepares, or provides various coding operations. The inclusion of the coding module 470 thus provides substantial improvements to the functionality of the video coding device 400 and enables transformation of the video coding device 400 into different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.

[0126] Memory 460 may include one or more disks, tape drives, and solid-state drives, and may be used as overflow data storage devices for storing programs when such programs are selected for execution, and for storing instructions and data read during program execution. Memory 460 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (SRAM).

[0127] 5 is a schematic block diagram of an apparatus 500 that may be used as either or both of the source device 12 and the destination device 14 from FIG. 1A according to an example embodiment. The apparatus 500 may implement the techniques of the present application. The apparatus 500 may take the form of a computing system including multiple computing devices, or may take the form of a single computing device, such as a mobile phone, tablet computer, laptop computer, notebook computer, desktop computer, etc.

[0128] Processor 502 of device 500 can be a central processing unit. Alternatively, processor 502 can be any other type of device or devices now existing or later developed that are capable of manipulating or processing information. While the disclosed implementations can be performed with a single processor as shown, e.g., processor 502, advantages in speed and efficiency can be achieved using more than one processor.

[0129] The memory 504 in the device 500 can be implemented as a read-only memory (ROM) device or a random-access memory (RAM) device. Any other suitable type of storage device can be used as the memory 504. The memory 504 can be used to store code and / or data 506 accessed by the processor 502 using the bus 512. The memory 504 can also be used to store an operating system 508 and application programs 510. The application programs 510 can include at least one program that enables the processor 502 to perform the methods described herein. For example, the application programs 510 can include multiple applications 1-N and further include a video coding application that performs the methods described herein. The device 500 can also include additional memory in the form of secondary storage 514, which can be, for example, a memory card used with a mobile computing device. Because a video communication session can contain a significant amount of information, it can be stored in whole or in part in the storage 514 and loaded into the memory 504 as needed for processing.

[0130] The device 500 may also include one or more output devices, such as a display 518. The display 518, in one example, may be a touch-sensitive display that combines a display with a touch-sensitive element operable to detect touch input. The display 518 may be coupled to the processor 502 via the bus 512. Other output devices that allow a user to program or otherwise use the device 500 may be provided in addition to or instead of the display 518. When the output device is or includes a display, the display may be implemented in various ways, including by a liquid crystal display (LCD), a cathode ray tube (CRT) display, a plasma display, or a light emitting diode (LED) display, such as an organic LED (OLED) display.

[0131] The device 500 may also include or communicate with an image sensing device 520, such as a camera or any other now existing or later developed image sensing device 520 that is capable of sensing images, such as an image of a user operating the device 500. The image sensing device 520 may be positioned so that it is pointed toward the user operating the device 500. In an example, the position and optical axis of the image sensing device 520 may be configured such that the field of view is immediately adjacent to the display 518 and includes the area where the display 518 is viewable.

[0132] Device 500 may also include or communicate with a sound sensing device 522, such as a microphone or any other now existing or later developed sound sensing device that can detect sound near device 500. Sound sensing device 522 may be positioned such that it is pointed toward a user operating device 500 and may be configured to receive sounds, such as speech or other utterances, made by the user as the user operates device 500.

[0133] While FIG. 5 depicts the processor 502 and memory 504 of device 500 as being integrated into a single device, other configurations are available. The operations of processor 502 can be distributed across multiple machines (each machine having one or more processors) that may be coupled directly or across a local area or other network. Memory 504 can be distributed across multiple machines, such as network-based memory or memory in multiple machines that perform the operations of device 500. While depicted here as a single bus, bus 512 of device 500 may include multiple buses. Furthermore, secondary storage 514 may be directly coupled to other components of device 500 or may be accessible over a network and may include a single integrated unit such as a memory card or multiple units, such as multiple memory cards. Device 500 can be implemented in a wide variety of configurations in this manner.

[0134] This disclosure is concerned with intra prediction as part of the video coding mechanism.

[0135] Intra prediction can be used when there is no reference picture available or when inter-predictive coding is not used for the current block or picture. Reference samples for intra prediction are usually derived from previously coded (or reconstructed) neighboring blocks in the same picture. For example, in both H.264 / AVC and H.265 / HEVC, boundary samples of neighboring blocks are used as references for intra prediction. There are many different intra prediction modes to cover different textures or structural features. Each mode uses a different prediction signal derivation method. For example, H.265 / HEVC supports a total of 35 intra prediction modes, as shown in Figure 6.

[0136] For intra prediction, the composite boundary samples of neighboring blocks are used as references. The encoder selects the best luma intra prediction mode for each block from 35 options: 33 directional prediction modes, DC mode, and planar mode. The mapping between intra prediction direction and intra prediction mode number is specified in Figure 6.

[0137] 7, block "CUR" is the current block to be predicted, and the gray samples along the boundaries of adjacent constructed blocks are used as reference samples. The prediction signal can be derived by mapping the reference samples according to a specific method indicated by the intra prediction mode.

[0138] Video coding may be performed based on color spaces and color formats. For example, color video plays an important role in multimedia systems, and various color spaces are used to efficiently represent colors. A color space specifies a color numerically using multiple components. A well-known color space is the RGB color space, in which a color is represented as a combination of three primary color component values ​​(i.e., red, green, and blue). For color video compression, the YCbCr color space is widely used, as described in A. Ford and A. Roberts, "Color space conversions," University of Westminster, London, Tech. Rep., August 1998.

[0139] YCbCr can be easily converted from the RGB color space via a linear transformation, and redundancies between different components, i.e., cross-component redundancies, are greatly reduced in the YCbCr color space. One advantage of YCbCr is backward compatibility with black-and-white TVs, since the Y signal carries luminance information. Furthermore, chrominance bandwidth can be reduced by subsampling the Cb and Cr components in a 4:2:0 chroma sampling format, with significantly less subjective impact than subsampling in the RGB color space. These advantages have made YCbCr the dominant color space in video compression. There are also other color spaces, such as YCoCg, that are used in video compression. In this disclosure, luma (or L or Y) and two chromas (Cb and Cr) are used to represent the three color components in video compression schemes, regardless of the actual color space used.

[0140] For example, if the chroma format sampling structure is 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array. The nominal vertical and horizontal relative positions of the luma and chroma samples within a picture are shown in Figure 8.

[0141] FIG. 9 (including FIG. 9A and FIG. 9B) is a schematic diagram illustrating an example of a mechanism for performing cross-component linear model (CCLM) intra prediction 900. FIG. 9 illustrates an example of 4:2:0 sampling. FIG. 9 shows example locations of samples of a current block and left and top samples for CCLM mode. The white squares are samples of the current block, and the shaded circles are reconstructed samples. FIG. 9A illustrates an example of neighboring reconstructed pixels of a co-located luma block. FIG. 9B illustrates an example of neighboring reconstructed pixels of a chroma block. When the video format is YUV4:2:0, there is one 16x16 luma block and two 8x8 chroma blocks.

[0142] CCLM intra prediction 900 is a type of cross-component intra prediction. Accordingly, CCLM intra prediction 900 may be performed by intra prediction unit 46 of encoder 20 and / or intra prediction unit 74 of decoder 30. CCLM intra prediction 900 predicts chroma samples 903 in a chroma block 901. Chroma samples 903 appear at integer positions, shown as intersecting lines. The prediction is based in part on neighboring reference samples, shown as black circles. Unlike intra prediction modes, chroma samples 903 are not predicted solely based on neighboring chroma reference samples 905, shown as reconstructed chroma samples (Rec'C). Chroma samples 903 are also predicted based on luma reference samples 913 and neighboring luma reference samples 915. Specifically, a CU includes a luma block 911 and two chroma blocks 901. A model correlating chroma samples 903 and luma reference samples 913 within the same CU is generated. The linear coefficients of the model are determined by comparing adjacent luma reference samples 915 with adjacent chroma reference samples 905 .

[0143] When the luma reference sample 913 is reconstructed, the luma reference sample 913 is represented as a reconstructed luma sample (Rec'L). When the adjacent chroma reference sample 905 is reconstructed, the adjacent chroma reference sample 905 is represented as a reconstructed chroma sample (Rec'C).

[0144] As shown, luma block 911 contains four times as many samples as chroma block 901. Specifically, chroma block 901 contains N samples by N samples, while luma block 911 contains 2N samples by 2N samples. Thus, luma block 911 has four times the resolution of chroma block 901. For prediction to operate on luma reference sample 913 and adjacent luma reference sample 915, luma reference sample 913 and adjacent luma reference sample 915 are downsampled to provide an accurate comparison with adjacent chroma reference sample 905 and chroma sample 903. Downsampling is the process of reducing the resolution of a group of sample values. For example, if the YUV4:2:0 format is used, luma samples may be downsampled by a factor of four (e.g., width by 2 and height by 2). YUV is a color encoding system that uses a color space in terms of a luma component Y and two chrominance components U and V.

[0145] To reduce cross-component redundancy, there is a cross-component linear model (CCLM, LM mode, also called CCIP mode) prediction mode, whereby chroma samples are predicted based on the reconstructed luma samples of the same coding unit (CU) by using a linear model as follows:

number

[0146] In one example, the parameters α and β are derived by minimizing the regression error between neighboring reconstructed luma samples around the current luma block and neighboring reconstructed chroma samples around the chroma block, as follows:

number

[0147] This disclosure relates to using luma samples to predict chroma samples by intra prediction as part of a video coding mechanism. A cross-component linear model (CCLM) prediction mode is added as an additional chroma intra prediction mode. At the encoder side, an additional rate-distortion cost check for the chroma components is added to select the chroma intra prediction mode.

[0148] In general, when a CCLM prediction mode (LM prediction mode for short) is applied, the video encoder 20 and the video decoder 30 may perform the following steps: The video encoder 20 and the video decoder 30 may downsample neighboring luma samples; The video encoder 20 and the video decoder 30 may derive linear parameters (i.e., α and β) (also referred to as scaling parameters or parameters of the cross-component linear model (CCLM) prediction mode); The video encoder 20 and the video decoder 30 may downsample the current luma block and derive a prediction (e.g., a prediction block) based on the downsampled luma block and the linear parameters.

[0149] There can be various methods for downsampling.

[0150] 10 is a conceptual diagram illustrating an example of luma and chroma positions of downsampled samples of a luma block for generating a prediction block for the chroma block. As illustrated in FIG. 10, a chroma sample represented by a filled (i.e., black) triangle is predicted from two luma samples represented by two filled circles by applying a [1,1] filter. The [1,1] filter is an example of a 2-tap filter.

[0151] 11 is a conceptual diagram illustrating another example of luma and chroma positions of downsampled samples of a luma block for generating a prediction block. As shown in FIG. 11, a chroma sample represented by a filled (i.e., black) triangle is predicted from six luma samples represented by six filled circles by applying a 6-tap filter.

[0152] 12-15 are schematic diagrams illustrating example downsampling mechanisms 1200, 1300, 1400, and 1500 to support cross-component intra prediction, e.g., according to CCLM intra prediction 900, mechanism 1600, MDLM intra prediction using CCIP_A mode 1700 and CCIP_L mode 1800, and / or MMLM intra prediction as illustrated in graph 1900. Accordingly, mechanisms 1200, 1300, 1400, and 1500 may be performed by intra prediction unit 46 and / or intra prediction unit 74 of codec system 10 or 40, intra prediction unit 46 of encoder 20, and / or intra prediction unit 74 of decoder 30. Specifically, mechanisms 1200, 1300, 1400, and 1500 may be used during step 2210 of method 200, during step 2320 of method 230 or step 2520 of method 250 at a decoder, and during step 2420 of method 240 or step 2620 of method 260 at an encoder, respectively. Details of Figures 12-15 are provided in International Application No. PCT / US2019 / 041526, filed December 7, 2019, which is incorporated herein by reference.

[0153] In mechanism 1200 of FIG. 12 , two rows 1218 and 1219 of neighboring luma reference samples are downsampled, and three columns 1220, 1221, and 1222 of neighboring luma reference samples are downsampled. Rows 1218 and 1219 and columns 1220, 1221, and 1222 are directly adjacent to luma block 1211, which shares a CU with the chroma block being predicted according to cross-component intra prediction. After downsampling, rows 1218 and 1219 of neighboring luma reference samples become a single row 1216 of downsampled neighboring luma reference samples. Furthermore, columns 1220, 1221, and 1222 of neighboring luma reference samples are downsampled to obtain a single column 1217 of downsampled neighboring luma reference samples. Furthermore, luma samples of luma block 1211 are downsampled to generate downsampled luma reference sample 1212. The downsampled luma reference sample 1212 and the downsampled neighboring luma reference samples from row 1216 and column 1217 may then be used for cross-component intra prediction according to equation (1). It should be noted that the dimensions of rows 1218 and 1219 and columns 1220, 1221, and 1222 may extend beyond the luma block 1211 as shown in FIG. 12. For example, the number of above-neighboring luma reference samples in each row 1218 / 1219, which may be denoted as M, is greater than the number of luma samples in a row of luma block 1211, which may be denoted as W. Furthermore, the number of left-neighboring luma reference samples in each column 1220 / 1221 / 1222, which may be denoted as N, is greater than the number of luma samples in a column of luma block 1211, which may be denoted as H.

[0154] In an example, mechanism 1200 may be implemented as follows: For luma block 1211, two upper neighboring rows 1218 and 1219, denoted as A1 and A2, are used for downsampling to obtain downsampled neighboring row 1216, denoted as A. A[i] is the i-th sample in A, A1[i] is the i-th sample in A1, and A2[i] is the i-th sample in A2. In a specific example, a 6-tap downsampling filter can be applied to neighboring rows 1218 and 1219 to obtain downsampled neighboring row 1216 according to equation (4):

number

[0155] Furthermore, left adjacent columns 1220, 1221, and 1222 are used for downsampling to obtain a downsampled adjacent column 1217, denoted as L1, L2, and L3, where L[i] is the i-th sample in L, L1[i] is the i-th sample in L1, L2[i] is the i-th sample in L2, and L3[i] is the i-th sample in L3. In a specific example, a 6-tap downsampling filter can be applied to adjacent columns 1220, 1221, and 1222 to obtain a downsampled adjacent column 1217 according to equation (5):

number

[0156] Mechanism 1300 of Figure 13 is substantially similar to mechanism 1200 of Figure 12. Mechanism 1300 includes luma block 1311, which is similar to luma block 1211, rows 1218 and 1219, and columns 1220, 1221, and 1222, respectively, and adjacent rows 1318 and 1319 and columns 1320, 1321, and 1322 of neighboring luma reference samples. The difference is that rows 1318 and 1319 and columns 1320, 1321, and 1322 do not extend beyond luma block 1311. As seen in mechanism 1200, luma block 1311, rows 1318 and 1319, and columns 1320, 1321, and 1322 are downsampled to generate downsampled luma reference sample 1312 and column 1317 and row 1316, which contain downsampled neighboring luma reference samples. Column 1317 and row 1316 do not extend beyond the block of downsampled luma reference sample 1312. In other respects, downsampled luma reference sample 1312, column 1317, and row 1316 are substantially similar to downsampled luma reference sample 1212, column 1217, and row 1216, respectively.

[0157] Mechanism 1400 of Figure 14 is similar to mechanisms 1200 and 1300, but uses a single row 1418 of neighboring luma reference samples instead of two rows. Mechanism 1400 also uses three columns 1420, 1421, and 1422 of neighboring luma reference samples. Row 1418 and columns 1420, 1421, and 1422 are directly adjacent to luma block 1411, which shares a CU with the chroma block being predicted according to cross-component intra prediction. After downsampling, row 1418 of neighboring luma reference samples becomes row 1416 of downsampled neighboring luma reference samples. Furthermore, columns 1420, 1421, and 1422 of neighboring luma reference samples are downsampled to obtain single column 1417 of downsampled neighboring luma reference samples. Furthermore, the luma samples of luma block 1411 are downsampled to generate downsampled luma reference samples 1412. The downsampled luma reference sample 1412 and the downsampled neighboring luma reference samples from row 1416 and column 1417 may then be used for cross-component intra prediction according to equation (1).

[0158] During downsampling, rows and columns are stored in memory in a line buffer. By omitting row 1319 during downsampling and instead using a single row of values ​​1418, the memory usage of the line buffer is significantly reduced. However, it is recognized that the downsampled neighboring luma reference samples from row 1316 are substantially similar to the downsampled neighboring luma reference samples from row 1416. As such, omitting row 1319 during downsampling and instead using a single row 1418 results in a reduction in the memory usage of the line buffer, and thus better processing speed, increased parallelism, fewer memory requirements, and the like, without sacrificing accuracy and therefore coding efficiency. Thus, in one exemplary embodiment, a single row 1418 of neighboring luma reference samples is downsampled for use in cross-component intra prediction.

[0159] In an example, mechanism 1400 may be implemented as follows: For luma block 1411, an upper neighboring row 1418, denoted as A1, is used for downsampling to obtain a downsampled neighboring row 1416, denoted as A, where A[i] is the i-th sample in A and A1[i] is the i-th sample in A. In a specific example, a 3-tap downsampling filter can be applied to neighboring row 1418 to obtain downsampled neighboring row 1416 according to equation (6):

number

[0160] Furthermore, left neighboring columns 1420, 1421, and 1422 are denoted as L1, L2, and L3 and are used for downsampling to obtain a downsampled neighboring column 1417 denoted as L. L[i] is the i-th sample in L, L1[i] is the i-th sample in L1, L2[i] is the i-th sample in L2, and L3[i] is the i-th sample in L3. In a specific example, a 6-tap downsampling filter can be applied to neighboring columns 1420, 1421, and 1422 to obtain a downsampled neighboring column 1417 according to equation (7):

number

[0161] It should be noted that the mechanism 1400 is not limited to the downsampling filter described. For example, instead of using the 3-tap downsampling filter described in equation (6), the samples can also be fetched directly as seen in equation (8) below:

number

[0162] Mechanism 1500 of Figure 15 is similar to mechanism 1300, but uses a single row 1518 of neighboring luma reference samples and a single column 1520 of neighboring luma reference samples instead of two rows 1318 and 1319 and three columns 1320, 1321, and 1322, respectively. Row 1518 and column 1520 are directly adjacent to luma block 1511, which shares a CU with the chroma block being predicted according to cross-component intra prediction. After downsampling, row 1518 of neighboring luma reference samples becomes row 1516 of downsampled neighboring luma reference samples. Furthermore, column 1520 of neighboring luma reference samples is downsampled to obtain a single column 1517 of downsampled neighboring luma reference samples. The downsampled neighboring luma reference samples from row 1516 and column 1517 may then be used for cross-component intra prediction according to equation (1).

[0163] Mechanism 1500 omits row 1319 and columns 1321 and 1322 during downsampling, and instead uses a single row 1518 and a single column 1520 of values, significantly reducing line buffer memory usage. However, it is recognized that the downsampled neighboring luma reference samples from row 1316 and column 1317 are substantially similar to the downsampled neighboring luma reference samples from row 1516 and column 1517, respectively. As such, omitting row 1319 and columns 1321 and 1322 during downsampling and instead using a single row 1518 and column 1520 results in reduced line buffer memory usage, and thus better processing speed, increased parallelism, fewer memory requirements, and the like, without sacrificing accuracy and therefore coding efficiency. Thus, in another exemplary embodiment, a single row 1518 of neighboring luma reference samples and a single column 1520 of neighboring luma reference samples are downsampled for use in cross-component intra prediction.

[0164] In an example, mechanism 1500 may be implemented as follows: For luma block 1511, an upper neighboring row 1518, denoted as A1, is used for downsampling to obtain a downsampled neighboring row 1516, denoted as A, where A[i] is the i-th sample in A and A1[i] is the i-th sample in A. In a specific example, a 3-tap downsampling filter can be applied to neighboring row 1518 to obtain downsampled neighboring row 1516 according to equation (9):

number

[0165] Furthermore, the left adjacent column 1520 is denoted as L1 and is used for downsampling to obtain the downsampled adjacent column 1517 denoted as L, where L[i] is the i-th sample in L and L1[i] is the i-th sample in L1. In a specific example, a 2-tap downsampling filter can be applied to the adjacent column 1520 to obtain the downsampled adjacent column 1517 according to equation (10):

number

[0166] In an alternative example, mechanism 1500 may be modified to use an L2 column (e.g., column 1321) instead of an L1 column (e.g., column 1520) when downsampling. In such a case, a 2-tap downsampling filter can be applied to adjacent column L2 to obtain downsampled adjacent column 1517 according to equation (11). It should be noted that mechanism 1500 is not limited to the downsampling filters described. For example, instead of using the 2-tap and 3-tap downsampling filters described in equations (9) and (10), samples can also be fetched directly as seen in equations (11) and (12) below:

number

[0167] It should further be noted that mechanisms 1400 and 1500 are also applicable when the dimensions of rows 1418, 1416, 1518, 1516 and / or columns 1420, 1421, 1422, 1417, 1520 and / or 1517 extend beyond the corresponding luma blocks 1411 and / or 1511 (e.g., as illustrated in FIG. 12).

[0168] In the Joint Exploration Model (JEM), there are two CCLM modes: single-model CCLM mode and multi-model CCLM mode (MMLM). As the names suggest, the single-model CCLM mode uses one linear model to predict chroma samples from luma samples for the entire CU, while in MMLM, there can be two linear models. In MMLM, the neighboring luma samples and neighboring chroma samples of the current block are classified into two groups, and each group is used as a training set to derive a linear model (i.e., a specific α and a specific β are derived for a specific group). Furthermore, the samples of the current luma block are also classified based on the same rules for classifying neighboring luma samples.

[0169] 16 is a graph illustrating an example mechanism 1600 for determining linear model parameters to support CCLM intra prediction. To derive the linear model parameters α and β, the neighboring reconstructed luma samples above and to the left may be downsampled to obtain a one-to-one relationship with the neighboring reconstructed chroma samples above and to the left. In mechanism 1200, α and β used in equation (1) are determined based on the minimum and maximum values ​​of the downsampled neighboring luma reference samples. Two points (two pairs of luma and chroma values, or two sets of luma and chroma values) (A, B) are the minimum and maximum values ​​among the set of neighboring luma samples, as illustrated in FIG. 16. This is an alternative approach to determining α and β based on minimizing the regression error.

[0170] As shown in FIG. 16, the straight line is represented by the equation Y=αx+β, and the linear model parameters α and β are obtained according to the following equations (13) and (14):

number

[0171] The example mechanism 1600 uses maximum / minimum luma values ​​and corresponding chroma values ​​to derive linear model parameters. Only two points (a point is represented by a pair of luma and chroma values) are selected from adjacent luma samples and adjacent chroma samples to derive the linear model parameters. The example mechanism 1600 does not apply to some video sequences with some noise.

[0172] Multiway Linear Models In addition to both the upper (top) adjacent sample and the left adjacent sample being usable to jointly calculate the linear model parameters, they can also be alternatively used in two other CCIP (cross-component intra prediction) modes, referred to as CCIP_A and CCIP_L modes, which can also be denoted as multi-directional linear models (MDLMs) for simplicity.

[0173] 17 and 18 are schematic diagrams illustrating an example of a mechanism for performing MDLM intra prediction. MDLM intra prediction operates in a manner similar to CCLM intra prediction 900. Specifically, MDLM intra prediction uses both cross-component linear model prediction (CCIP)_A mode 1700 and CCIP_L mode 1800 when determining the linear model parameters α and β. For example, MDLM intra prediction may calculate the linear model parameters α and β using CCIP_A mode 1700 and CCIP_L mode 1800. In another example, MDLM intra prediction may use CCIP_A mode 1700 or CCIP_L mode 1800 to determine the linear model parameters α and β.

[0174] In CCIP_A mode, only the upper neighboring samples are used to calculate the linear model parameters. To obtain more reference samples, the upper neighboring samples are usually expanded to (W+H). As shown in Figure 17, W=H, where W denotes the width of each luma or chroma block and H denotes the height of each luma or chroma block.

[0175] In CCIP_L mode, only the left-neighboring samples are used to calculate the linear model parameters. To obtain more reference samples, the left-neighboring samples are usually extended to (H+W). As shown in Figure 18, W=H, where W denotes the width of each luma or chroma block and H denotes the height of each luma or chroma block.

[0176] CCIP mode (i.e., CCLM or LM mode) and MDLM (CCIP_A and CCIP_L) can be used together, or alternatively, for example, only CCIP is used in the codec, or only MDLM is used in the codec, or both CCIP and MDLM are used in the codec.

[0177] Multiple model CCLM In addition to single-model CCLM, there is another mode called multi-model CCLM mode (MMLM). As the name suggests, single-model CCLM mode uses one linear model to predict chroma samples from luma samples for the entire CU, while in MMLM, there can be two models. In MMLM, the neighboring luma samples and neighboring chroma samples of the current block are classified into two groups, and each group is used as a training set to derive a linear model (i.e., a specific α and a specific β are derived for a specific group). Furthermore, the samples of the current luma block are also classified based on the same rules for classifying neighboring luma samples.

[0178] 19 is a graph illustrating an example mechanism 1900 for determining linear model parameters to support MMLM intra prediction. MMLM intra prediction illustrated in graph 1900 is a type of cross-component intra prediction. MMLM intra prediction is similar to CCLM intra prediction. The difference is that in MMLM, neighboring reconstructed luma samples are placed into two groups by comparing the associated luma value (e.g., Rec'L) with a threshold. CCLM intra prediction is then performed on each group to determine linear model parameters α and β and complete the corresponding linear model according to Equation (1). The classification of neighboring reconstructed luma samples into two groups may be performed according to the following Equation (15):

[0179] In the example, the threshold is calculated as the average value of neighboring reconstructed luma samples. L Neighboring reconstructed luma samples with [x,y]<=Threshod are classified into group 1, and Rec' L Neighboring reconstructed luma samples with [x,y]>Threshod are classified into group 2.

number

[0180] As shown by graph 1900, linear model parameters α1 and β1 can be calculated for the first group, and linear model parameters α2 and β2 can be calculated for the second group. As a specific example, such values ​​can be α1=2, β1=1, α2=½, and β2=−1, with a threshold luma value of 17. MMLM intra prediction can then select the resulting model that results in the smallest residual samples and / or the greatest coding efficiency.

[0181] As mentioned above, the example mechanisms for performing different CCLM intra predictions discussed herein use maximum / minimum luma values ​​and corresponding chroma values ​​to derive linear model parameters, and an improved mechanism for performing CCLM intra prediction that achieves robust linear model parameters is desirable.

[0182] If more than one point has a maximum value or more than one point has a minimum value, pairs of points will be selected based on the chroma values ​​of the corresponding points.

[0183] If more than one point has a maximum value or more than one point has a minimum value, the average chroma value of the luma samples with the maximum values ​​will be set as the corresponding chroma value for the maximum luma value, and the average chroma value of the luma samples with the minimum values ​​will be set as the corresponding chroma value for the minimum luma value.

[0184] Not only one pair of points (minimum and maximum) is selected: specifically, the N points with the larger luma values ​​and the M points with the smaller luma values ​​will be used to calculate the linear model parameters.

[0185] Not only a pair of points are selected: specifically, N points with luma values ​​in the range of [MaxValue-T1, MaxValue] and M points with luma values ​​in the range of [MinValue, MinValue+T2] will be selected as points for calculating linear model parameters.

[0186] Not only the above and left neighboring samples are used to obtain the maximum / minimum value, but also some extended neighboring samples, such as the bottom left neighboring sample and the top right neighboring sample.

[0187] According to the above exemplary improved mechanism, more robust linear model parameters can be achieved along with improving the coding efficiency of CCLM intra prediction.

[0188] In this disclosure, an improved mechanism for obtaining the maximum / minimum luma values ​​and corresponding chroma values ​​from a pair of luma and chroma samples is described in detail below.

[0189] It should be noted that the improved mechanism can also be used in MDLM and MMLM.

[0190] In this disclosure, an improved mechanism is provided to obtain the maximum and minimum luma values ​​and corresponding chroma values ​​to derive linear model parameters, which may result in more robust linear model parameters.

[0191] In the example, where the set of pairs of luma and chroma samples is {(p0,q0),(p1,q1),(p2,q2),...,(p i ,q i ),···,(p V-1 ,q V-1 )}. In this case, p i is the luma value of the i-th point, and q i is the chroma value of the i-th point, where the set of luma points is P={p0,p1,p2,...,pi ,···,p V-1}, and the set of chroma points is denoted as Q={q0,q1,···,q i ,···,q V-1}

[0192] First improved mechanism: more than one extreme point and point pairs are selected according to chroma values In the first improved mechanism, when more than one point has a maximum / minimum value, a pair of points will be selected based on the chroma values ​​of corresponding points. The pair of points with the smallest chroma value difference will be selected as the pair of points for deriving the linear model parameters.

[0193] For example, assume that the fifth, seventh, and eighth points have the largest luma values, the fourth and sixth points have the smallest luma values, and |q7-q4| is the smallest value among |q5-q4|, |q5-q6|, |q7-q4|, |q7-q6|, |q8-q4|, and |q8-q6|, then the seventh and fourth points will be selected to derive the linear model parameters.

[0194] It should be noted here that in addition to using the minimum chroma value difference, the first improved mechanism can also use the maximum chroma value difference. For example, assume that the fifth, seventh, and eighth points have the maximum luma values, the fourth and sixth points have the minimum luma values, and |q5-q6| is the maximum value among |q5-q4|, |q5-q6|, |q7-q4|, |q7-q6|, |q8-q4|, and |q8-q6|. In this case, the fifth and sixth points will be selected to derive linear model parameters.

[0195] It should be noted that the improved mechanism can also be used in MDLM and MMLM.

[0196] Second improved mechanism: use of more than one pole, average chroma value In the second improved mechanism, when more than one point has a maximum / minimum value, the average chroma value will be used. The chroma value corresponding to the maximum luma value is the average chroma value of the points with the maximum luma value. The chroma value corresponding to the minimum luma value is the average chroma value of the points with the minimum luma value.

[0197] For example, if the fifth, seventh, and eighth points have the maximum luma values ​​and the fourth and sixth points have the minimum luma values, then the chroma value corresponding to the maximum luma value is the average of q5, q7, and q8. The chroma value corresponding to the minimum luma value is the average of q4 and q6.

[0198] It should be noted that the improved mechanism can also be used in MDLM and MMLM.

[0199] Third improved mechanism: (based on the number of points, more than one point), more than one point greater / less than one point is used with the average value In the third improved mechanism, N points will be used to calculate the maximum luma value and corresponding chroma value. The selected N points have luma values ​​greater than other points. The average luma value of the selected N points will be used as the maximum luma value, and the average chroma value of the selected N points will be used as the chroma value corresponding to the maximum luma value.

[0200] The M points will be used to calculate a minimum luma value and corresponding chroma values. The selected M points have luma values ​​smaller than the other points. The average luma value of the selected M points will be used as the minimum luma value, and the average chroma value of the selected M points will be used as the chroma value corresponding to the minimum luma value.

[0201] For example, if the 5th, 7th, 8th, 9th, and 11th points have larger luma values ​​than the other points, and the 4th, 6th, 14th, and 18th points have smaller luma values, then p5, p7, p8, p9, and p 11is the maximum luma value used for the linear model parameters, and 11 The average value of p4, p6, p is the chroma value corresponding to the maximum luma value. 14 and p 18 The average of q is the minimum luma value used for the linear model parameters, and q 14 and q 18 The average value of is the chroma value that corresponds to the minimum luma value.

[0202] Note that M and N may or may not be equal, for example, M=N=2.

[0203] It should be noted that M and N can be adaptively defined based on the block size, for example, M=(W+H)>>t, N=(W+H)>>r, where t and r are the number of right shift bits, such as 2, 3, and 4.

[0204] In an alternative implementation, if (W+H)>T1, M and N are set to specific values ​​M1, N1. Otherwise, M and N are set to specific values ​​M2, N2, where M1 and N1 may or may not be equal. M2 and N2 may or may not be equal. For example, if (W+H)>16, then M=2, N=2. If (W+H)<=16, then M=1, N=1.

[0205] It should be noted that the improved mechanism is also available for MDLM and MMLM.

[0206] Fourth improved mechanism: (actively, more than one point based on luma value threshold), more than one point greater / less than one is used with average value In the fourth improved mechanism, N points will be used to calculate the maximum luma value and corresponding chroma values. The selected N points have luma values ​​in the range of [MaxlumaValue-T1, MaxlumaValue]. The average luma value of the selected N points will be used as the maximum luma value, and the average chroma value of the selected N points will be used as the chroma value corresponding to the maximum luma value. In the example, MaxlumaValue represents the maximum luma value in the set P.

[0207] In the fourth improved mechanism, M points will be used to calculate the minimum luma value and corresponding chroma value. The selected M points have luma values ​​in the range of [MinlumaValue, MinlumaValue+T2]. The average luma value of the selected M points will be used as the minimum luma value, and the average chroma value of the selected M points will be used as the chroma value corresponding to the minimum luma value. In the example, MinlumaValue represents the minimum luma value in the set P.

[0208] For example, the 5th, 7th, 8th, 9th, and 11th points are [L max -T1,L max ]. The 4th, 6th, 14th, and 18th points are points with luma values ​​in the range [L min ,L min +T2]. max denotes the maximum luma value in the set P, and L min represents the minimum luma value in the set P. Then p5, p7, p8, p9 and p 11 is the maximum luma value used for the linear model parameters, and 11 The average of p4, p6, and p7 is the maximum chroma value corresponding to the maximum luma value. 14 and p 18 The average of q is the minimum luma value used for the linear model parameters, and q 14 and q 18The average of is the minimum chroma value that corresponds to the minimum luma value.

[0209] Note that M and N may or may not be equal.

[0210] Note that T1 and T2 may or may not be equal.

[0211] It should be noted that the improved mechanism is also available for MDLM and MMLM.

[0212] Fifth improved mechanism: Use of extended adjacent samples In the existing mechanism, only the upper and left neighboring samples are used to obtain point pairs for the purpose of searching point pairs to derive linear model parameters. In the fifth improved mechanism, some extended samples can be used to increase the number of point pairs so as to improve the robustness of the linear model parameters.

[0213] For example, the upper right and lower left neighboring samples are also used to derive the linear model parameters.

[0214] For example, as shown in Figure 20, in the existing single-mode CCLM mechanism, the downsampled upper-neighboring luma sample is represented by A', the downsampled left-neighboring luma sample is represented by L', the upper-neighboring chroma sample is represented by Ac', and the left-neighboring chroma sample is represented by Lc'.

[0215] As shown in Fig. 21, in the fifth improved mechanism, the neighboring samples will be extended to the top-right and bottom-left samples, which means that the reference samples A, L and Ac, Lc may be used to obtain the maximum / minimum luma values ​​and corresponding chroma values.

[0216] Here, M>W, N>H.

[0217] It should be noted that the improved mechanism can also be used in MDLM and MMLM.

[0218] In the existing CCIP or LM mechanisms, only a pair of points is used to obtain the maximum / minimum luma value and the corresponding chroma value.

[0219] In the proposed improved mechanism, not only pairs of points are used.

[0220] If more than one point has a maximum value or more than one point has a minimum value, pairs of points will be selected based on the chroma values ​​of the corresponding points.

[0221] If more than one point has a maximum value or more than one point has a minimum value, the corresponding chroma value for the maximum luma value is the average chroma value of the luma samples with the maximum values, and the corresponding chroma value for the minimum luma value is the average chroma value of the luma samples with the minimum values.

[0222] Not only one pair of points is selected: specifically, the N points with the larger values ​​and the M points with the smaller values ​​will be used to derive the linear model parameters.

[0223] Not only pairs of points are selected: specifically, N points with values ​​in the range [MaxValue-T1, MaxValue] and M points with values ​​in the range [MinValue, MinValue+T2] will be selected as points for deriving linear model parameters.

[0224] Not only the above and left neighboring samples are used to obtain the maximum / minimum value, but also some extended neighboring samples, such as the bottom left neighboring sample and the top right neighboring sample.

[0225] All the above improved mechanisms result in obtaining more robust linear model parameters.

[0226] All the above improved mechanisms are also available in MMLM.

[0227] All the above improved mechanisms are also available in MDLM, except for improved mechanism 5.

[0228] It should be noted that the improved mechanism proposed in this disclosure is used to obtain the maximum / minimum luma values ​​and corresponding chroma values ​​for the purpose of deriving linear model parameters for chroma intra prediction. The improved mechanism is applied to the intra prediction module or intra prediction process. Therefore, it exists on both the decoder side and the encoder side. Moreover, the improved mechanism for obtaining the maximum / minimum luma values ​​and corresponding chroma values ​​may be implemented in the same way on both the encoder and the decoder.

[0229] For a chroma block, to obtain its prediction using LM mode, the corresponding downsampled luma sample is first obtained, and then the maximum / minimum luma value and the corresponding chroma value in the reconstructed neighboring samples are obtained to derive linear model parameters. Then, a prediction of the current chroma block (i.e., a prediction block) is obtained using the derived linear model parameters and the downsampled luma block.

[0230] The method for cross-component prediction of a block according to embodiment 1 of the present disclosure is related to the first improved mechanism above.

[0231] The method for cross-component prediction of a block according to the second embodiment of the present disclosure is related to the second improved mechanism above.

[0232] The method for cross-component prediction of a block according to the third embodiment of the present disclosure is related to the third improved mechanism above.

[0233] The method for cross-component prediction of a block according to the fourth embodiment of the present disclosure is related to the fourth improved mechanism above.

[0234] The method for cross-component prediction of a block according to the fifth embodiment of the present disclosure is related to the above fifth improved mechanism.

[0235] 22 is a flowchart of another example method 220 for cross-component prediction of a block (e.g., a chroma block) in accordance with some embodiments of this disclosure. Accordingly, the method may be performed by video encoder 20 and / or video decoder 30 of codec system 10 or 40. In particular, the method may be performed by intra prediction unit 46 of video encoder 20 and / or intra prediction unit 74 of video decoder 30.

[0236] In step 2201, a downsampled luma block is obtained. It can be understood that the spatial resolution of the luma block is usually greater than that of the chroma block, and the luma block (i.e., the reconstructed luma block) is downsampled to obtain the downsampled luma block. As shown in Figures 9, 12-15, luma blocks 911, 1211, 1311, 1411, and 1511 correspond to chroma block 901.

[0237] In step 2203, a maximum luma value and a minimum luma value are determined from a set of downsampled reconstructed neighboring luma samples, where the reconstructed neighboring luma samples include a plurality of reconstructed luma samples above the luma block and / or a plurality of reconstructed luma samples to the left of the luma block, and corresponding chroma values ​​are also determined.

[0238] In step 2205, the linear model parameters are calculated. For example, the linear model parameters are calculated based on the maximum luma value and the corresponding chroma value and the minimum luma value and the corresponding chroma value using equations (13) and (14).

[0239] In step 2207, a prediction block for the chroma block 901 is obtained based at least on one or more linear model parameters. A predicted chroma value for the chroma block 901 is generated based on the one or more linear model parameters and the downsampled luma blocks 1212, 1312, 1412, and 1512. The predicted chroma value for the chroma block 901 is derived using equation (1).

[0240] A method for cross-component prediction of a block according to embodiment 1 of the present disclosure (corresponding to the first improved mechanism for LM mode) is provided with reference to FIG.

[0241] The above first improved mechanism will be used to derive the maximum / minimum luma values ​​and corresponding chroma values. When more than one point has the maximum / minimum value, a pair of points will be selected based on the chroma values ​​of corresponding points. The pair of points with the minimum chroma value difference (having the maximum / minimum luma value) will be selected as the pair of points for deriving linear model parameters.

[0242] It should be noted that in addition to using the minimum chroma value difference, the first improved mechanism can also use the maximum chroma value difference.

[0243] For details, see the improved mechanism 1 presented above.

[0244] Improved mechanism 1 can also be used in MDML and MMLM. For example, for MDLM / MMLM, only the maximum / minimum luma values ​​and corresponding chroma values ​​are used to derive the linear model parameters. Improved mechanism 1 is used to derive the maximum / minimum luma values ​​and corresponding chroma values.

[0245] A method for cross-component prediction of a block according to embodiment 2 of the present disclosure (corresponding to the second improved mechanism for LM mode) is provided with reference to FIG.

[0246] The difference between the second embodiment and the first embodiment is as follows.

[0247] If more than one point has a maximum / minimum value, the average chroma value will be used. The chroma value corresponding to the maximum luma value is the average chroma value of the points with the maximum luma value. The chroma value corresponding to the minimum luma value is the average chroma value of the points with the minimum luma value.

[0248] For more details, see Improved Mechanism 2.

[0249] The improved mechanism 2 can also be used for MDLM and MMLM. For example, for MDLM / MMLM, only the maximum / minimum luma values ​​and corresponding chroma values ​​are used to derive the linear model parameters. The improved mechanism 2 is used to derive the maximum / minimum luma values ​​and corresponding chroma values.

[0250] A method for cross-component prediction of a block according to embodiment 3 (corresponding to the third improved mechanism) of the present disclosure is provided with reference to FIG.

[0251] The difference between the third embodiment and the first embodiment is as follows.

[0252] N points will be used to calculate the maximum luma value and corresponding chroma value. The selected N points have luma values ​​greater than the other points. The average luma value of the selected N points will be used as the maximum luma value, and the average chroma value of the selected N points will be used as the chroma value corresponding to the maximum luma value.

[0253] M points will be used to calculate the minimum luma value and corresponding chroma value. The selected M points have luma values ​​smaller than the other points. The average luma value of the selected M points will be used as the minimum luma value, and the average chroma value of the selected M points will be used as the chroma value corresponding to the minimum luma value.

[0254] For details, see Improved Mechanism 3 above.

[0255] Improved Mechanism 3 can also be used for MDLM and MMLM. For example, for MDLM / MMLM, only the maximum / minimum luma values ​​and corresponding chroma values ​​are used to derive the linear model parameters. Improved Mechanism 3 is used to derive the maximum / minimum luma values ​​and corresponding chroma values.

[0256] A method for cross-component prediction of a block according to embodiment 4 (corresponding to the fourth improved mechanism) of the present disclosure is provided with reference to FIG.

[0257] The difference between the fourth embodiment and the first embodiment is as follows.

[0258] N pairs of points will be used to calculate the maximum luma value and corresponding chroma value. The selected N pairs of points have luma values ​​that are in the range of [MaxlumaValue-T1,MaxlumaValue]. The average luma value of the selected N pairs of points will be used as the maximum luma value, and the average chroma value of the selected N pairs of points will be used as the chroma value that corresponds to the maximum luma value.

[0259] The M pairs of points will be used to calculate the minimum luma value and corresponding chroma value. The selected M pairs of points have luma values ​​that are in the range of [MinlumaValue, MinlumaValue+T2]. The average luma value of the selected M pairs of points will be used as the minimum luma value, and the average chroma value of the selected M pairs of points will be used as the chroma value that corresponds to the minimum luma value.

[0260] For details, see Improved Mechanism 4 above.

[0261] Improved mechanism 4 can also be used for MDLM and MMLM. For example, for MDLM / MMLM, only the maximum / minimum luma values ​​and corresponding chroma values ​​are used to derive the linear model parameters. Improved mechanism 4 is used to derive the maximum / minimum luma values ​​and corresponding chroma values.

[0262] A method for cross-component prediction of a block according to embodiment 5 (corresponding to the fifth improved mechanism) of the present disclosure is provided with reference to FIG.

[0263] The difference between the fifth embodiment and the first embodiment is as follows.

[0264] Some extended samples can be used to increase the number of point pairs to improve the robustness of the linear model parameters.

[0265] For example, the upper right and lower left neighboring samples are also used to derive the linear model parameters.

[0266] For details, see Improved Mechanism 5 above.

[0267] Improved mechanism 5 can also be used in MMLM. For example, for MMLM, only the maximum / minimum luma values ​​and corresponding chroma values ​​are used to derive the linear model parameters. Improved mechanism 5 is used to derive the maximum / minimum luma values ​​and corresponding chroma values.

[0268] 23 is a flowchart of an example method 230 for decoding video data. At step 2310, the luma blocks 911, 1211, 1311, 1411, and 1511 corresponding to the chroma block 901 are determined.

[0269] In step 2320, a set of downsampled reconstructed neighboring luma samples is determined, where the reconstructed neighboring luma samples include a plurality of reconstructed luma samples above the luma block and / or a plurality of reconstructed luma samples to the left of the luma block.

[0270] In step 2330, two pairs of luma values ​​and chroma values ​​are determined according to the N downsampled neighboring luma samples and the N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples, and / or the M downsampled neighboring luma samples and the M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples, where the minimum value of the N downsampled neighboring luma samples is not smaller than the luma value of the remaining downsampled neighboring luma samples in the downsampled sample set of the reconstructed neighboring luma samples, and the maximum value of the M downsampled neighboring luma samples is not greater than the luma value of the remaining downsampled neighboring luma samples in the downsampled sample set of the reconstructed neighboring luma samples, and M and N are positive integers greater than 1. In particular, a first pair of luma and chroma values ​​is determined according to N downsampled adjacent luma samples from the set of downsampled samples and N reconstructed adjacent chroma samples corresponding to the N downsampled adjacent luma samples, and a second pair of luma and chroma values ​​is determined according to M downsampled adjacent luma samples from the set of downsampled samples and M reconstructed adjacent chroma samples corresponding to the M downsampled adjacent luma samples.

[0271] In step 2340, one or more linear model parameters are determined based on the two pairs of luma and chroma values.

[0272] In step 2350, a prediction block for chroma block 901 is determined based at least on one or more linear model parameters, for example, a predicted chroma value for chroma block 901 is generated based on the linear model parameters and downsampled luma blocks 1212, 1312, 1412, and 1512.

[0273] In step 2360, the chroma block 901 is reconstructed based on the prediction block, for example, by adding the prediction block to the residual block to reconstruct the chroma block 901.

[0274] It should be noted that for MDLM intra prediction using CCIP_A mode 1700, the set of reconstructed neighboring luma samples includes reconstructed luma samples above the luma block but does not include reconstructed luma samples to the left of the luma block. For MDLM intra prediction using CCIP_L mode 1800, the set of reconstructed neighboring luma samples does not include reconstructed luma samples above the luma block but includes reconstructed luma samples to the left of the luma block. For CCLM intra prediction, the set of reconstructed neighboring luma samples includes reconstructed luma samples above the luma block and reconstructed luma samples to the left of the luma block.

[0275] 24 is a flowchart of an exemplary method 240 for encoding video data. At step 2410, the luma blocks 911, 1211, 1311, 1411, and 1511 corresponding to the chroma block 901 are determined.

[0276] In step 2420, a set of downsampled reconstructed neighboring luma samples is determined, where the reconstructed neighboring luma samples include a plurality of reconstructed luma samples above the luma block and / or a plurality of reconstructed luma samples to the left of the luma block.

[0277] In step 2430, two pairs of luma and chroma values ​​are determined according to the N downsampled neighboring luma samples and N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples and / or M downsampled neighboring luma samples and M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples. The minimum value of the N downsampled neighboring luma samples is not smaller than the luma value of the remaining downsampled neighboring luma samples in the downsampled sample set of the reconstructed neighboring luma samples. The maximum value of the M downsampled neighboring luma samples is not greater than the luma value of the remaining downsampled neighboring luma samples in the downsampled sample set of the reconstructed neighboring luma samples, and M and N are positive integers greater than 1. In particular, a first pair of luma and chroma values ​​is determined according to N downsampled adjacent luma samples from the set of downsampled samples and N reconstructed adjacent chroma samples corresponding to the N downsampled adjacent luma samples, and a second pair of luma and chroma values ​​is determined according to M downsampled adjacent luma samples from the set of downsampled samples and M reconstructed adjacent chroma samples corresponding to the M downsampled adjacent luma samples.

[0278] In step 2440, one or more linear model parameters are determined based on the two pairs of luma and chroma values.

[0279] In step 2450, a prediction block for chroma block 901 is determined based on one or more linear model parameters, for example, a predicted chroma value for chroma block 901 is generated based on the linear model parameters and downsampled luma blocks 1212, 1312, 1412, and 1512.

[0280] In step 2460, the chroma block 901 is coded based on the prediction block. Residual data between the chroma block and the prediction block is coded, and a bitstream including the coded residual data is generated. For example, the prediction block is subtracted from the chroma block 901 to obtain a residual block (residual data), and a bitstream including the coded residual data is generated.

[0281] It should be noted that for MDLM intra prediction using CCIP_A mode 1700, the set of reconstructed neighboring luma samples includes reconstructed luma samples above the luma block but does not include reconstructed luma samples to the left of the luma block. For MDLM intra prediction using CCIP_L mode 1800, the set of reconstructed neighboring luma samples does not include reconstructed luma samples above the luma block but includes reconstructed luma samples to the left of the luma block. For CCLM intra prediction, the set of reconstructed neighboring luma samples includes reconstructed luma samples above the luma block and reconstructed luma samples to the left of the luma block.

[0282] 25 is a flowchart of an example method 250 for decoding video data. At step 2510, the luma block 911 that corresponds to the chroma block 901 is determined.

[0283] In step 2520, a set of downsampled reconstructed neighboring luma samples is determined, where the reconstructed neighboring luma samples include a plurality of reconstructed luma samples above the luma block and / or a plurality of reconstructed luma samples to the left of the luma block.

[0284] In step 2530, if the N downsampled neighboring luma samples with the largest values ​​and / or the M downsampled neighboring luma samples with the smallest values ​​are included in the downsampled sample set of reconstructed neighboring luma samples, two pairs of luma and chroma values ​​are determined according to the N downsampled neighboring luma samples with the largest values ​​and the N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples with the largest values, and / or the M downsampled neighboring luma samples with the smallest values ​​and the M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples with the smallest values, where M, N are positive integers greater than 1. In particular, the two pairs of luma and chroma values ​​are as follows: 1. N downsampled neighboring luma samples with the largest values ​​and N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples with the largest values, as well as one downsampled neighboring luma sample with the smallest value and one reconstructed neighboring chroma sample corresponding to the downsampled neighboring luma sample with the smallest value; 2. One downsampled neighboring luma sample having the maximum value and one reconstructed neighboring chroma sample corresponding to the downsampled neighboring luma sample having the maximum value, and M downsampled neighboring luma samples having the minimum value and M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples having the minimum value; 3. N downsampled neighboring luma samples with the largest values ​​and N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples with the largest values, and M downsampled neighboring luma samples with the smallest values ​​and M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples with the smallest values, where M and N are positive integers greater than 1; is determined according to at least one of:

[0285] In step 2540, one or more linear model parameters are determined based on the two pairs of luma and chroma values.

[0286] In step 2550, a prediction block is determined based on one or more linear model parameters, for example, a predicted chroma value for chroma block 901 is generated based on the linear model parameters and downsampled luma blocks 1212, 1312, 1412, and 1512.

[0287] In step 2560, the chroma block 901 is reconstructed based on the prediction block, for example, by adding the prediction block to the residual block to reconstruct the chroma block 901.

[0288] It should be noted that for MDLM intra prediction using CCIP_A mode 1700, the set of reconstructed neighboring luma samples includes reconstructed luma samples above the luma block but does not include reconstructed luma samples to the left of the luma block. For MDLM intra prediction using CCIP_L mode 1800, the set of reconstructed neighboring luma samples does not include reconstructed luma samples above the luma block but includes reconstructed luma samples to the left of the luma block. For CCLM intra prediction, the set of reconstructed neighboring luma samples includes reconstructed luma samples above the luma block and reconstructed luma samples to the left of the luma block.

[0289] 26 is a flowchart of an example method 260 for encoding video data. At step 2610, the luma block 911 that corresponds to the chroma block 901 is determined.

[0290] In step 2620, a set of downsampled reconstructed neighboring luma samples is determined, where the reconstructed neighboring luma samples include a plurality of reconstructed luma samples above the luma block and / or a plurality of reconstructed luma samples to the left of the luma block.

[0291] In step 2630, if the N downsampled neighboring luma samples with the largest values ​​and / or the M downsampled neighboring luma samples with the smallest values ​​are included in the downsampled sample set of reconstructed neighboring luma samples, two pairs of luma and chroma values ​​are determined according to the N downsampled neighboring luma samples with the largest values ​​and the N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples with the largest values, and / or the M downsampled neighboring luma samples with the smallest values ​​and the M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples with the smallest values, where M, N are positive integers greater than 1. In particular, the two pairs of luma and chroma values ​​are as follows: 1. N downsampled neighboring luma samples with the largest values ​​and N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples with the largest values, as well as one downsampled neighboring luma sample with the smallest value and one reconstructed neighboring chroma sample corresponding to the downsampled neighboring luma sample with the smallest value; 2. One downsampled neighboring luma sample having the maximum value and one reconstructed neighboring chroma sample corresponding to the downsampled neighboring luma sample having the maximum value, and M downsampled neighboring luma samples having the minimum value and M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples having the minimum value; 3. N downsampled neighboring luma samples with the largest values ​​and N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples with the largest values, and M downsampled neighboring luma samples with the smallest values ​​and M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples with the smallest values, where M and N are positive integers greater than 1; is determined according to at least one of:

[0292] In step 2640, one or more linear model parameters are determined based on the two pairs of luma and chroma values.

[0293] In step 2650, a prediction block for chroma block 901 is determined based on one or more linear model parameters, for example, a predicted chroma value for chroma block 901 is generated based on the linear model parameters and downsampled luma blocks 1212, 1312, 1412, and 1512.

[0294] In step 2660, the chroma block 901 is coded based on the prediction block. Residual data between the chroma block and the prediction block is coded, and a bitstream including the coded residual data is generated. For example, the prediction block is subtracted from the chroma block 901 to obtain a residual block (residual data), and a bitstream including the coded residual data is generated.

[0295] It should be noted that for MDLM intra prediction using CCIP_A mode 1700, the set of reconstructed neighboring luma samples includes reconstructed luma samples above the luma block but does not include reconstructed luma samples to the left of the luma block. For MDLM intra prediction using CCIP_L mode 1800, the set of reconstructed neighboring luma samples does not include reconstructed luma samples above the luma block but includes reconstructed luma samples to the left of the luma block. For CCLM intra prediction, the set of reconstructed neighboring luma samples includes reconstructed luma samples above the luma block and reconstructed luma samples to the left of the luma block.

[0296] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this aspect, computer-readable media may generally correspond to (1) tangible computer-readable storage media that is non-transitory, or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0297] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio waves, and microwaves, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio waves, and microwaves are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory, tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where a disk typically reproduces data magnetically, while a disc reproduces data optically with a laser. Combinations of the above should also be included within the scope of computer-readable media.

[0298] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein, may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Moreover, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a composite codec. Alternatively, the techniques may be implemented entirely in one or more circuit or logic elements.

[0299] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as noted above, the various units may be combined into a codec hardware unit or may be provided by a collection of interoperating hardware units including one or more processors as described above, along with appropriate software and / or firmware.

[0300] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples should be considered illustrative rather than limiting, and the intention should not be limited to the details provided herein. For example, various elements or components may be combined or integrated into other systems, or certain features may be omitted or not implemented.

[0301] Moreover, techniques, systems, subsystems, and methods described and illustrated in various embodiments as individual or separate may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items illustrated or discussed as coupled or in direct communication with each other may also be indirectly coupled or in communication through some interface, device, or intermediate component, whether electrical, mechanical, or otherwise. Other examples of modifications, substitutions, and alterations will be ascertainable by those skilled in the art and may be made without departing from the spirit and scope of the present disclosure.

Claims

1. 1. A data structure of an encoded bitstream of video data, comprising: the bitstream includes information indicating a selected intra prediction mode for a chroma block, the selected intra prediction mode being a cross-component intra prediction_L (CCIP_L) mode; the bitstream further includes coded data of a residual block between the chroma block and a prediction block of the chroma block; The prediction block of the chroma block is obtained based on one or more parameters of a linear model, the one or more parameters of the linear model being determined based on a first pair of luma and chroma values ​​and a second pair of luma and chroma values, the first pair of luma and chroma values ​​being calculated using N downsampled neighboring luma samples in a set of downsampled samples of reconstructed neighboring luma samples and N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples, N being a positive integer greater than 1, and a minimum value of the N downsampled neighboring luma samples is a minimum value of each of a first remaining downsampled neighboring luma samples in the set that is different from the N downsampled neighboring luma samples. the second pair of luma and chroma values ​​is calculated using M downsampled neighboring luma samples in the set and M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples, where M is a positive integer greater than 1, and the maximum value of the M downsampled neighboring luma samples is less than or equal to the luma value of each of second remaining downsampled neighboring luma samples in the set that are different from the M downsampled neighboring luma samples, and in the CCIP_L mode, the reconstructed neighboring luma samples include reconstructed luma samples to the left of the luma block but not include reconstructed luma samples above a luma block corresponding to the chroma block; the bitstream further includes data describing a division of a picture into luma blocks and chroma blocks; the chroma blocks are reconstructed by adding the residual blocks to the prediction blocks, and each picture is reconstructed in accordance with the data describing a division of the picture into luma and chroma blocks; The bitstream is used in a process in which a coding device performs a number of operations, the number of operations including: performing entropy decoding on the bitstream to obtain the information indicative of a selected intra-prediction mode; determining a luma block corresponding to a chroma block; determining a downsampled set of reconstructed nearby luma samples, wherein the selected intra prediction mode is a cross-component intra prediction_L (CCIP_L) mode, and in the CCIP_L mode, the reconstructed nearby luma samples include a plurality of reconstructed luma samples to the left of the luma block but not a plurality of reconstructed luma samples above the luma block; calculating a first pair of luma and chroma values ​​using N downsampled neighboring luma samples in the set and N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples, where N is a positive integer greater than 1, and a minimum value of the N downsampled neighboring luma samples is greater than or equal to the luma value of each of a first remaining downsampled luma samples in the set that is different from the N downsampled neighboring luma samples; calculating a second pair of luma and chroma values ​​using M downsampled neighboring luma samples in the set and M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples, where M is a positive integer greater than 1, and a maximum value of the M downsampled neighboring luma samples is less than or equal to the luma value of each of a second remaining downsampled luma samples in the set that is different from the M downsampled neighboring luma samples; determining one or more parameters of a linear model based on the first pair of luma and chroma values ​​and the second pair of luma and chroma values; determining a prediction block for the chroma block based on the one or more parameters of the linear model; and reconstructing the chroma block based on the residual block and the prediction block; having Data structure of the coded bitstream.

2. the downsampled set of reconstructed neighboring luma samples consists of the N downsampled neighboring luma samples and the M downsampled neighboring luma samples; The sum of N and M is equal to the total number of downsampled neighboring luma samples included in the set. The data structure of the encoded bitstream of claim 1 .

3. the luma value of the first pair of luma and chroma values ​​is an average luma value of the N downsampled neighboring luma samples; a chroma value of the first pair of luma and chroma values ​​is an average chroma value of the N reconstructed neighboring chroma samples corresponding to the N downsampled neighboring luma samples; the luma value of the second pair of luma and chroma values ​​is an average luma value of the M downsampled neighboring luma samples; a chroma value of the second pair of luma and chroma values ​​is an average chroma value of the M reconstructed neighboring chroma samples corresponding to the M downsampled neighboring luma samples.

3. A data structure of an encoded bitstream according to claim 1 or 2.

4. M and N are equal, A data structure of an encoded bitstream according to any one of claims 1 to 3.

5. M=N=2, The data structure of the coded bitstream of claim 4.

6. The reconstructed neighboring luma samples are a bottom-left neighboring luma sample outside the luma block and a luma sample below the bottom-left neighboring luma sample outside the luma block; A data structure of an encoded bitstream according to any one of claims 1 to 5.

7. the reconstructed luma samples to the left of the luma block are reconstructed neighboring luma samples proximate each left boundary. A data structure of an encoded bitstream according to any one of claims 1 to 6.

8. the reconstructed neighboring luma samples exclude luma samples to the left of the top-left neighboring luma sample outside the luma block. A data structure of an encoded bitstream according to any one of claims 1 to 7.

9. the set of downsampled samples of the reconstructed neighboring luma samples is obtained by downsampling on the reconstructed neighboring luma samples. A data structure of an encoded bitstream according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Luma-based chroma intra-prediction for video coding

    US20170359597A1

  • Linear model chroma intra prediction for video coding

    US20180077426A1

  • Method and apparatus of advanced intra prediction for chroma components in video coding

    WO2017140211A1

  • Linear model prediction mode with sample accessing for video coding

    WO2018118940A1

  • New sample sets and new down-sampling schemes for linear component sample prediction

    WO2019162413A1