Method and apparatus for cross-component prediction for video coding

By using numerical values ​​based on external brightness and chromaticity sample values ​​in video encoding and decoding technology, and combining filter shape to predict the internal chromaticity sample values ​​of video blocks, the problem of insufficient redundancy utilization of span components in the prior art is solved, and higher encoding and decoding efficiency and bit rate reduction are achieved.

CN120077639APending Publication Date: 2025-05-30BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069317.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-09-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are difficult to effectively utilize cross-component redundancy when compressing video data, resulting in low encoding and decoding efficiency.

Method used

By determining the weighting coefficient based on the external luminance sample value and chrominance sample value, and predicting the internal chrominance sample value of the video block in combination with the filter shape, thereby improving the encoding and decoding efficiency.

Benefits of technology

It achieves higher encoding and decoding efficiency, reduces the bit rate of video data, and maintains video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077639A_ABST
    Figure CN120077639A_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for decoding video data, comprising: obtaining a video block from a bitstream; obtaining an internal brightness sample point value of the video block, and an external brightness sample point value and an external chrominance sample point value of an external region of the video block; determining a set of weighting coefficients corresponding to a filter shape by using values based on the external luma sample values and the external chroma sample values, where the filter shape and the set of weighting coefficients are configured for predicting chroma sample values based on a plurality of corresponding luma sample values, and determining a chroma sample value based on the external luma sample values. The numerical value comprises a non-downsampling value of the external brightness sample point values, or the non-downsampling value of the external brightness sample point values and a downsampling value of at least one external brightness sample point value in the external brightness sample point values; predicting an internal chroma sample value for the video block based on the internal luma sample value using the filter shape and the set of weighting coefficients; and, using the predicted internal chroma sample values to obtain a decoded video block.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 411,325, filed on September 29, 2022. The entire content of the U.S. Provisional Application is incorporated herein by reference in its entirety. Technical Field

[0002] Aspects of the present disclosure generally relate to image / video coding and compression, and more particularly, to methods and apparatuses for cross - component prediction techniques. Background Art

[0003] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Experts Group (MPEG) coding, etc. Video coding generally uses prediction methods (e.g., inter - frame prediction, intra - frame prediction, etc.), which utilize the redundancy present in video images or sequences. An important goal of video coding techniques is to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality. Summary of the Invention

[0004] A simplified summary of one or more aspects in accordance with the present disclosure is presented below in order to provide a basic understanding of these aspects. This summary of the invention is not an extensive overview of all contemplated aspects, is not intended to identify key or critical elements of all aspects, nor is it intended to delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description presented later.

[0005] According to one aspect of the present disclosure, there is provided a method for decoding video data, including: obtaining a video block from a bitstream; obtaining internal luminance sample values of the video block, external luminance sample values of an external region of the video block, and external chrominance sample values; determining a set of weighting coefficients corresponding to a filter shape by using a value based on the external luminance sample values and the external chrominance sample values, wherein the filter shape and the set of weighting coefficients are configured to predict chrominance sample values based on a plurality of corresponding luminance sample values, and the value includes a non-downsampled value of the external luminance sample values, or the non-downsampled value of the external luminance sample values together with a downsampled value of at least one of the external luminance sample values; predicting internal chrominance sample values of the video block based on the internal luminance sample values by using the filter shape and the set of weighting coefficients; and obtaining a decoded video block by using the predicted internal chrominance sample values.

[0006] According to another aspect of the present disclosure, there is provided a method for decoding video data, including: obtaining a video block from a bitstream; obtaining internal luminance sample values of the video block, external luminance sample values of a first external region of the video block, and external chrominance sample values; determining multiple sets of weighting coefficients corresponding to multiple filters based on the external luminance sample values and the external chrominance sample values of the first external region, wherein each of the multiple filters is configured to predict chrominance sample values based on a plurality of corresponding luminance sample values; predicting internal chrominance sample values of the video block by using at least one of the multiple filters based on the internal luminance sample values; and obtaining a decoded video block by using the predicted internal chrominance sample values.

[0007] According to another aspect of the present invention, there is provided a method for encoding video data, including: obtaining a video block; obtaining internal luminance sample values of the video block, external luminance sample values of an external region of the video block, and external chrominance sample values; determining a set of weighting coefficients corresponding to a filter shape by using a value based on the external luminance sample values and the external chrominance sample values, wherein the filter shape and the set of weighting coefficients are configured to predict chrominance sample values based on a plurality of corresponding luminance sample values, and the value includes a non-downsampled value of the external luminance sample values, or the non-downsampled value of the external luminance sample values together with a downsampled value of at least one of the external luminance sample values; predicting internal chrominance sample values of the video block based on the internal luminance sample values by using the filter shape and the set of weighting coefficients; and generating a bitstream including an encoded video block by using the predicted internal chrominance sample values.

[0008] According to another aspect of the present disclosure, a method for encoding video data is provided, including: obtaining a video block; obtaining internal luminance sample values of the video block, external luminance sample values and external chrominance sample values of a first external region of the video block; determining multiple sets of weighting coefficients corresponding to multiple filters based on the external luminance sample values and external chrominance sample values of the first external region, wherein each of the multiple filters is configured to predict chrominance sample values based on multiple corresponding luminance sample values; predicting internal chrominance sample values of the video block by using at least one of the multiple filters based on the internal luminance sample values; and generating a bitstream including the decoded video block by using the predicted internal chrominance sample values.

[0009] According to another aspect of the present disclosure, a computer system is provided, including: one or more processors; and one or more storage devices storing computer-executable instructions, which when executed cause the one or more processors to perform the operations of the method of the present disclosure.

[0010] According to another aspect of the present disclosure, a computer program product storing computer-executable instructions is provided, which when executed cause one or more processors to perform the operations of the method of the present disclosure.

[0011] According to another aspect of the present disclosure, a computer-readable storage medium storing instructions is provided, which when executed by a computing device having one or more processors cause the one or more processors to perform the decoding method of the present disclosure and the operation of saving a bitstream to be decoded by the decoding method of the present disclosure.

[0012] According to another aspect of the present disclosure, a computer-readable storage medium storing instructions is provided, which when executed by a computing device having one or more processors cause the one or more processors to perform the encoding method of the present disclosure and the operation of the bitstream generated by the encoding method of the present disclosure.

[0013] According to another aspect of the present disclosure, a computer-readable medium storing a bitstream is provided, wherein the bitstream will be decoded by performing the operations of the method of the present disclosure.

[0014] According to another aspect of the present disclosure, a computer-readable medium storing a bitstream is provided, wherein the bitstream is obtained by performing the operations of the method of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Aspects disclosed below will be described in conjunction with the accompanying drawings, and the accompanying drawings are provided for illustration and not limitation of the disclosed aspects.

[0016] Figure 1 A block diagram of a general block-based hybrid video coding system is shown.

[0017] Figures 2A to 2E Five splitting types are shown, including quadtree splitting, horizontal binary splitting, vertical binary splitting, horizontal ternary splitting, and vertical ternary splitting.

[0018] Figure 3 An overall block diagram of a block-based video decoder is shown.

[0019] Figure 4 An example of the positions of the left and upper samples of the current block and the samples of the current block involved in the CCLM mode is shown.

[0020] Figures 5A to 5C An example of deriving CCLM parameters is shown.

[0021] Figure 6 An example of classifying adjacent samples into two groups based on the value threshold Threshold is shown.

[0022] Figure 7 An example of classifying adjacent samples into two groups based on the inflection point is shown.

[0023] Figure 8A and Figure 8B The effect of the scaling adjustment parameter "u" is shown.

[0024] Figure 8C The reconstructed luma samples in the same position are shown.

[0025] Figure 8D The adjacent reconstructed samples are shown.

[0026] Figures 8E to 8H The steps of intra-mode derivation on the decoder side are shown.

[0027] Figure 9 An example of four reference lines adjacent to the prediction block is shown.

[0028] Figure 10A An exemplary mode for the convolutional cross-component model (CCCM) is shown.

[0029] Figure 10B An exemplary reference region composed of 6 chroma samples above and to the left of the PU is shown.

[0030] Figure 10C and Figure 10D A schematic diagram of the correlation between chroma samples and one or more luma samples is shown.

[0031] Figure 11 Shows an example of using 6 taps in a multiple linear regression (MLR) model according to one or more aspects of the present disclosure.

[0032] Figure 12 Shows exemplary different filter shapes and / or tap numbers according to one or more aspects of the present disclosure.

[0033] Figure 13 Shows an example where FLM can use only top or left luminance and / or chrominance samples (extended) for parameter derivation.

[0034] Figure 14 Shows an example where FLM can use different lines for parameter derivation.

[0035] Figures 15A to 15D Shows some examples of 1-tap / 2-tap pre-operations.

[0036] Figure 16 Shows examples of different shapes / numbers of filter taps.

[0037] Figure 17 Shows examples of different shapes / numbers of filter taps.

[0038] Figure 18A and Figure 18B Shows examples of different shapes / numbers of filter taps.

[0039] Figures 19A to 19G Shows examples of different groups of filter taps.

[0040] Figure 20A and Figure 20B Shows an example of 2-fold training for implicit filter shape derivation.

[0041] Figure 21 Shows the workflow of a method for decoding video data according to one or more aspects of the present disclosure.

[0042] Figure 22 Shows the workflow of a method for decoding video data according to one or more aspects of the present disclosure.

[0043] Figure 23 Shows the workflow of a method for encoding video data according to one or more aspects of the present disclosure.

[0044] Figure 24 Shows the workflow of a method for encoding video data according to one or more aspects of the present disclosure.

[0045] Figure 25An exemplary computing system is shown in accordance with one or more aspects of the present disclosure. Detailed Description

[0046] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. However, it will be apparent to one of ordinary skill in the art that various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0047] It should be noted that the terms "first", "second", etc. used in the specification, claims, and drawings of the present disclosure are used to distinguish objects and not to describe any particular order or sequence. It should be understood that the data used in this way may be interchanged under appropriate conditions, so that the embodiments of the present disclosure described herein may be implemented in an order other than those shown in the drawings or described in the present disclosure.

[0048] The first version of the VVC standard was finalized in July 2020, providing approximately 50% bitrate savings or equivalent perceptual quality compared to the previous-generation video coding standard HEVC. Although the VVC standard provides significant coding and decoding improvements over its predecessors, there is evidence that higher coding and decoding efficiency can be achieved using additional coding and decoding tools. Recently, under the cooperation of ITU-T VCEG and ISO / IEC MPEG, the Joint Video Exploration Team (JVET) began exploring advanced technologies that can significantly improve the coding and decoding efficiency of VVC. In April 2021, a software code library called the Enhanced Compression Model (ECM) was established for future video coding and decoding exploration work. The ECM reference software is based on the VVC Test Model (VTM) developed by JVET for VVC, where several existing modules (e.g., intra / inter prediction, transform, loop filter, etc.) were further extended and / or improved. In the future, any new coding and decoding tools beyond the VVC standard will need to be integrated into the ECM platform and tested using the JVET Common Test Conditions (CTC).

[0049] Similar to all previous video coding and decoding standards, ECM is built on a block-based hybrid video coding and decoding framework. Figure 1The block diagram of a general block-based hybrid video coding system is shown. The input video signal is processed block by block (referred to as coding units (CUs)). In ECM-1.0, a CU can be up to 128×128 pixels. However, the same as in VVC, a coding tree unit (CTU) is split into CUs based on a quadtree / binary tree / trinary tree to adapt to varying local characteristics. In the multi-type tree structure, a CTU is first split according to the quadtree structure. Then, each quadtree leaf node can be further split according to the binary tree and trinary tree structures. As Figure 2A , Figure 2B , Figure 2C , Figure 2D and Figure 2E shown, there are five partitioning types, quadtree split, vertical binary split, horizontal binary split, vertical extended quadtree split, and horizontal extended quadtree split.

[0050] In Figure 1In it, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra prediction") uses pixels of samples (referred to as reference samples) from adjacent blocks that have been encoded and decoded in the same video picture / strip to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also referred to as "inter prediction" or "motion-compensated prediction") uses the reconstructed pixels from the encoded and decoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), and the one or more motion vectors indicate the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, an additional reference picture index is sent, which is used to identify from which reference picture in the reference picture buffer the temporal prediction signal is derived. After spatial prediction and / or temporal prediction, a mode decision block in the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. Then, the predicted block is subtracted from the current video block; and the prediction residual is decorrelated and quantized using a transform. The quantized residual coefficients are dequantized and inverse-transformed to form the reconstructed residual, which is then added back to the predicted block to form the reconstructed signal of the CU. Before the reconstructed CU is placed in the reference picture buffer and used for encoding and decoding future video blocks, further loop filtering, such as a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF), can be applied to the reconstructed CU. To form the output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and the quantized residual coefficients are all sent to the entropy coding unit for further compression and packing to form the bitstream. It should be noted that the terms "block" or "video block" as used herein can be part of a frame or picture, particularly a rectangular (square or non-square) portion. For example, referring to HEVC and VVC, a block or video block can be or correspond to a coding tree unit (CTU), coding unit (CU), prediction unit (PU), or transform unit (TU), and / or can be or correspond to the corresponding block, e.g., coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB), and / or correspond to a sub-block.

[0051] Figure 3A general block diagram of a block-based video decoder is shown. The video bitstream is first entropy decoded at the entropy decoding unit. The coding mode and prediction information are sent to the spatial prediction unit (if intra coding) or the temporal prediction unit (if inter coding) to form a prediction block. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual block. Then, the prediction block and the residual block are added together. The reconstructed block may be further loop-filtered before being stored in the reference picture memory. Then, the reconstructed video in the reference picture memory is sent out to drive the display device and for predicting future video blocks.

[0052] The main focus of this disclosure is to further enhance the coding efficiency of the coding tools for cross-component prediction and cross-component linear model (CCLM) applied in ECM. Some related coding tools in ECM are briefly reviewed below. Thereafter, some deficiencies in the existing designs of CCLM are discussed. Finally, solutions are provided to improve the existing CCLM prediction design.

[0053] Cross-component linear model prediction

[0054] To reduce cross-component redundancy, the cross-component linear model (CCLM) prediction mode is used in VVC. For this CCLM prediction mode, chroma samples are predicted based on the reconstructed luma samples of the same CU by using the following linear model: pred C (i,j) = α·rec L ′(i,j)+β (1) where, pred C (i,j) represents the predicted chroma sample in the CU, and rec L '(i,j) represents the downsampled reconstructed luma sample of the same CU obtained by downsampling the reconstructed luma sample rec L (i,j). The above α and β are linear model parameters derived from up to four adjacent chroma samples and their corresponding downsampled luma samples (which can be called adjacent luma-chroma sample pairs). Assuming the size of the current chroma block is W×H, W’ and H’ are obtained as follows: - When applying the LM mode, W’ = W, H’ = H; - When applying the LM-A mode, W’ = W + H; - When applying the LM-L mode, H’ = H + W; where, in the LM mode, the upper sample and the left sample of the CU are used together to calculate the linear model coefficients; in the LM_A mode, only the upper sample of the CU is used to calculate the linear model coefficients; and in the LM_L mode, only the left sample of the CU is used to calculate the linear model coefficients.

[0055] If the positions of the adjacent samples above the chrominance block are represented as S[0, -1]... S[W’ - 1, -1], and the positions of the adjacent samples to the left of the chrominance block are represented as S[-1, 0] … S[-1, H’ - 1], then the positions of four adjacent chrominance samples are selected as follows: – When the LM mode is applied and both the adjacent samples above and to the left are available, select S[W’ / 4, -1], S[3*W’ / 4, -1], S[-1, H’ / 4], S[-1, 3*H’ / 4] as the positions of the four adjacent chrominance samples; – When the LM-A mode is applied or only the adjacent samples above are available, select S[W’ / 8, -1], S[3*W’ / 8, -1], S[5*W’ / 8, -1], S[7*W’ / 8, -1] as the positions of the four adjacent chrominance samples; – When the LM-L mode is applied or only the adjacent samples to the left are available, select S[-1, H’ / 8], S[-1, 3*H’ / 8], S[-1, 5*H’ / 8], S[-1, 7*H’ / 8] as the positions of the four adjacent chrominance samples;

[0056] Four adjacent luma samples corresponding to the selected positions are obtained through a downsampling operation, and the four obtained adjacent luma samples are compared four times to find two larger values: x 0 A and x 1 A , and two smaller values: x 0 B and x 1 B . The chrominance sample values corresponding to the two larger values and the two smaller values are respectively represented as y 0 A , y 1 A , y 0 B and y 1 B . Then X a 、X b 、Y a and Y b are derived as: X a =(x 0 A +x 1 A +1)>>1; X b =(x 0 B+x 1 B + 1) >> 1; Y a = (y 0 A + y 1 A + 1) >> 1; Y b = (y 0 B + y 1 B + 1) >> 1 (2)

[0057] Finally, the linear model parameters α and β are obtained according to the following formulae. β = Y b - α·X b (4)

[0058] Figure 4 An example showing the positions of the left and above samples of the current block involved in the CCLM mode and the samples of the current block, including the positions of the left and above samples of the N×N chrominance block in the CU and the positions of the left and above samples of the 2N×2N luma block in the CU.

[0059] The division operation for calculating the parameter α is implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented in exponential notation. For example, diff is approximated using 4 bits of significant part and an exponent. Thus, the table for 1 / diff is reduced to 16 elements for 16 valid bit values, as follows: DivTable[] = {0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0} (5)

[0060] This will have the following advantages: both reducing the computational complexity and decreasing the memory size required to store the required table.

[0061] In addition to the above template and the left template being used together to calculate the linear model coefficients, they can alternatively be used for two other LM modes, called the LM_A mode and the LM_L mode.

[0062] In the LM_T mode, only the above template is used to calculate the linear model coefficients. To obtain more samples, the above template is extended to (W + H) samples. In the LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H + W) samples.

[0063] In the LM_LT mode, the left template and the upper template are used to calculate the linear model coefficients.

[0064] To match the chrominance sample positions for 4:2:0 video sequences, two types of downsampling filters are applied to the luma samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "type 0" and "type 2" content respectively.

[0065] Note that when the upper reference line is at the CTU boundary, only one luma line (the general line buffer in intra prediction) is used to produce the downsampled luma samples.

[0066] This parameter calculation is performed as part of the decoding process, not just as an encoder search operation. Therefore, the α and β values are not transmitted to the decoder using syntax.

[0067] For chrominance intra mode coding and decoding, a total of 8 intra modes are allowed for chrominance intra mode coding and decoding. These modes include five traditional intra modes and three cross-component linear modes (CCLM, LM_A, and LM_L). The chrominance mode signaling and derivation process are shown in Table 1. The chrominance mode coding and decoding directly depend on the intra prediction mode of the corresponding luma block. Since a separate block partitioning structure for the luma component and the chrominance component is enabled in the I slice, a chrominance block can correspond to multiple luma blocks. Therefore, for the chrominance DM mode, the intra prediction mode of the corresponding luma block covering the center position of the current chrominance block is directly inherited. Table 1—Derivation of Chrominance Prediction Modes According to Luma Modes When CCLM is Enabled

[0068] As shown in Table 2, a single binarization table is used regardless of the value of sps_cclm_enabled_flag. Table 2—Unified Binarization Table for Chrominance Prediction Modes Value of intra_chroma_pred_mode Binary string 4 00 0 0100 1 0101 2 0110 3 0111 5 10 6 110 7 111

[0069] In Table 2, the first binary number indicates whether it is in the normal mode (0) or the LM mode (1). If it is in the LM mode, the next binary number indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next binary number indicates whether it is LM_L (0) or LM_A (1). For this case, when the sps_cclm_enabled_flag is 0, the first binary number of the binarized table corresponding to intra_chroma_pred_mode can be discarded before entropy encoding / decoding. Or, in other words, the first binary number is inferred to be 0 and thus is not encoded / decoded. This single binarized table is used for both cases where the sps_cclm_enabled_flag is equal to 0 and 1. The first two binary numbers in Table 22 are context-encoded using their own context models, while the remaining binary numbers are bypass-encoded.

[0070] In addition, to reduce the luminance-chrominance delay in the binary tree, when the 64×64 luminance encoding / decoding tree node is split using non-split (and ISP is not used for 64×64 CUs) or QT, the chroma CUs in the 32×32 / 32×16 chrominance encoding / decoding tree nodes are allowed to use CCLM in the following ways: — If the 32×32 chroma node is not split or is split using QT split, all chroma CUs in the 32×32 node can use CCLM. — If the 32×32 chroma node is split using horizontal BT and the 32×16 child node is not split or is split using vertical BT split, all chroma CUs in the 32×16 chroma node can use CCLM.

[0071] In all other luminance and chrominance encoding / decoding tree split conditions, CCLM is not allowed for chroma CUs.

[0072] During the ECM development process, the simplified derivation (min-max approximation) of α and β was removed. Instead, a linear least squares solution between the causal reconstruction data of the downsampled luminance samples and the causal chrominance samples was used to derive the model parameters α and β. where Rec C (i) and Rec’ L (i) indicate the reconstructed chrominance samples and the downsampled luminance samples around the target block, and I indicates the total number of samples of the adjacent data.

[0073] The LM_A and LM_L modes are also referred to as the multi-directional linear model (MDLM). Figure 5A An example showing the operation of MDLM when the block content cannot be predicted based on the L-shaped reconstruction region is presented. Figure 5BShows MDLM_L that derives CCLM parameters using only left reconstructed samples. Figure 5C Shows MDLM_T that derives CCLM parameters using only top reconstructed samples.

[0074] The integerization of the above-discussed least mean square (LMS) (refer to equations (8)-(9)) has been proposed as an improvement for CCLM. The initial integerization design of LMS CCLM was first proposed in JCTVC-C206. Then the method was improved through a series of simplifications, including JCTVC-F0233 / I0178 that reduces the precision n α from 13 to 7, JCTVC-I0151 that reduces the maximum multiplier bit width, and JCTVC-H0490 / I0166 that reduces the division LUT entries from 64 to 32, ultimately resulting in the ECM LMS version.

[0075] As discussed in equation (1), the integerization design utilizes the linear relationship to model the correlation between the luminance signal and the chrominance signal. The chrominance value is predicted from the reconstructed luminance values of the co-located blocks.

[0076] The luminance and chrominance components have different sampling rates in YUV420 sampling. The sampling rate of the chrominance component is half of the sampling rate of the luminance component and has a 0.5 pixel phase difference in the vertical direction. The reconstructed luminance needs to be downsampled in the vertical direction and subsampled in the horizontal direction to match the size of the chrominance signal. For example, the downsampling can be achieved by the following equation: Rec L ′(i,j) = (rec L (2i,2j) + rec L (2i,2j + 1)) >> 1 (10)

[0077] Floating-point operations are required in equation (8) to calculate the linear model parameter α to maintain high data precision. And when α is represented by a floating-point value, the floating-point multiplication involves equation (1). In this part, an integer implementation of the algorithm is designed. Specifically, the fractional part of the parameter α is quantized with n α bit data precision. The parameter value α is represented by the amplified and rounded integer value α' and a' = a × (1 << n α ). Then the linear model of equation (1) becomes: pred C [x,y] = (α' · Rec L '[x,y] >> n α ) + β' (11) where β′ is the rounded value of the floating-point β and α' can be calculated as follows.

[0078] It is proposed to replace the division operation in formula (12) with table lookup and multiplication. A 2 First, it is scaled down to reduce the table size. A 1 It is also scaled down to avoid product overflow. Then, in A 2 only the most significant bits defined by the value are retained, and the other bits are set to zero. The approximate value A 2 ' can be calculated as: where [... ] means a rounding operation, and can be calculated as: where bdepth(A 2 ) represents the bit depth of the value A 2 .

[0079] The same operation is performed on A 1 as follows:

[0080] Considering the quantized representations of A 1 and A 2 , formula (12) can be rewritten as follows. where is represented as a lookup table with a length of to avoid division.

[0081] In the simulation, the constant parameters are set as: ·n α is equal to 13, which is a trade - off between data precision and computational cost. ● is equal to 6, resulting in a lookup table size of 64. The table size can be further reduced to 32 by scaling up A 2 when bdepth(A 2 ) < 6 (e.g., A 2 < 32). ●n table is equal to 15, resulting in a 16 - bit data representation of the table elements. ● is set to 15 to avoid product overflow and maintain 16 - bit multiplication.

[0082] Finally, α' is clipped to [-2 -15 , 2 15-1] to maintain the 16-bit multiplication in Equation (11). Using this clipping, when n α equals 13, the actual value a is restricted to [-4, 4), which helps prevent error magnification.

[0083] Using the calculated parameter α', parameter β ′ is calculated as follows: where, since the value I is a power of 2, the division in the above formula can be simply replaced by a shift.

[0084] Similar to what was discussed above regarding Equation (1), in HM6.0, an intra prediction mode called LM is applied to use the reconstruction of the co-located luma PU to predict the chroma PU based on a linear model. The parameters of the linear model include the slope (a >> k) and the y-intercept (b), which are derived from adjacent luma and chroma pixels using the least mean square solution. The value of the predicted sample predSamples[x, y] (where x, y = 0…nS-1, and nS specifies the block size of the current chroma PU) is derived as follows: predSamples[x, y] = Clip1 C (((p Y ’[x, y] * a) >> k) + b), where x, y = 0..nS-1 (17) where, P Y ’[x, y] is the reconstructed pixel from the corresponding luma component. When the coordinates x and y are equal to or greater than 0, P Y ’ is the reconstructed pixel from the co-located luma PU. When x or y is less than 0, P Y ’ is the reconstructed adjacent pixel of the co-located luma PU.

[0085] Some intermediate variables L, C, LL, LC, k2, and k3 during the derivation are derived as: k2 = Log2((2 * nS) >> k3) (18-5) k3 = Max(0, BitDepth C + Log2(nS) - 14) (18-6)

[0086] Therefore, the variables a, b, and k can be derived as: a1 = (LC << k2) – L * C (19-1) a2 = (LL << k2) – L * L (19-2) k1 = Max(0, Log2(abs(a2)) - 5) – Max(0, Log2(abs(a1)) - 14) + 2 (19 - 3) a1s = a1 >> Max(0, Log2(abs(a1)) - 14) (19 - 4) a2s = abs(a2 >> Max(0, Log2(abs(a2)) - 5)) (19 - 5) a3 = a2s < 1? 0 : Clip3(-2 15 , 2 15 -1, a1s * lmDiv + (1 << (k1 - 1)) >> k1)(19 - 6) a = a3 >> Max(0, Log2(abs(a3)) - 6) (19 - 7) k = 13 – Max(0, Log2(abs(a)) - 6) (19 - 8) b = (L – ((a * C) >> k1) + (1 << (k2 - 1))) >> k2, (19 - 9) where lmDiv is specified in a 63 - entry lookup table (i.e., Table 3), which is generated online by the following formula: lmDiv(a2s) = ((1 << 15) + a2s / 2) / a2s. (20) Table 3 - Explanation of lmDiv

[0087] In formula (19 - 6), a1s is a 16 - bit signed integer and lmDiv is a 16 - bit unsigned integer. Therefore, a 16 - bit multiplier and 16 - bit storage are required. It is proposed to reduce the bit - depth of the multiplier to the internal bit - depth and the size of the lookup table, as detailed below.

[0088] By changing formula (19 - 4) as follows, the bit - depth of a1s is reduced to the internal bit - depth: a1s = a1 >> Max(0, Log2(abs(a1)) – (BitDepth C – 2)) (21) The value of lmDiv with the internal bit - depth is implemented and stored in the lookup table by the following formula (22): lmDiv(a2s) = ((1 << (BitDepth C -1)) + a2s / 2) / a2s. (22)

[0089] Table 4 shows an example of an internal bit depth of 10. Table 4 - Description of lmDiv with an internal bit depth equal to 10 a2s 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 lmDiv 512 256 171 128 102 85 73 64 57 51 47 43 39 37 34 32 a2s 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 lmDiv 30 28 27 26 24 23 22 21 20 20 19 18 18 17 17 16 a2s 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 lmDiv 16 15 15 14 14 13 13 13 12 12 12 12 11 11 11 11 a2s 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 lmDiv 10 10 10 10 10 9 9 9 9 9 9 9 8 8 8

[0090] Equations (19-3) and (19-8) are also modified as follows: k1 = Max(0, Log2(abs(a2)) - 5) – Max(0, Log2(abs(a1)) – (BitDepth C – 2)) and (23-1) k = BitDepth C – 1 – Max(0, Log2(abs(a)) - 6). (23-2)

[0091] It is also proposed to reduce the entries from 63 to 32 and reduce the bits per entry from 16 to 10, as shown in Table 5. By doing so, a memory saving of almost 70% can be achieved. The corresponding changes for Equations (19-6), (20), and (19-8) are as follows: a3 = a2s < 32? 0 : Clip3(-2 15 , 2 15 -1, a1s * lmDiv + (1 << (k1 - 1)) >> k1) (24-1) lmDiv(a2s) = ((1 << (BitDepth C + 4)) + a2s / 2) / a2s (24-2) k = BitDepth C + 4 – Max(0, Log2(abs(a)) - 6). (24-3) Table 5 - Description of lmDiv with an internal bit depth equal to 10 a2s 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 lmDiv 512 496 482 468 455 443 431 420 410 400 390 381 372 364 356 349 a2s 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 lmDiv 341 334 328 321 315 309 303 298 293 287 282 278 273 269 264 260

[0092] Multi-model linear model prediction

[0093] In ECM-1.0, a multi-model LM (MMLM) prediction mode is proposed. For this MMLM prediction mode, chroma samples are predicted based on the reconstructed luma samples of the same CU by using the following two linear models: where pred C (i, j) represents the predicted chroma sample in the CU, and FeC L ′(i, j) represents the downsampled reconstructed luma sample of the same CU. The threshold Threshold is calculated as the average of adjacent reconstructed luma samples.Figure 6 An example of classifying adjacent samples into two groups based on the value Threshold is given. For each group, the parameter α i and β i (where i is equal to 1 and 2 respectively) are derived based on the linear relationship between the luminance value and the chrominance value from two samples, which are the minimum luminance sample A(X A , Y A ) and the maximum luminance sample B(X B , Y B ) within the group. Here, X A , Y A are the x - coordinate (i.e., luminance value) and y - coordinate (i.e., chrominance value) of sample A, and X B , Y B are the x - coordinate and y - coordinate of sample B. The linear model parameters α and β are obtained according to the following formulae. β = y A -αx A (26)

[0094] This method is also called the minimum - maximum method. The division in the above formula can be avoided and replaced by multiplication and shift.

[0095] For square encoding / decoding blocks, the above two formulae are directly applied. For non - square encoding / decoding blocks, the adjacent samples on the longer boundary are first subsampled to have the same number of samples as the shorter boundary.

[0096] Except for the scenario where the upper template and the left template are used together to calculate the linear model coefficients, these two templates can also be alternately used in two other MMLM modes (called MMLM_A and MMLM_L modes).

[0097] In MMLM_A mode, only the pixel samples in the upper template are used to calculate the linear model coefficients. To obtain more samples, the upper template is extended to size (w + w). In MMLM_L mode, only the pixel samples in the left template are used to calculate the linear model coefficients. To obtain more samples, the left template is extended to size (H + H).

[0098] Note that when the upper reference line is at the CTU boundary, only one luminance row (which is stored in the row buffer for intra - frame prediction) is used to create the downsampled luminance samples.

[0099] For chrominance intra mode coding and decoding, a total of 11 intra modes are allowed for chrominance intra mode coding and decoding. Those modes include five traditional intra modes and six cross-component linear model modes (CCLM, LM_A, LM_L, MMLM, MMLM_A, and MMLM_L). The chrominance mode signaling and derivation process are shown in Table 6. The chrominance mode coding and decoding directly depend on the intra prediction mode of the corresponding luma block. Since a separate block splitting structure for the luma component and the chrominance component is enabled in the I slice, one chrominance block can correspond to multiple luma blocks. Therefore, for the chrominance DM mode, the intra prediction mode of the corresponding luma block covering the center position of the current chrominance block is directly inherited. Table 6 - Derivation of Chrominance Prediction Modes According to Luma Modes When MMLM is Enabled

[0100] The MMLM and LM modes can also be used together in an adaptive manner. For MMLM, the two linear models are as follows: where pred C (i, j) represents the predicted chrominance sample in the CU, and rec L ′(i, j) represents the downsampled reconstructed luma sample of the same CU. The Threshold can be simply determined based on the luma and chrominance averages and their minimum and maximum values. Figure 7 An example of classifying adjacent samples into two groups based on the inflection point T indicated by the arrow is shown. The linear model parameters α 1 and β 1 are derived from the linear relationship between the luma values and chrominance values of two samples, which are the minimum luma sample A (X A , Y A ) and Threshold (X T , Y T ). The linear model parameters α 2 and β 2 are derived from the linear relationship between the luma values and chrominance values of two samples, which are the maximum luma sample B (X B , Y B ) and Threshold (X T , Y T ). Here, X A , Y A are the x - coordinate (i.e., luma value) and y - coordinate (i.e., chrominance value) of sample A, and X B , Y B are the x - coordinate and y - coordinate of sample B. The linear model parameters α i and β i(where i equals 1 and 2 respectively) is obtained according to the following formula. β 2 = Y T - α 2 X T (28)

[0101] For square encoding / decoding blocks, the above formula is directly applied. For non-square encoding / decoding blocks, the adjacent samples on the longer boundary are first subsampled to have the same number of samples as the shorter boundary.

[0102] Except for the scenario where the upper template and the left template are used together to determine the linear model coefficients, these two templates can also be used alternatively in two other MMLM modes (referred to as MMLM_A and MMLM_L modes respectively).

[0103] In the MMLM_A mode, only the pixel samples in the upper template are used to calculate the linear model coefficients. To obtain more samples, the upper template is extended to the size (W + W). In the MMLM_L mode, only the pixel samples in the left template are used to calculate the linear model coefficients. To obtain more samples, the left template is extended to the size (H + H).

[0104] Note that when the upper reference line is at the CTU boundary, only one luminance row (which is stored in the row buffer for intra prediction) is used to generate the downsampled luminance samples.

[0105] For chrominance intra mode encoding / decoding, there is a conditional check for selecting the LM mode (CCLM, LM_A, and LM_L) or the multi-model LM mode (MMLM, MMLM_A, and MMLM_L). The conditional check is as follows: where BlkSizeThres LM represents the minimum block size of the LM mode, and BlkSizeThres MM represents the minimum block size of the MMLM mode. The symbol d represents a predetermined threshold. In one example, d can take the value of 0. In another example, d can take the value of 8.

[0106] For chroma intra mode coding and decoding, a total of 8 intra modes are allowed for chroma intra mode coding and decoding. Those modes include five traditional intra modes and three cross-component linear modes. The chroma mode signaling and derivation process are shown in Table 1. It should be noted that for a given CU, if it is coded in the linear model mode, whether it is the traditional single model LM mode or the MMLM mode is determined based on the above condition check. Different from the situation shown in Table 3, the separate MMLM mode is not signaled. The chroma mode coding and decoding directly depend on the intra prediction mode of the corresponding luma block. Since a separate block splitting structure for the luma component and the chroma component is enabled in the I slice, a chroma block can correspond to multiple luma blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luma block covering the center position of the current chroma block is directly inherited.

[0107] During the ECM development, a scaling (slope) adjustment of the CCLM was proposed as a further improvement, e.g., as described in JVET-Y0055 / Z0049.

[0108] As discussed above, the CCLM uses a model with 2 parameters to map luma values to chroma values. The scaling parameter "a" and the bias parameter "b" are defined for the following mapping: chromaVal = a * lumaVal + b (30)

[0109] A signal is proposed to send an adjustment "u" of the scaling parameter to update the model to the following form: chromaVal = a' * lumaVal + b' (31) where a' = a + u, and b' = b - u * y r .

[0110] With this selection, the mapping function is tilted or rotated around the point with the luma value y r It is recommended to use the average value of the reference luma samples used in model creation as y r so as to provide a meaningful modification to the model. Figures 8A to 8B The effect of the scaling adjustment parameter "u" is shown, where Figure 8A shows the model created without the scaling adjustment parameter "u", and Figure 8B shows the model created with the scaling adjustment parameter "u".

[0111] In one example, the scaling adjustment parameter is provided as an integer between -4 and 4 (closed interval) and is signaled in the bitstream. The unit of the scaling adjustment parameter is 1 / 8 of the chroma sample value per luma sample value (for 10-bit content).

[0112] In one example, the CCLM models ("LM_CHROMA_IDX" and "MMLM_CHROMA_IDX") that can be used to reference sample points above and to the left of the block being used are adjusted, but not for the "unilateral" mode. This selection is based on a trade-off between decoding efficiency and complexity.

[0113] When applying scaling adjustment to the multi-mode CCLM models, both models can be adjusted and thus up to two scaling updates can be signaled for a single chroma block.

[0114] To implement scaling adjustment at the encoder, the encoder can perform a SATD-based search for the optimal value of the scaling update for Cr and a similar SATD-based search for Cb. If either result is a non-zero scaling adjustment parameter, the combined scaling adjustment pair (the SATD-based update for Cr, the SATD-based update for Cb) is included in the list for the RD check for the TU.

[0115] Fusion of intra-chroma prediction modes

[0116] During the ECM development, JVET-Y0092 / Z0051 proposed the fusion of chroma intra modes.

[0117] The intra prediction modes enabled for the chroma component in ECM-4.0 are six cross-component linear model (LM) modes, including the CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L, and MMLM_T modes, the direct mode (DM), and four default chroma intra prediction modes. The four default modes are given by the list {0, 50, 18, 1}, and if the DM mode already belongs to the list, the modes in the list will be replaced by mode 66.

[0118] The decoder-side intra mode derivation (DIMD) method for luma intra prediction is included in ECM-4.0. First, the horizontal gradient and vertical gradient are calculated for each reconstructed luma sample point of the L-shaped template in the second adjacent rows and columns of the current block to construct a Histogram of Gradients (HoG). Then, the two intra prediction modes with the maximum histogram magnitude value and the second maximum histogram magnitude value are mixed with the planar mode to generate the final prediction value for the current luma block.

[0119] To improve the coding and decoding efficiency of chroma intra prediction, two methods are proposed, including decoder-side derivation of chroma intra prediction mode (DIMD chroma) and the fusion of non-LM modes and MMLM_LT mode.

[0120] In the first embodiment, the DIMD chroma mode is proposed. The proposed DIMD chroma mode uses the DIMD derivation method to derive the intra prediction mode of the current block based on the co-located reconstructed luma samples. Specifically, the horizontal gradient and the vertical gradient are calculated for each co-located reconstructed luma sample of the current chroma block to construct the HoG, as Figure 8C shown. Then, the intra prediction mode with the maximum histogram magnitude value is used to perform the intra prediction of the current chroma block.

[0121] When the intra prediction mode derived according to the DIMD chroma mode is the same as the intra prediction mode derived according to the DM mode, the intra prediction mode with the second largest histogram magnitude value is used as the DIMD chroma mode.

[0122] A CU-level flag is signaled to indicate whether the proposed DIMD chroma mode is applied, as shown in Table 7. Table 7. Binarization process for intra_chroma_pred_mode in the proposed method intra_chroma_pred_mode Binary string Intra-chroma mode 0 1100 List[0] 1 1101 List[1] 2 1110 List[2] 3 1111 List[3] 4 10 DIMD chroma 5 0 DM

[0123] In the second embodiment, the fusion of intra prediction modes of chroma is proposed, where the DM mode and four default modes can be fused with the MMLM_LT mode as follows: pred = (w0 * pred0 + w1 * pred1 + (1 << (shift - 1))) >> shift where pred0 is the prediction value obtained by applying a non-LM mode, pred1 is the prediction value obtained by applying the MMLM_LT mode, and pred is the final prediction value of the current chroma block. The two weights w0 and w1 are determined by the intra prediction modes of the neighboring chroma blocks and are set to be equal to 2. Specifically, when the neighboring blocks above and to the left are both coded with the LM mode, {w0, w1} = {1, 3}; when the neighboring blocks above and to the left are both coded with non-LM modes, {w0, w1} = {3, 1}; otherwise, {w0, w1} = {2, 2}.

[0124] For syntax design, if a non-LM mode is selected, a flag is signaled to indicate whether the fusion is applied. And the proposed fusion is only applied to I slices.

[0125] In the third embodiment, the fusion of the DIMD chrominance mode and the chrominance intra prediction mode is combined. Specifically, the DIMD chrominance mode described in the first embodiment is applied, and for I slices, the DM mode, the four default modes, and the DIMD chrominance mode can be fused with the MMLM_LT mode using the weights described in the second embodiment, while for non-I slices, only the DIMD chrominance mode can be fused with the MMLM_LT mode using equal weights.

[0126] In the fourth embodiment, the fusion of the DIMD chrominance mode with reduced processing and the chrominance intra prediction mode is combined. Specifically, the DIMD chrominance mode with reduced processing derives the intra mode based on the neighboring reconstructed Y, Cb, and Cr samples in the second adjacent rows and columns, as Figure 8D shown. Other components are the same as those in the third embodiment.

[0127] In one embodiment, when DIMD is applied, two intra modes are derived from the reconstructed neighboring samples, and the two predicted values are combined with the planar mode predicted value, where the weights are derived according to the gradient, as described in JVET-O0449. The division operation in the weight derivation is performed using the same lookup table (LUT)-based integerization scheme used by CCLM. For example, the division operation in the direction calculation Orient = G y / G x is calculated by the following LUT-based scheme: x = Floor(Log2(Gx)) normDiff = ((Gx << 4) >> x) & 15 x += (3 + (normDiff!= 0)? 1 : 0) Orient = (Gy * (DivSigTable[normDiff] | 8) + (1 << (x - 1))) >> x where DivSigTable

[16] = {0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}.

[0128] The derived intra modes are included in the main list of the intra most probable modes (MPM), so the DIMD process is performed before constructing the MPM list. The main derived intra mode of the DIMD block is stored in blocks and used for constructing the MPM list of neighboring blocks.

[0129] Figures 8E to 8H The steps of decoder-side intra mode derivation are shown, where the intra prediction direction is estimated in the absence of intra mode signaling. As Figure 8EThe first step shown includes estimating the gradient of each sample point (for light gray sample points as shown in Figure 8E ). The second step shown in Figure 8F includes mapping the gradient value to the closest predicted direction within [2, 66]. The third step shown in Figure 8G includes: selecting 2 predicted directions, where for each predicted direction, the absolute gradients Gx and Gy of adjacent pixels with that direction are summed, and the top 2 directions are selected. The fourth step shown in Figure 8H includes: enabling weighted intra prediction with the selected directions.

[0130] Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. In Figure 9 , an example of 4 reference lines is depicted, where the sample points of segments A and F are not extracted from the reconstructed adjacent sample points, but are filled with the closest sample points from segments B and E respectively. HEVC picture intra prediction uses the closest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 3) are used.

[0131] The index (mrl_idx) of the selected reference line is signaled and used to generate the intra prediction value. For reference line idx greater than 0, in the absence of the residual mode, only the additional reference line modes in the MPM list are included and only the mpm index is signaled. The reference line index is signaled before the intra prediction mode, and the planar mode is excluded from the intra prediction mode when a non-zero reference line index is signaled.

[0132] MRL is disabled for the blocks of the first line inside the CTU to prevent the use of extended reference sample points outside the current CTU line. In addition, PDPC is disabled when additional lines are used. For the MRL mode, the derivation of the DC value for the non-zero reference line index in the DC intra prediction mode is aligned with the derivation of reference line 0. MRL requires storing 3 adjacent luminance reference lines in the CTU to generate the prediction. The cross-component linear model (CCLM) tool also requires 3 adjacent luminance reference lines for its downsampling filter. The definition of MRL using the same 3 lines is aligned with CCLM to reduce the storage requirements for the decoder.

[0133] During the development of ECM, a convolutional cross-component model (CCCM) for chrominance intra modes was proposed.

[0134] It is proposed to apply the convolutional cross-component model (CCCM) to predict chrominance sample points based on the reconstructed luminance sample points in a similar spirit to what the current CCLM mode does. Similar to CCLM, when chrominance subsampling is used, the reconstructed luminance sample points are downsampled to match the lower-resolution chrominance grid.

[0135] Moreover, similar to CCLM, there is an option to use a single model or multi-model variant of CCCM. The multi-model variant uses two models, one model is derived for samples above the average luminance reference value, and the other model is derived for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for a PU with at least 128 available reference samples.

[0136] The proposed convolutional 7-tap filter includes a 5-tap + sign shape spatial component, a non-linear term, and a bias term. The input of the spatial 5-tap component of the filter includes the central (C) luminance sample co-located with the chrominance sample to be predicted and its upper / north (N), lower / south (S), left / west (W), and right / east (E) adjacent samples, as Figure 10A shown.

[0137] The non-linear term P is expressed as the square of the central luminance sample C and scaled to the sample value range of the content:

[0138] P = (C * C + midVal) >> bitDepth

[0139] That is, for 10-bit content, it is calculated as:

[0140] P = (c * c + 512) >> 10

[0141] The bias term B represents a scalar offset between the input and the output (similar to the offset term in CCLM) and is set to the middle chrominance value (set to 512 for 10-bit content).

[0142] The output of the filter is calculated as the convolution between the filter coefficients ci and the input values, and is clipped to the range of valid chrominance samples:

[0143] predChromaVal = c 0 C + c 1 N + c 2 S + c 3 E + c 4 W + c 5 P + c 6 B

[0144] The filter coefficients c are calculated by minimizing the MSE between the predicted chrominance samples and the reconstructed chrominance samples in the reference region i . Figure 10BA reference region consisting of 6 line chroma samples above and to the left of the PU is shown. The reference region extends one PU width to the right and one PU height below the PU boundary. The adjustment region is adjusted to include only available samples. An extension of the region shown in blue is needed to support the "side samples" of the + shaped spatial filter and to fill it when in an unavailable region.

[0145] MSE minimization is performed by computing the autocorrelation matrix for the luminance input and the cross - correlation vector between the luminance input and the chrominance output. The autocorrelation matrix is LDL - decomposed and back substitution is used to compute the final filter coefficients. This process roughly follows the computation of the ALF filter coefficients in ECM, however, LDL - decomposition is chosen instead of Cholesky decomposition to avoid using square - root operations. The proposed method uses only integer arithmetic.

[0146] The use of the PU - level flag for signaling the mode is coded / decoded with CABAC. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered a sub - mode of CCLM. That is, the CCCM flag is signaled only when the intra - prediction mode is LM_CHROMA_IDX (to enable single - mode CCCM) or MMLM_CHROMA_IDX (to enable multi - model CCCM).

[0147] The encoder performs two new RD checks in the chrominance prediction mode loop, one for checking the single - model CCCM mode and one for checking the multi - model CCCM mode.

[0148] In existing CCLM or MMLM designs, adjacent reconstructed luminance - chrominance sample pairs are classified into one or more sample groups based on the value Threshold which only considers the luminance DC value. That is, the luminance - chrominance sample pairs are classified by only considering the intensity of the luminance samples. However, the luminance component usually retains rich texture and the current luminance sample may be highly correlated with adjacent luminance samples. This inter - sample correlation (AC correlation) can be beneficial for the classification of luminance - chrominance sample pairs and can bring additional coding / decoding efficiency.

[0149] In addition, as Figure 10C shown, CCLM assumes that a given chrominance sample is only related to the corresponding luminance sample (L0.5, which can be at a fractional luminance sample position) and uses simple linear regression (SLR) with ordinary least squares (OLS) estimation to predict the given chrominance sample. However, as Figure 10D shown, in some video content, a chrominance sample may be related to multiple luminance samples simultaneously (AC or DC related), so a multiple linear regression (MLR) model can further improve the prediction accuracy.

[0150] Although the CCCM mode can improve intra-frame prediction efficiency, there is room for further improvement in its performance. At the same time, some parts of the existing CCCM mode also need to be simplified for efficient codec hardware implementation, or improved to achieve better codec efficiency. In addition, it is necessary to further improve the trade-off between its implementation complexity and its codec efficiency benefits.

[0151] Edge classification linear model (ELM)

[0152] In order to improve the coding efficiency of the luma component and the chroma component, a classifier that considers luma edge or AC information is introduced, which is in contrast to the above-mentioned embodiment in which only the luma DC value is considered. In addition to the existing band classification MMLM, the present disclosure also provides an exemplary classifier. The process of generating a linear prediction model for different sample groups can be similar to CCLM or MMLM (e.g., via least squares, or a simplified minimum-maximum method, etc.), but with different metrics for classification. Different classifiers can be used to classify adjacent luma samples (e.g., adjacent luma samples of adjacent luma-chroma sample pairs) and / or luma samples corresponding to the chroma samples to be predicted. Luma samples corresponding to chroma samples can be obtained by downsampling operations to match the positions of corresponding chroma samples for 4:2:0 video sequences. For example, luma samples corresponding to chroma samples can be obtained by performing downsampling operations on more than one (e.g., 4) reconstructed luma samples corresponding to chroma samples (e.g., located around chroma samples). Alternatively, for example, in the case of a 4:4:4 video sequence, the luma sample may be obtained directly from the reconstructed luma sample. Alternatively, the luma sample may be obtained from a corresponding luma sample in the reconstructed luma sample located at a corresponding co-located position of the corresponding chroma sample. For example, the luma sample to be classified may be obtained from one of the four reconstructed luma samples corresponding to the chroma sample, the reconstructed luma sample being located at an upper left position of the four reconstructed luma samples, and the upper left position may be considered as a co-located position of the chroma sample.

[0153] The first classifier can classify the luminance samples according to the edge strength of the luminance samples. For example, a direction (e.g., 0 degrees, 45 degrees, or 90 degrees, etc.) can be selected to calculate the edge strength. The direction can be formed by the current sample and adjacent samples along this direction (e.g., the adjacent sample is 45 degrees above the right of the current sample). The edge strength can be calculated by subtracting the adjacent sample from the current sample. The edge strength can be quantized into one of M segments by M - 1 thresholds, and the first classifier can classify the current sample using M classes. Alternatively or additionally, N directions can be formed by the current sample and N adjacent samples along N directions. The N edge strengths can be calculated by subtracting the N adjacent samples from the current sample respectively. Similarly, if each of the N edge strengths can be quantized into one of M segments by M - 1 thresholds, the first classifier can classify the current sample using MN classes.

[0154] The second classifier can be used for classification according to local patterns. For example, the current luminance sample Y0 can be compared with its adjacent N luminance samples Yi. If the value of Y0 is greater than the value of Yi, the score can be incremented by one, otherwise, the score can be decremented by one. The score can be quantized to form K classes. The second classifier can classify the current sample into one of the K classes. For example, the adjacent luminance samples can be obtained from four neighbors located above, to the left, to the right, and below the current luminance sample, i.e., there are no diagonal neighbors.

[0155] It can be envisioned that multiple first classifiers, second classifiers, or different instances of the first classifier or second classifier, or other classifiers described herein can be combined. For example, the first classifier can be combined with an existing classifier based on MMLM thresholds. For another example, instance A of the first classifier can be combined with another instance B of the first classifier, where instance A and instance B adopt different directions (e.g., vertical and horizontal directions respectively).

[0156] Those skilled in the art will understand that although the existing CCLM design in the VVC standard is used as the basic CCLM method in the specification, the proposed cross-component method described in the present disclosure can also be applied to other predictive coding and decoding tools with a similar design concept. For example, for chrominance from luminance (CfL) in the AV1 standard, the proposed method can also be applied by dividing the luminance / chrominance sample pairs into multiple sample groups.

[0157] Those skilled in the art will understand that Y / Cb / Cr can also be represented as Y / U / V in the field of video coding and decoding. If the video data is in RGB format, for example, by simply mapping the YUV symbols to GBR, the proposed method can also be applied.

[0158] Filter-based linear model (FLM)

[0159] The following describes a filter-based linear model (FLM) using an MLR model to account for the possibility that a chroma sample can be related to multiple luma samples simultaneously.

[0160] For a chroma sample to be predicted, the reconstructed co-located and neighboring luma samples can be used to predict the chroma sample to capture the inter-sample correlation between the co-located luma samples, neighboring luma samples, and chroma samples. The reconstructed luma samples are linearly weighted and combined with an "offset" to generate the predicted chroma sample (C: predicted chroma sample, L i : the i-th reconstructed co-located or neighboring luma sample, α i : filter coefficient, β: offset, N: filter taps), as shown in the following formula (32-1). Note that the linear weighted + offset value directly forms the predicted chroma sample (which can be adaptively low-pass, high-pass according to the video content), and then added to the residual to form the reconstructed chroma sample.

[0161] In some embodiments similar to CCCM, the offset term can also be implemented as the intermediate chroma value B (512 for 10-bit content) multiplied by another coefficient, as shown in the following formula (32-2).

[0162] For a given CU, the top and left reconstructed luma and chroma samples can be used to derive or train the FLM parameters (α i , β). Similar to CCLM, α i and β can be derived via OLS. The top and left training samples are collected, and a pseudo-inverse matrix is calculated on both the encoder and decoder sides to derive the parameters, and then these parameters are used to predict the chroma samples in a given CU. Let N denote the number of filter taps applied to the luma samples, M denote the total number of top and left reconstructed luma-chroma sample pairs for training the parameters, denote the luma sample with the i-th sample pair and the j-th filter tap, C i denote the chroma sample with the i-th sample pair, and the following formula shows the pseudo-inverse matrix A + and the derivation of the parameters. Figure 11 Shows an example where N is 6 (6 taps), M is 8, the top 2 rows and left 3 columns of luma samples, and the top 1 row and left 1 column of chroma samples are used to derive or train the parameters. b = Ax x=(A T A) -1 A T b = A + b (33)

[0163] Note that the chroma sample can be predicted only by α without the offset β, and the offset β can be a subset of the proposed method. i The proposed ELM / FLM / GLM (described below) can be directly extended to the CfL design in the AV1 standard, which explicitly transmits the model parameters (α, β). For example, in the (1-tap case), α and / or β are derived at the encoder at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / Subblock / sample level and signaled to the decoder for the CfL mode.

[0164] To further improve the codec performance, additional designs can be used in the FLM prediction. As

[0165] shown and discussed above, a 6-tap luma filter is used for the FLM prediction. However, although the multi-tap filter can well fit the training data (e.g., the reconstructed luma and chroma samples adjacent to the top and left), in some cases, the training data does not capture all the features of the test data, which may lead to overfitting and may not well predict the test data (i.e., the chroma block samples to be predicted). In addition, different filter shapes can well adapt to different video block contents, resulting in more accurate predictions. Figure 11

[0166] To solve this problem, the filter shape and the number of filter taps can be predefined or signaled or switched at the sequence parameter set (SPS), adaptive parameter set (APS), picture parameter set (PPS), picture header (PH), slice header (SH), region, CTU, CU, subblock, or sample level. A set of filter shape candidates can be predefined, and the selection of the set of filter shape candidates can be signaled or switched at the SPS, APS, PPS, PH, SH, region, CTU, CU, subblock, or sample level. Different components (e.g., U and V) can have different filter switching controls. For example, a set of filter shape candidates (e.g., indicated by indices 0 to 5) can be predefined, and as Figure 11 Figure 11As shown, the filter shape (1, 2) can represent a 2-tap luma filter, the filter shape (1, 2, 4) can represent a 3-tap luma filter, etc. The filter shape selection for the U and V components can be switched in the PH or at the CU or CTU level. Note that N taps can represent N taps with or without the offset β described herein. An example is given in Table 8 below. Table 8 - Exemplary Signaling and Switching for Different Filter Shapes

[0167] The FLM or CCCM filter shape can include a non-linear term. For example, for the CCCM filter, the chroma sample value can be predicted using the following formula: predChromaVal = c 0 C + c 1 N + c 2 S + c 3 E + c 4 W + c 5 P + c 6 B, P = (C * C + midVal) >> bitDepth where the filter corresponds to the weighted coefficients c for the center (C) luma sample value that is co-located with the chroma sample to be predicted, and its adjacent samples above / north (N), below / south (S), right / east (E), and left / west (W). 0 , c 1 ,... c 6 , the non-linear term P, and the bias term B. The value used to derive the non-linear term P can be a combination of the current luma sample and adjacent luma samples, but is not limited to C * C. For example, P can be derived as follows: P = (Q * R + midVal) >> bitDepth where Q and R represent the values used to derive the non-linear term P.

[0168] Q and R can be a linear combination of the current luma sample and adjacent luma samples in the downsampled domain (e.g., Q and R are pre-computed luma samples obtained through a weighted average operation) or without any downsampling process.

[0169] For example, each of Q and R can be selected from one of the N, S, E, W, and C luma sample values, e.g., Q * R = C * N, C * S, C * E, C * W, S * N, or N * N, etc.; or Q and R can both be equal to the average of the N, S, E, and W luma sample values, i.e., Q = R = (N + S + E + W) / 4; or, Q is equal to the C luma sample value, and R is equal to the average of the N, S, E, and W luma sample values, i.e., Q = C, and R = (N + S + E + W) / 4.

[0170] Different values (Q and R) used to derive the non - linear terms are considered different filter shapes and can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub - block / sample level. A set of filter shape candidates can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub - block / sample level.

[0171] Different chroma types and / or color formats can have different predefined filter shapes and / or taps. For example, as Figure 12 shown, the predefined filter shapes (1, 2, 4, 5) can be used for 4:2:0 type - 0, the predefined filter shapes (0, 1, 2, 4, 7) can be used for 4:2:0 type - 2, and the predefined filter shapes (1, 4) can be used for 4:2:2, and the predefined filter shapes (0, 1, 2, 3, 4, 5) can be used for 4:4:4.

[0172] In another aspect of the present disclosure, unavailable luma samples and chroma samples used to derive the MLR model can be filled from available reconstructed samples. For example, if a 6 - tap (0, 1, 2, 3, 4, 5) filter as Figure 12 shown is used, for a CU located at the left picture boundary, the left column including samples (0, 3) is unavailable (beyond the picture boundary), so samples (0, 3) are filled as duplicates from samples (1, 4) to apply the 6 - tap filter. Note that the filling process can be applied to both training data (top and left adjacent reconstructed luma and chroma samples) and test data (luma and chroma samples in the CU).

[0173] One or more shapes / numbers of filter taps can be used for FLM prediction, examples are shown in Figure 16 、 Figure 17 and Figures 18A to 18B . One or more sets of filter taps can be used for FLM prediction, examples are shown in Figures 19A to 19G .

[0174] Filter shape candidates can be implicitly derived without explicitly signaling bits. For example, the filter shape candidates can be filter shape candidates for FLM or GLM (as discussed below). In another example, the filter shape candidates can be cross - shape filters for CCCM, Figure 16 、 Figure 17 、 Figure 18A 、 Figure 18B and Figures 19A to 19GAny filter shown in, or other filters mentioned in, the present disclosure. Since longer filter taps always theoretically fit the training data (template region) better, but may overfit, the well-known "N-fold cross-validation" technique in the machine learning area can be used to train the filter coefficients. This technique divides the available training data into N groups, and uses some of the groups for training and the other groups for validation.

[0175] The following example relates to the implicit filter shape derivation for FLM prediction: Step 1: Determine M filter shape candidates for predicting the chroma sample values of the current CU; Step 2: Divide the available L-shaped template region outside the CU into N regions, labeled as R 0 , R 1 , … R N-1 , that is, divide the training data into N groups for N-fold training, where the luminance sample values and chroma sample values of the available template region are known values; Step 3: Apply each of the M filter shape candidates to a part of the available template region, that is, to one or more of all N regions R 0 , R 1 , … R N-1 ; Step 4: Derive M sets of filter coefficients corresponding to the M filter shape candidates, labeled as F 0 , F 1 , … F M-1 ; Step 5: Apply the derived sets of F 0 , F 1 , … F M-1 filter coefficients to another part of the available template region to predict the chroma sample values based on the corresponding luminance sample values, where the another part of the available template region is different from the part of the available template region mentioned in Step 3; Step 6: For each of the M filter shapes, accumulate the errors between the predicted chroma sample values and the known chroma sample values in the another part of the available template region by the sum of absolute differences (SAD), sum of squared differences (SSD), or sum of absolute transform differences (SATD), labeled as E 0 , E 1 , … E M-1 ; Step 7: Sort and select the K smallest errors, labeled as E’ 0 , E’ 1 , … E’ K-1 , which correspond to K filter shapes and K sets of filter coefficients; and Step 8: Select one filter shape candidate from the K filter shape candidates for chrominance prediction applied to the current CU. If K is greater than 1, the decoder can still receive a signal indicating the applied filter from the encoder. However, if K is 1, the signaling can be omitted, and the filter with the minimum cumulative error is determined as the applied filter.

[0176] Figure 20A and Figure 20B shows an example of 2-fold training for implicit filter shape derivation. For the current chrominance CU prediction (blue area), Figure 20A shows that in the template area, the even row area R 0 (yellow) is used to train / derive 4 sets of filter coefficients, and the odd row area R 1 (red) is used for verification / comparison and sorting cost of 4 sets of filter coefficients; Figure 20B shows that in the template area, R 0 (yellow) and R 1 (red) are interleaved. It should be understood that R 0 and R 1 can be swapped in these examples.

[0177] In one example, one of the 4 filter shape candidates is selected as the applied filter, while the L-shaped template area is divided into even and odd rows or columns. The steps include: Pre-define 4 filter shape candidates for the current CU (e.g., 4 filter shape candidates from Figure 16 ); Divide the available L-shaped template area (e.g., 6 chrominance rows and columns for CCCM, note that in the CCCM design, each chrominance sample involves 6 luminance samples for downsampling) into 2 areas labeled R 0 , R 1 , where, for example, R 0 is composed of even rows or columns, and R 1 is composed of odd rows or columns ( Figure 20A shows the following example: in the template area, the even row area R 0 is used to train / derive 4 sets of filter coefficients, and the area of odd rows R 1 is used for verification / comparison and sorting cost of 4 sets of filter coefficients); Independently apply the 4 filter shape candidates to a part of the available template area (e.g., area R 0 ); Derive 4 sets of filter coefficients for the 4 filter shapes respectively, labeled F 0 , F 1,…F 3 ; Apply the derived F 0 ,F 1 ,…F 3 to another part of the available template region (e.g., region R 1 ) to predict the corresponding chrominance sample values; For each of the 4 filters, accumulate the errors between the predicted chrominance sample values and the known chrominance sample values in another part of the available template region (e.g., region R 1 ) using SAD, SSD or SATD, and label them as E 0 ,E 1 ,…E 3 ; Select one filter shape candidate among the 4 filter shape candidates that has the minimum accumulated error in E 0 ,E 1 ,…E 3 (labeled as E’ 0 ). It corresponds to a filter shape and a set of filter coefficients. In this example, only the filter shape candidate with the minimum accumulated error will be determined as the applied filter, and then the decoder does not need to receive a signal indicating the applied filter.

[0178] In one example, the L-shaped template region can be divided into interleaved parts, where K = 2. The steps include: Determine 4 filter shape candidates for the current CU; Divide the available L-shaped template region (e.g., 6 chrominance rows or columns for CCCM) into 2 regions labeled as R 0 , R 1 , where the luma samples in R 0 and R 1 are interleaved as shown in the following table, for example: And Figure 20B shows an example of the interleaving of R 0 and R 1 in the template region; Apply the 4 filter shape candidates independently to a part of the available template region (e.g., region R 0 ); Derive 4 sets of filter coefficients for the 4 filter shapes, denoted as F 0 ,F 1 ,…F 3 ; Apply the derived F 0 ,F 1,…F 3 The sets of filter coefficients are respectively applied to another part of the available template region (e.g., region R 1 ) to predict the corresponding chrominance sample values; For each of the 4 filter shapes, the error between the predicted chrominance sample values and the known chrominance sample values in another part of the available template region (e.g., region R 1 ) is respectively accumulated by SAD, SSD or SATD, denoted as E 0 ,E 1 ,…E 3 ; For E 0 ,E 1 ,…E 3 , the 2 smallest accumulated errors are sorted and selected, denoted as E’ 0 ,E’ 1 , which corresponds to 2 filter shapes and 2 sets of filter coefficients; and Based on the signal received from the encoder, one filter shape candidate out of 2 filter shape candidates is selected to be applied to the current CU for chrominance prediction.

[0179] Note that the method of implicit filter shape derivation can also be used to determine whether to introduce a non - linear term in the CCCM filter coefficients (considering the cases with / without non - linear terms as different filter shapes).

[0180] Although examples have been shown above for the CCCM filter, it should be understood that the non - linear term P can also be included in the FLM filter (e.g., the 3*2 filter as shown in Figure 11 ) and be derived in a similar way as discussed above.

[0181] In one example, the step of dividing the available L - shaped template region can be omitted. In this example, M sets of filter coefficients can be derived based on the sample values from the available template region and then respectively applied back to the available template region to predict the corresponding chrominance sample values for accumulated errors.

[0182] As mentioned above, the MLR model (linear equation) must be derived at both the encoder and the decoder. According to one or more aspects of the present disclosure, several methods are proposed to derive the pseudo - inverse matrix A + , or directly solve the linear equation. Other known methods can be applied, such as the Newton method, Cayley - Hamilton method, and Eigendecomposition mentioned at https: / / en.wikipedia.org / wiki / Invertible_matrix.

[0183] In the present disclosure, for simplicity, A + can be represented as A - 1. A linear equation can be solved as follows: 1. Solve for A through the adjugate matrix (adjA) - 1. Closed-form analytical solution: Shown below are a general nxn form, a 2x2, and a 3x3 case. If FLM uses 3x3, then 2 scalers plus an offset need to be solved for. b = Ax, x = (A T A) -1 A T b = A + b, labeled as A -1 b By removing the (n - 1)x(n - 1) submatrix of the j-th row and i-th column 2. Gauss-Jordan elimination Gauss-Jordan elimination can be used to solve the linear equation by means of the augmented matrix [A I n and a series of elementary row operations to obtain the reduced row echelon form [I|X]. Shown below are 2x2 and 3x3 examples. 3. Cholesky decomposition To solve Ax = b, A can first be decomposed by the Cholesky-Crout algorithm, resulting in an upper triangular matrix and a lower triangular matrix, and a forward substitution followed by a back substitution can be applied in sequence to obtain the solution. Shown below is a 3x3 example.

[0184] In addition to the above examples, some conditions require special handling. For example, if some conditions result in the inability to solve the linear equation, default values can be used to fill the chrominance prediction values. The default values can be predefined or signaled or switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, when predefined as 1 << (bit Deph - 1), meanC, meanL, or meanC - meanL (average current chrominance or other chrominance, from available luminance values, or a subset of the FLM reconstructed adjacent regions).

[0185] The following example represents a situation where matrix A cannot be solved, and in which case, default prediction values can be assigned to the entire current block: 1. Solve by closed form (analytical, conjugate matrix), but A is singular (i.e., detA = 0); 2. Solve by Cholesky decomposition, but A cannot be Cholesky decomposed, g jj <REG_SQR, where REG_SQR is a small value that can be predefined or signaled or switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.

[0186] Figure 11 Illustrates a typical situation where the FLM parameters are derived using the top 2 and / or left 3 luminance lines and the top 1 and / or left 1 chrominance line. However, as mentioned above, due to the reconstruction quality of different block contents and different neighboring samples, using different regions for parameter derivation can bring coding and decoding benefits. Several ways to select the application region for parameter derivation are proposed below: 1. Similar to MDLM, FLM derivation can use only the top or left luminance and / or chrominance samples to derive parameters. Whether to use FLM, FLM_L, or FLM_T can be predefined or signaled or switched at the SPS / DPS / VPS / SEI / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Assuming the size of the current chrominance block is W×H, then W’ and H’ are obtained as follows: — When applying the FLM mode, W’ = W, H’ = H; — When applying the FLM_T mode, W’ = W + We; where We represents the extended top luminance / chrominance samples; — When applying the FLM_L mode, H’ = H + He; where He represents the extended left luminance / chrominance samples. The number of extended luminance / chrominance samples (We, He) can be predefined or signaled or switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, (We, He) = (H, W) is predefined as VVC CCLM, or (W, H) is predefined as ECM CCLM. Unavailable (We, He) luminance / chrominance samples can be repetitively filled from the nearest (horizontal, vertical) luminance / chrominance samples. Figure 13Shows an illustration of FLM_L and FLM_T (e.g., at 4 taps). When applying FLM_L or FLM_T, only the H’ or W’ luminance / chrominance samples are used for parameter derivation respectively. 2. Similar to MRL, different line indices can be predefined or signaled or switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level to indicate the selected luminance-chrominance sample pairs. This can benefit from the different reconstruction qualities of different line samples. Figure 14 Shows that similar to MRL, FLM can use different lines for parameter derivation (e.g., at 4 taps). For example, FLM can use the luminance and / or chrominance samples of light blue / yellow in index 1. 3. The extended CCLM region and uses the top N lines and / or the left M lines in full for parameter derivation. Figure 14 Shows that all dark blue, light blue, and yellow regions can be used simultaneously. Training with a larger region (data) can result in a more robust MLR model.

[0187] It should be understood that in the present disclosure, the luminance sample values in the external region of the video block to be decoded can be referred to as “external luminance sample values”, and the chrominance sample values in the external region can be referred to as “external chrominance sample values”.

[0188] The corresponding syntax can be defined as follows for FLM prediction. Among them, FLC represents fixed-length code, TU represents truncated unary code, EGk represents exponential-golomb code with order k, where k can be fixed or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level, SVLC represents signed EG0, and UVLC represents unsigned EG0. Table 9 - Example of FLM syntax

[0189] Note: The binarization of each syntax element can be changed.

[0190] Based on the existing linear model design, a new method for cross-component prediction is proposed to further improve the coding and decoding accuracy and efficiency. The main aspects of the proposed method are detailed as follows.

[0191] Although the FLM discussed above provides optimal flexibility (resulting in optimal performance), if the number of filter taps increases, many unknown parameters need to be solved. When the inverse matrix is larger than 3x3, the closed-form derivation is not appropriate (too many multipliers), and iterative methods such as Cholesky are required, which burdens the decoder processing cycles. In this section, a pre-operation before applying the linear model is proposed, including using the sample gradient to utilize the correlation between the luminance AC information and the chrominance intensity. With the help of the gradient, the number of filter taps can be effectively reduced.

[0192] Note that the methods / examples in this section can be combined / reused from any of the designs discussed above, including but not limited to classification, filter shape, matrix derivation (with special processing), application area, syntax. In addition, the methods / examples listed in this section can also be applied to any of the designs discussed above to have better performance with some complexity trade-offs.

[0193] Note that the reference sample / training template / reconstructed neighboring region as used herein generally refers to the luminance sample used to derive the MLR model parameters, which are then applied to the internal luminance samples in a CU to predict the chrominance samples in the CU.

[0194] According to the proposed method, instead of directly using the luminance sample intensity value as the input of the linear model, a pre-operation (e.g., pre-linear weighting, sign, scaling / abs, threshold, ReLU) can be applied to degrade the dimension of the unknown parameters. In one example, the pre-operation can include: calculating the sample difference based on the luminance sample value. As understood by those skilled in the art, the sample difference can be characterized as a gradient, and thus in some embodiments this new method is also referred to as the Gradient Linear Model (GLM).

[0195] Note that the following detailed description discusses the scenarios where the proposed pre-operation can be repeated for the SLR model (also known as the 1-tap case) / combined with the SLR model, and repeated for the MLR model (also known as the multi-tap case, e.g., 2-taps) / combined with the MLR model.

[0196] For example, instead of applying 2-taps to 2 luminance samples, a pre-operation can be performed on the 2 luminance samples, and then a simpler 1-tap can be applied to reduce the complexity. Figures 15A to 15D Some examples of 1-tap / 2-tap (with offset) pre-operations are shown, where the 2-tap coefficients are represented as (a, b). Note that as Figures 15A to 15DEach of the circles shown represents a schematic chrominance position in the YUV4:2:0 format. As discussed above, in the YUV 4:2:0 format, the luminance samples corresponding to a chrominance sample can be obtained by performing a downsampling operation on more than one (e.g., 4) reconstructed luminance samples corresponding to (e.g., surrounding) the chrominance sample. In other words, the chrominance position can correspond to one or more luminance samples including co-located luminance samples. Different 1-tap modes are designed for different gradient directions and different "interpolated" luminance samples (weighting different luminance positions) are used for gradient calculation. For example, a typical filter [1, 0, -1; 1, 0, -1] is shown in Figure 15A , Figure 15C and Figure 15D which represents the following operation: where rec L represents the reconstructed luminance sample value and Rec L ″(i,j) represents the pre-operation luminance sample value. Also note that the 1-tap filters shown in Figure 15A , Figure 15C and Figure 15D can be understood as an alternative to the downsampling filter with modified filter coefficients used in CCLM (refer to formulas (6)-(7)).

[0197] The pre-operation can be based on gradient, edge direction (detection), pixel intensity, pixel change, pixel variance, Roberts / Prewitt / compass / Sobel / Laplacian operator, high-pass filter (by calculating gradient or other related operators), low-pass filter (by performing weighted average operation), etc. The edge direction detectors listed in the examples can be extended to different edge directions. For example, a 1-tap (1, -1) or 2-tap (a, b) is applied along different directions to detect different edge gradients. The filter shape / coefficients can be symmetric with respect to the chrominance position, as in Figures 15A to 15D Example (420 type 0 case).

[0198] The pre-operation parameters (coefficients, signs, scaling / abs, thresholding, ReLU) can be fixed or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Note that in the examples, if multiple coefficients are applied to a sample (e.g., -1, 4), they can be combined (e.g., 3) to reduce operations.

[0199] In one example, the pre - operation can involve calculating the sample differences of the luminance sample values. Optionally, the pre - operation can include performing downsampling through a weighted average operation. In some cases, the pre - operation can be repeatedly applied. For example, a template filter can be applied to a template using a low - pass smoothing FIR filter [1, 2, 1] / 4 or [1, 2, 1; 1, 2, 1] / 8 (i.e., downsampling) to remove outliers, and then a 1 - tap GLM filter can be applied to calculate the sample differences to derive a linear model. It is conceivable that the sample differences can also be calculated and then downsampling can be enabled.

[0200] In one example, the pre - operation coefficients (finally applied (e.g., 3) or intermediate applied (e.g., - 1, 4) to each luminance sample) can be limited to powers of 2 to save multipliers.

[0201] In one aspect of the present disclosure, the proposed new method can be repeatedly used in combination with the CCLM discussed above, which utilizes a simple linear regression (SLR) model and uses a corresponding luminance sample value to predict the chrominance sample value. This is also referred to as the 1 - tap case. In this case, deriving the linear model further includes: deriving a scaling parameter α and an offset parameter β by using the pre - operation adjacent luminance sample values and adjacent chrominance sample values. Alternatively, the linear model can be rewritten as: C = α·L + β (35) where L here represents the "pre - operated" luminance sample. The parameter derivation of the 1 - tap GLM can reuse the CCLM design, but considers the directional gradient (which can have a high - pass filter). In one example, the scaling parameter α can be derived by utilizing a division lookup table detailed below to achieve simplification.

[0202] In one example, when combining the GLM with the SLR model, the scaling parameter α and the offset parameter β can be derived by utilizing the min - max method discussed above. Specifically, the scaling parameter α and the offset parameter β can be derived as follows: comparing the pre - operated adjacent luminance sample values to determine the minimum luminance sample value Y A and the maximum luminance sample value Y B ; determining the corresponding chrominance sample values X A and X B for the minimum luminance sample value Y A and the maximum luminance sample value Y B respectively; and, based on the minimum luminance sample value Y A the maximum luminance sample value Y B and the corresponding chrominance sample values X A and X B derive the scaling parameter α and the offset parameter β according to the following formula: β = Y A -αX A . (36)

[0203] In one example, when combining the GLM with the SLR model, the scaling adjustment discussed above can be reused. In this case, the encoder can determine the scaling adjustment value signaled in the bitstream (e.g., "u") and add the scaling adjustment value to the derived scaling parameter α. The decoder can determine the scaling adjustment value (e.g., "u") from the bitstream and add the scaling adjustment value to the derived scaling parameter α. The added value is ultimately used to predict the internal chrominance sample values.

[0204] In one aspect of the present disclosure, the proposed new method can be reused for / combined with the FLM, which utilizes a multi-linear regression (MLR) model and uses multiple luminance sample values to predict chrominance sample values. This is also referred to as the multi-tap case, e.g., 2-tap. In this case, the linear model can be rewritten as:

[0205] In this case, multiple scaling parameters α and offset parameter β can be derived by using pre-operated adjacent luminance sample values and adjacent chrominance sample values. In one example, the offset parameter β is optional. In one example, at least one of the multiple scaling parameters α can be derived by utilizing the sample difference. Additionally, another one of the multiple scaling parameters α can be derived by utilizing the downsampled luminance sample values. In one example, at least one of the multiple scaling parameters α can be derived by utilizing the horizontal or vertical sample difference calculated based on the downsampled adjacent luminance sample values. In other words, the linear model can incorporate multiple scaling parameters α associated with different pre-operations.

[0206] Implicit filter shape derivation

[0207] In one example, instead of explicitly signaling the selected filter shape index, the orientation-dependent filter shape used can be derived at the decoder to save bit overhead. For example, at the decoder, multiple direction gradient filters can be applied to each reconstructed luminance sample of the L-shaped template for the i-th adjacent row and column of the current block. Then, the filtered values (gradients) can be accumulated separately for each of the multiple direction gradient filters. In the example, the accumulated value is the accumulated value of the absolute values of the corresponding filtered values. After the accumulation, the direction of the direction gradient filter with the maximum accumulated value can be determined as the derived (luminance) gradient direction. For example, a gradient histogram (HoG) can be constructed to determine the maximum value. The derived direction can be further applied as the direction for predicting the chrominance samples in the current block.

[0208] The following example relates to the repeated use of a decoder - side intra - mode derivation (DIMD) method for luma intra - prediction included in ECM - 4.0: Step 1: Apply 2 - direction gradient filters (3x3 hor / ver Sobel) to each reconstructed luma sample of the L - shaped template for the second - adjacent rows and columns of the current block; Step 2: Accumulate the filtered values (gradients) of each direction gradient filter by SAD (Sum of Absolute Differences); Step 3: Construct a gradient histogram (HoG) based on the accumulated filtered values; and Step 4: The maximum value in the HoG is determined as the derived (luma) gradient direction, and based on this direction, the GLM filter can be determined.

[0209] In one example, if the shape candidates are [-1, 0, 1; -1, 0, 1] (horizontal) and [1, 2, 1; -1, -2, -1] (vertical), when the maximum value is associated with the horizontal shape, the shape [-1, 0, 1; -1, 0, 1] is used for GLM - based chroma prediction.

[0210] The gradient filter used to derive the gradient direction can be the same as or different from the GLM filter in shape. For example, both filters can be horizontal [-1, 0, 1; -1, 0, 1], or the two filters can have different shapes, and the GLM filter can be determined based on the gradient filter.

[0211] The proposed GLM can be combined with the MMLM or ELM discussed above. When combined with classification, each group can share or have its own filter shape, where the syntax indicates the shape for each group. For example, as an exemplary classifier, the horizontal gradient grad_hor can be classified into a first group, which corresponds to a first linear model, and the vertical gradient grad_ver can be classified into a second group, which corresponds to a second linear model. In one example, the horizontal luma mode can be generated only once.

[0212] Further possible classifiers are also provided as follows. Using a classifier, the adjacent and intra - luma - chroma sample pairs of the current video block can be classified into multiple groups based on one or more thresholds. Note that, as discussed above, each adjacent / intra - chroma sample and its corresponding luma sample can be referred to as a luma - chroma sample pair. One or more thresholds are associated with the intensity of the adjacent / intra - luma samples. In this case, each group in the multiple groups corresponds to a corresponding one of the multiple linear models.

[0213] When combined with an MMLM classifier, the following operations can be performed: classify adjacent reconstructed luminance-chrominance sample pairs of the current video block into two groups based on a threshold Threshold; derive different linear models for different groups, where the derivation process can be GLM-simplified, i.e., use the above pre-operations to reduce the number of taps; similarly classify luminance-chrominance sample pairs (internal luminance-chrominance sample pairs within a CU, where each internal luminance-chrominance sample pair in the internal luminance-chrominance sample pairs includes an internal chrominance sample value predicted using the derived linear model) within the CU into two groups based on the threshold Threshold; apply different linear models to the reconstructed luminance samples in different groups; and predict the chrominance samples in the CU based on the different classified linear models. Here, Threshold can be the average value of adjacent reconstructed luminance samples. Note that the number of categories (2) can be extended to multiple categories (e.g., equally divide based on the minimum / maximum of adjacent reconstructed (downsampled) luminance samples, fixed or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level) by increasing the number of Thresholds.

[0214] In one example, instead of the MMLM luminance DC intensity, the filtered values of the FLM / GLM applied to adjacent luminance samples are used for classification. For example, if a 1-tap (1, -1) GLM is applied, the average AC value (physical meaning) is used. The processing can be: classify adjacent reconstructed luminance-chrominance sample pairs into K groups based on one or more filter shapes, one or more filtered values, and K - 1 Thresholds Ti; derive different MLR models for different groups, where the derivation process can be GLM-simplified, i.e., use the above pre-operations to reduce the number of taps; similarly classify luminance-chrominance sample pairs (internal luminance-chrominance sample pairs within a CU, where each internal luminance-chrominance sample pair in the internal luminance-chrominance sample pairs includes an internal chrominance sample value predicted using the derived linear model) within the CU into K groups based on one or more filter shapes, one or more filtered values, and K - 1 Thresholds Ti; apply different linear models to the reconstructed luminance samples in different groups; predict the chrominance samples in the CU based on the different classified linear models. Here, Threshold can be predefined (e.g., 0, or can be a table) or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, Threshold can be the average AC value (filtered value) of adjacent reconstructed (which can be downsampled) luminance samples (for 2 groups), or equally divide based on the minimum / maximum AC (for K groups).

[0215] It is also proposed to combine GLM with an ELM classifier. As Figures 15A to 15D shown, a filter shape (e.g., 1-tap) can be selected to calculate the edge strength. The direction is determined as the direction of calculating the sample difference between the current sample and the samples of N adjacent samples (e.g., all 6 luminance samples). For example, Figure 15A the filter (shape [1, 0, 1; 1, 0, 1]) at the upper middle in indicates the horizontal direction because the sample difference between samples in the horizontal direction can be calculated, while the filter below it (shape [1, 2, 1; -1, -2, -1]) indicates the vertical direction because the sample difference between samples in the vertical direction can be calculated. The positive and negative coefficients in each filter can calculate the sample difference. Then, the process may include: calculating an edge strength through a filter value (e.g., an equivalent value); quantizing the edge strength into M segments through M-1 thresholds Ti; classifying the current sample using K classes (e.g., where K == M); deriving different MLR models for different groups, where the derivation process can be GLM-simplified, i.e., using the above pre-operations to reduce the number of taps; classifying the luminance-chrominance sample pairs inside the CU into K groups; applying different MLR models to the reconstructed luminance samples in different groups; and predicting the chrominance samples in the CU based on different classified MLR models. Note that the filter shape for classification can be the same as or different from the filter shape for MLR prediction. The number of both the threshold M-1 and the threshold values Ti can be fixed or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. In addition, other classifiers / combined classifiers discussed in ELM can also be used for GLM.

[0216] If the number of classified samples in a group is less than a quantity (e.g., predefined 4), the default values mentioned when discussing the matrix derivation for the MLR model can be applied to the group parameters (α i ,, β). If the corresponding adjacent reconstructed samples are not available for the selected LM mode, the default values can be applied. For example, when the MMLM_L mode is selected but the left sample is invalid.

[0217] Several methods related to the simplification of GLM are introduced below to further improve the coding and decoding efficiency.

[0218] Matrix / parameter derivation in FLM requires floating-point operations (e.g., closed-form division), which are expensive for decoder hardware, so a fixed-point design is needed. For the 1-tap GLM case, it can be regarded as a modified luminance reconstruction sample generation of CCLM (e.g., horizontal gradient direction, from CCLM[1, 2, 1; 1, 2, 1] / 8 to GLM[-1, 0, 1; -1, 0, 1]). The original CCLM process can be repeated for GLM, including fixed-point operations, MDLM downsampling, division tables, applied size limits, min-max approximation, and scaling adjustment. For all entries, the 1-tap GLM can have its own configuration or share the same design as CCLM. For example, a simplified min-max method is used to derive parameters (instead of LMS), and it is combined with scaling adjustment after the GLM model is derived. In this case, the center point (luminance value y r ) for the rotation slope becomes the average value of the reference luminance sample "gradient". Another example, when GLM is enabled for this CU, CCLM slope adjustment is inferred to be disabled, and there is no need to signal slope adjustment-related syntax.

[0219] This part, for example, adopts typical case reference samples (1 row above and 1 column to the left). Note that, as Figure 14 in, the extended reconstruction area can also use simplifications with the same spirit and can have syntax indicating specific areas (such as MDLM, MRL).

[0220] Note that the following aspects can be combined and applied jointly. For example, division processing is combined with reference sample downsampling and division tables.

[0221] When classification (MMLM / ELM) is applied, the same or different simplification operations can be applied to each group. For example, before applying a right shift, the samples of each group are filled to the target sample number respectively, and then the same derivation process and the same division table are applied.

[0222] Note that the implicit filter shape derivation method can also be used to determine whether to disable downsampling processing in the CCCM filter coefficients (regarded as different filter shapes with / without downsampling processing).

[0223] Fixed-point implementation

[0224] The 1-tap case can reuse the CCLM design. Division by n can be achieved by a right shift, and division by A 2 can be achieved by a LUT. Integerize the parameters (including n involved in the integerized design of LMS CCLM α , n tableThe intermediate parameters used for deriving the linear model (formulas (19)-(20)) can be the same as or have different values from those of the CCLM to achieve higher precision. The integerization parameters can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level, conditional on the sequence bit depth. For example, n table = bit depth + 4.

[0225] MDLM downsampling

[0226] When the GLM is combined with the MDLM, the existing total number of samples for parameter derivation may not be a power of 2 value and needs to be padded to a power of 2 to replace division with a right shift operation. For example, for an 8x4 chroma CU, the MDLM requires W + H = 12 samples, where only 8 samples of the MDLM_T are available (for reconstruction), and then the downsampled 4 samples (0, 2, 4, 6) can be equally filled. The code for implementing this operation is shown below:

[0227] Other padding methods can also be applied, such as repetitive / mirror padding relative to the last annotated sample (rightmost / bottommost).

[0228] The padding method of the GLM can be the same as or different from that of the CCLM.

[0229] Note that in the ECM version, for an 8x4 chroma CU, the MDLM_T / MDLM_L respectively require 2T / 2L = 16 / 8 samples. In this case, the same padding method can be applied to meet the target power of 2 number of samples.

[0230] Division LUT

[0231] The division LUT proposed for the CCLM / LIC (Local Illumination Compensation) in known standard development (such as AVC / HEVC / AV1 / VVC / AVS) can be used for GLM division. For example, for the case of bit depth = 10 (Table 4), the LUT in JCTVC-I0166 is reused. The division LUT can be different from that of the CCLM. For example, the CCLM uses min-max with DivTable as in formula 5, but the GLM uses a 32-entry LMS division LUT as in Table 5.

[0232] When GLM is combined with MMLM, the meanL value may not always be positive (e.g., using filtered / gradient values to classify groups), so sgn(mean L) needs to be extracted and abs(mean L) is used to find the division LUT. Note that the division LUTs for MMLM classification and parameter derivation can be different. For example, a lower-precision LUT (as the LUT in min-max) is used for average classification, and a higher-precision LUT (such as in LMS) is used for parameter derivation.

[0233] Size limit and latency constraint

[0234] Similar to the CCLM design, some size limitations can be applied to ELM / FLM / GLM. For example, the same constraints on the luminance-chrominance delay in the dual-tree can be applied.

[0235] The size limitations can be based on the CU area / width / height / depth. The thresholds can be predefined or signaled at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample levels. For example, for the chrominance CU area, the predefined threshold can be 128.

[0236] In one example, at least one pre-operation is performed in response to determining that a video block meets an enabling threshold, where the enabling threshold is associated with the area, width, height, or split depth of the video block. Specifically, the enabling threshold can define the minimum or maximum area, width, height, or split depth of the video block. As understood by those skilled in the art, the video block can include the current chrominance block and its co-located luminance block. It is also proposed to apply the above enabling threshold jointly to the current chrominance block and its co-located luminance block. For example, at least one pre-operation is performed in response to determining that the enabling threshold is met for both the current chrominance block and its co-located luminance block.

[0237] Line buffer reduction

[0238] Similar to the CCLM design, if the co-located luminance region of the current chrominance CU contains the first row within a CTU, the top template sample generation can be limited to 1 row to reduce the CTU row line buffer storage. Note that when the upper reference line is at the CTU boundary, only one luminance line (the general line buffer in intra prediction) is used for downsampled luminance samples.

[0239] For example, in Figure 13In [the context], if the co-located luma region of the current chroma CU contains the first row within a CTU, the top template can be limited to using only 1 row (but not 2 rows) for parameter derivation (other CUs can still use 2 rows). When processing the CTU line by line at the decoder hardware, this saves luma sample line buffer storage. Several methods can be used to achieve line buffer reduction. Note that the example of the limited "1" row can be extended to N rows with similar operations. Similarly, 2-tap or multi-tap can also apply such operations. When applying multi-tap, chroma samples may also need to apply operations.

[0240] For example, Figure 15A taking the 1-tap filter [1, 0, 1; 1, 0, 1] shown as an example for illustration. This filter can be reduced to [0, 0, 0; 1, 0, -1], that is, only the coefficients of the next row are used. Alternatively, the limited upper row luma samples can be filled from the next row luma samples (repetition, mirroring, 0, meanL, meanC, etc.).

[0241] Taking N = 4 as an example, that is, the video block is located at the top boundary of the current CTU, and the first 4 adjacent luma sample values and corresponding chroma sample values are used to derive the linear model. Note that the corresponding chroma sample values can refer to the corresponding top 4 rows of adjacent chroma sample values (for example, for the YUV 4:4:4 format). Alternatively, the corresponding chroma sample values can refer to the corresponding top 2 rows of adjacent chroma sample values (for example, for the YUV 4:2:0 format). In this case, the first 4 adjacent luma sample values and corresponding chroma sample values can be divided into two regions - the first region including valid sample values (for example, one nearest row of luma sample values and corresponding chroma sample values) and the second region including invalid sample values (for example, the other three rows of luma sample values and corresponding chroma sample values). Then, the coefficients of the filter corresponding to the sample positions that do not belong to the first region can be set to zero, such that only the sample values from the first region are used to calculate the sample difference. For example, as discussed above, in this case, the filter [1, 0, -1; 1, 0, -1] can be reduced to [0, 0, 0; 1, 0, -1]. Alternatively, the nearest sample value in the first region can be filled into the second region, such that the filled sample values can be used to calculate the sample difference.

[0242] Fusion of intra-chroma prediction modes

[0243] In one example, since GLM can be regarded as a special CCLM mode, the fusion design can be reused or have its own way. Multiple (two or more) weights can be applied to generate the final predicted value. For example, pred = (w0 * pred0 + w1 * pred1 + (1 << (shift - 1))) >> shift Among them, pred0 is the predicted value based on the non-LM mode, and pred1 is the predicted value based on the GLM, or pred0 is the predicted value based on one of the CCLM (including all MDLM / MMLM), and pred1 is the predicted value based on the GLM, or pred0 is the predicted value based on the GLM, and pred1 is the predicted value based on the GLM.

[0244] Different I / P / B slices can have different weight designs, w0 and w1, depending on whether the adjacent blocks are encoded using CCLM / GLM / other codec modes or block size / width / height.

[0245] For example, the design of the weights can be determined by the intra prediction mode of the adjacent chrominance blocks and shift is set to be equal to 2. Specifically, when both the upper adjacent block and the left adjacent block are encoded using the LM mode, then {w0, w1} = {1, 3}; when both the upper adjacent block and the left adjacent block are encoded using the non-LM mode, then {w0, w1} = {3, 1}; otherwise, {w0, w1} = {2, 2}. For non-I slices, both w0 and w1 can be set to be equal to 2.

[0246] For the syntax design, if the non-LM mode is selected, then a flag is signaled to indicate whether fusion is applied.

[0247] As described above, the GLM has a good gain-complexity trade-off because it can reuse the existing CCLM modules without introducing additional derivations. According to one or more aspects of the present disclosure, such a 1-tap design can be further extended or generalized.

[0248] In one aspect of the present disclosure, for the chrominance samples to be predicted, a single corresponding luminance sample L can be generated by combining the co-located luminance sample with the adjacent luminance samples. For example, the combination can be a combination of different linear filters, for example, the high-pass gradient filter (GLM) and the low-pass smoothing filter (e.g., the combination of [1, 2, 1; 1, 2, 1] / 8 FIR downsampling filter) that can typically be used in the CCLM; and / or the combination of a linear filter and a non-linear filter (e.g., having a power of n, e.g., L n , n can be a positive number, a negative number, or a +- fraction (e.g., +1 / 2 square root, or +3 cube, which can be rounded and rescaled to the bit depth dynamic range)).

[0249] In one aspect of the present disclosure, the combination can be applied repeatedly. For example, the combination of GLM and [1, 2, 1; 1, 2, 1] / 8 FIR can be applied to the reconstructed luminance samples, and then the non-linear power of 1 / 2 can be applied. For example, the non-linear filter can be implemented as a LUT (look-up table). For example, for bit depth = 10, power of n, n = 1 / 2, LUT[i] = (int)(sqrt(i)+0.5)<<5, i = 0 to 1023, where 5 is scaled to the dynamic range of bit depth = 10. When the linear filter cannot effectively handle the luminance-chrominance relationship, the non-linear filter can provide an option. Whether to use the non-linear term can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.

[0250] In one or more of the above aspects of the present disclosure, GLM can refer to a generalized linear model (which can be used to linearly or non-linearly generate a single luminance sample, and the generated single luminance sample can be fed into a CCLM linear model to derive the parameters of the CCLM linear model), and the linear / non-linear generation can be referred to as the general mode. Different gradients or general modes can be combined to form another mode. For example, the gradient mode can be combined with the CCLM downsampled value; the gradient mode can be combined with the non-linear L 2 value; the gradient mode can be combined with another gradient mode. The two gradient modes to be combined can have different directions or the same direction. For example, [1, 1, 1; -1, -1, -1] and [1, 2, 1; -1, 2, -1] (both of which have a vertical direction) can be combined, or [1, 1, 1; -1, -1, -1] and [1, 0, -1; 1, 0, -1] (both of which have vertical and horizontal directions) can be combined, as Figures 15A to 15D shown. The combination can include addition, subtraction, or linear weighting.

[0251] GLM applied to the downsampled domain

[0252] As described above, the pre-operation can be applied repeatedly, and GLM can be applied to the pre-linearly weighted / pre-operated samples. For example, as CCLM, a template filtering can be applied to the luminance samples to remove outliers using a low-pass smoothing FIR filter [1, 2, 1; 1, 2, 1] / 8 (i.e., the CCLM downsampling smoothing filter), and generate the downsampled luminance samples (one downsampled luminance sample corresponds to one chrominance sample). Thereafter, a 1-tap GLM can be applied to the smoothed downsampled luminance samples to derive the MLR model.

[0253] Some gradient filter patterns (e.g., 3x3 Sobel or Prewitt operators) can be applied to the downsampled luminance samples. The following table shows some gradient filter patterns.

[0254] The gradient filter patterns can be combined with other gradient / general filter patterns in the downsampled luminance domain. In one example, the combined filter pattern can be applied to the downsampled luminance samples. For example, the combined filter pattern can be derived by performing an addition or subtraction operation on the corresponding coefficients of the gradient filter pattern and a DC / low-pass based filter pattern (e.g., filter pattern [0, 0, 0; 0, 1, 0; 0, 0] or [1, 2, 1; 2, 4, 1; 1, 2, 1]). In another example, the combined filter pattern is derived by performing an addition or subtraction operation on the coefficients of the gradient filter pattern and a non-linear value (e.g., L 2 )). In another example, the combined filter pattern is derived by performing an addition or subtraction operation on the corresponding coefficients of the gradient filter pattern and another gradient filter pattern with a different or the same direction. In another example, the combined filter pattern is derived by performing a linear weighting operation on the coefficients of the gradient filter pattern.

[0255] The GLM applied to the downsampled domain can fit the CCCM framework, but may sacrifice high-frequency accuracy because of the low-pass smoothing applied before the GLM is applied.

[0256] GLM used as the input to CCCM

[0257] As shown above, CCCM applies luminance downsampling before convolution just like CCLM. When chroma subsampling is used, the reconstructed luminance samples are downsampled to match the lower-resolution chroma grid. Since the 1-tap GLM can also be regarded as changing the CCLM downsampling filter coefficients (e.g., from [1, 2, 1; 1, 2, 1] / 8 to [1, 2, 1, -1, -2, -1], i.e., from low-pass to high-pass), the GLM can be used as the input of CCCM. Specifically, the gradient filter of the GLM replaces the luminance downsampling filter ([1, 2, 1; 1, 2, 1] / 8) with gradient-based coefficients (e.g., [1, 2, 1, -1, -2, -1, -1]). In this case, the CCCM operation becomes "linear / non-linear combination of gradients", as shown in the following formula:

[0258] predChromaVal = c 0 C + c 1 N + c 2 S + c 3E + c 4 W + c 5 P + c 6 B Wherein, C, N, S, E, W, P are gradients of current or adjacent samples (compared with the original lower sample values for CCCM). The relevant GLM methods described in this disclosure can be applied in the same way before entering the CCCM convolution, for example, classification, separate Cb / Cr control, syntax, mode combination, PU size limit, etc.

[0259] Gradient-based coefficient replacement can be applied to specific CCCM taps. Moreover, not only high-pass can be used, but also low-pass / band-pass / all-pass coefficient replacement can be used. The replacement can be combined with the FLM / CCCM shape switching discussed above (resulting in different numbers of taps). For example, Figures 15A to 15D the gradient patterns in can be used for replacement. In one example, the operations for applying GLM as the input to CCCM include: predefining one or more coefficient candidates for CCCM / FLM downsampling; determining the CCCM / FLM filter shape and the number of filter taps for the CU; applying different CCLM downsampling coefficients to different filter taps, where the coefficients can be high-pass filters (GLM), or low-pass / band-pass / all-pass filters; generating downsampled luminance samples for the CCCM input samples (using the applied coefficients); and feeding the generated downsampled luminance samples into the CCCM processing.

[0260] Some examples for changing the CCLM downsampling filter coefficients are shown below:

[0261] Example 1:

[0262] The candidate filters are [1, 2, 1; 1, 2, 1] / 8 and [1, 0, -1; 1, 0, -1]; predChromaVal = c 0 C + c 1 N + c 2 S + c 3 E + c 4 W + c 5 P + c 6 B, using the typical CCCM cross shape, 7 taps; N, S, W, E use the filter [1, 2, 1; 1, 2, 1] / 8, maintaining the original CCLM downsampling filter; and C, P use the filter [1, 0, -1; 1, 0, -1], that is, the horizontal gradient filter, and then P physically represents the gradient ^2.

[0263] Example 2:

[0264] The candidate filters are [1, 2, 1; 1, 2, 1] / 8, [1, 0, -1; 1, 0, -1], [1, 2, 1; -1, -2, -1], [2, 1, -1; 1, -1, -2], and [-1, 1, 2; -2, -1, 1]; predChromaVal = c 0 C 0 +c 1 C 1 +c 2 C 2 +c 3 C 3 +c 4 C 4 +c 5 P + c 6 B; C 0 Use the filter [1, 2, 1; 1, 2, 1] / 8 and keep the original CCLM downsampling filter; C 1 Use the filter [1, 0, -1; 1, 0, -1]; C 2 Use the filter [1, 2, 1; -1, -2, -1]; C 3 Use the filter [2, 1, -1; 1, -1, -2]; C 4 Use the filter [-1, 1, 2; -2, -1, 1]; C 5 Use the filter [1, 2, 1; 1, 2, 1] / 8 and keep the original CCLM downsampling filter; C 0 To C 5 And P have the same downsampled luma position (in the typical CCCM cross shape, == C); and C 1 To C 4 Are generated by different Sobel-based gradient filters (as Figures 15A to 15D shown).

[0265] Example 3:

[0266] The candidate filters are [1, 2, 1; 1, 2, 1] / 8, [1, 0, -1; 1, 0, -1], [1, 2, 1; -1, -2, -1], [0, 1, 1; 0, 1, 1], and [1, 1, 0; 1, 1, 0]; predChromaVal = c 0 C 0 +c 1 C 1 +c 2 C 2 +c 3 C 3 +c 4 C 4 +c 5 P + c 6 B; C 0Use the filter [1, 2, 1; 1, 2, 1] / 8 and keep the original CCLM downsampling filter; C 1 Use the filter [1, 0, -1; 1, 0, -1]; C2 uses the filter [1, 2, 1; -1, -2, -1]; C 3 Use the filter [0, 1, 1; 0, 1, 1]; C4 uses the filter [1, 1, 0; 1, 1, 0]; C 5 Use the filter [1, 2, 1; 1, 2, 1] / 8 and keep the original CCLM downsampling filter; C 0 To C 5 And P have the same downsampled luma position (in the typical CCCM cross shape, ==C); C 1 To C 2 Are generated by Sobel-based gradient filters in different directions (as Figures 15A to 15D Shown); and C 3 To C 4 Are generated by low-pass filters.

[0267] Can be predefined at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level (as in the above examples) or signaled / switched which CCCM / FLM taps to apply coefficient replacement.

[0268] For each CCCM / FLM tap, coefficient candidates for CCCM / FLM downsampling can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.

[0269] Example 4:

[0270] The candidate filters are [1, 2, 1; 1, 2, 1] / 8, [1, 0, -1; 1, 0, -1], [1, 2, 1; -1, -2, -1], [2, 1, -1; 1, -1, -2] and [-1, 1, 2; -2, -1, 1]; predChromaVal = c 0 C + c 1 N + c 2 S + c 3 E + c 4 W + c 5 P + c 6 B, using the typical CCCM cross shape, 7 taps; C uses a downsampling filter that switches between 5 candidate filters; N, S, W, E use the filter [1, 2, 1; 1, 2, 1] / 8 and keep the original CCLM downsampling filter; and, P uses a downsampling filter that switches between 5 candidate filters.

[0271] Example 5:

[0272] The candidate filters are: [1, 2, 1; 1, 2, 1] / 8, [1, 0, -1; 1, 0, -1], [1, 2, 1; -1, -2, -1], [0, 1, 1; 0, 1, 1], and [1, 1, 0; 1, 1, 0]; predChromaVal = c 0 C + c 1 W + c 2 E + c 3 P + c 4 B, i.e., horizontal minus shape, 5-tap; C uses a downsampling filter that switches between 5 candidate filters; W and E use a downsampling filter that switches between 3 candidate filters: [1, 2, 1; 1, 2, 1] / 8, [0, 1, 1; 0, 1, 1], and [1, 1, 0; 1, 1, 0]; and P uses the filter [1, 2, 1; 1, 2, 1] / 8, maintaining the original CCLM downsampling filter.

[0273] In one or more aspects of the present disclosure, one or more syntaxes may be introduced to indicate information about the GLM. Examples of GLM syntax are shown in Table 10 below. Table 10 FLC: Fixed Length Code TU: Truncated Unary Code EGk: Exponential - Golomb code with order k, where k can be fixed or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub - block / sample level. SVLC: Signed EG0 UVLC: Unsigned EG0

[0274] Note that the binarization of each syntax element can be changed.

[0275] In one aspect of the present disclosure, the GLM on / off control for the Cb / Cr components can be performed jointly or separately. For example, at the CU level, 1 flag can be used to indicate whether GLM is active for this CU. If it is active, 1 flag can be used to indicate whether both Cb / Cr are active. If not both are active, 1 flag indicates that either Cb or Cr is active. When Cb and / or Cr are active, the filter index / gradient (general) mode can be signaled separately. All flags can have their own context models or be bypass coded / decoded.

[0276] In another aspect of the present disclosure, whether to signal the GLM on / off flag can depend on the luma / chroma coding mode, and / or CU size. For example, in the ECM5 chroma intra mode syntax, when applying MMLM or MMLM_L or MMLM_T, it can be inferred that GLM is off; when the CU area < A, where A can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level; if combined with CCCM, when CCCM is on, it can be inferred that GLM is off.

[0277] Note that when GLM is combined with MMLM, different models can share the same or have their own gradient / general mode.

[0278] When GLM is combined with CCCM / FLM, if the current CU is enabled for CCCM / FLM, the CU-level GLM enable flag can be inferred to be off. hasGlmFlag &=!pu.cccmFlag;

[0279] CCCM without downsampling processing

[0280] CCCM needs to process the downsampled luma reference value before calculating the model parameters and applying the CCCM model, which burdens the decoder processing cycle. In this part, CCCM without downsampling processing is proposed, including using non-downsampled luma reference values and / or different selections of non-downsampled luma references. One or more filter shapes can be used for the purposes described below.

[0281] Note that the methods / examples in this section can be combined / reused with the above methods, including but not limited to methods related to classification, filter shape, matrix derivation (with special handling), application area, and syntax. In addition, the methods / examples listed in this section can also be applied to the above methods / examples (more taps) to have better performance under certain complexity trade - offs.

[0282] In this disclosure, reference samples / training templates / reconstructed adjacent regions generally refer to luminance samples used to derive MLR model parameters, which are then applied to internal luminance samples in a CU to predict chrominance samples in the CU.

[0283] In one example, a convolutional 7 - tap filter can include a 5 - tap plus sign - shaped spatial component, a non - linear term, and a bias term. The input to the spatial 5 - tap component of the filter includes a central (C) non - downsampled luminance sample co - located with the chrominance sample to be predicted and its non - downsampled upper or north (N), lower or south (S), left or west (W), and right or east (E) adjacent samples, as Figure 10A shown.

[0284] The non - linear term P is expressed as the square of the central luminance sample C and scaled to the sample value range of the content:

[0285] P=(C * C + midVal)>>bitDepth

[0286] That is, for 10 - bit content, it is calculated as:

[0287] P=(C * C + 512)>>10

[0288] The bias term B represents a scalar offset between the input and output (similar to the offset term in CCLM) and is set to the middle chrominance value (512 for 10 - bit content).

[0289] The output of the filter is calculated as the convolution between the filter coefficients ci and the input values and is clipped to the range of valid chrominance samples:

[0290] predChromaVal = c 0 C + c 1 N + c 2 S + c 3 E + c 4 W + c 5 P + c 6 B

[0291] In another example, a convolutional 7-tap filter may include a 6-tap rectangular spatial component and a bias term. The input to the spatial 6-tap component of the filter includes a central (b) non-downsampled luma sample that is co-located with the chroma sample to be predicted, and its non-downsampled lower-left or southwest (d), lower-right or southeast (f), lower or south (e), left or west (a), and right or east (c) neighboring samples, as shown in Shape 1 of Figure 16 as shown.

[0292] The bias term B represents a scalar offset between the input and the output (similar to the offset term in CCLM), and is set to the middle chroma value (512 for 10-bit content).

[0293] The output of the filter is calculated as the convolution between the filter coefficients ci and the input values, and is clipped to the range of valid chroma samples:

[0294] predChromaVal = c 0 a + c 1 b + c 2 c + c 3 d + c 4 e + c 5 f + c 6 B

[0295] In yet another example, a convolutional 8-tap filter may consist of a 6-tap rectangular spatial component, a non-linear term, and a bias term. The input to the spatial 6-tap component of the filter contains a central (b) non-downsampled luma sample that is co-located with the chroma sample to be predicted, and its non-downsampled lower-left or southwest (d), lower-right or southeast (f), lower or south (e), left or west (a), and right or east (c) neighboring samples, as shown in Shape 1 of Figure 16 as shown.

[0296] The non-linear term P is represented as the square of the central luma sample (b), and is scaled to the range of the sample values of the content:

[0297] P = (b * b + midVal) >> bitDepth

[0298] That is, for 10-bit content, it is calculated as:

[0299] P = (b * b + 512) >> 10

[0300] The bias term B represents a scalar offset between the input and the output (similar to the offset term in CCLM), and is set to the middle chroma value (512 for 10-bit content).

[0301] The output of the filter is calculated as the convolution between the filter coefficients c i and the input values, and is clipped to the range of valid chroma samples:

[0302] predChromaVal = c 0 a + c 1 b + c 2 c + c 3 d + c 4 e + c 5 f + c 6 P + c 7 B

[0303] It should be understood that the above examples are merely examples, and other implementations may not deviate from the present disclosure. For example, using more or fewer taps and having Figure 16 any shape shown (where the chroma samples are represented as circles).

[0304] Filter shape

[0305] One or more shapes / numbers of filter taps can be used for CCCM prediction, as Figure 16 , Figure 17 and Figures 18A to 18B shown. One or more sets of filter taps can be used for FLM prediction, examples of which are shown in Figures 19A to 19G . The selected luma reference values are non-downsampled. One or more predefined shapes / numbers of filter taps can be used for CCCM prediction based on prior decoded information at the TB / CB / slice / picture / sequence level.

[0306] Although multi-tap filters can fit well to the training data (i.e., the top / left adjacent reconstructed luma / chroma samples), in some cases, this training data does not capture the full characteristics of the test data and may lead to overfitting and may not predict the test data (i.e., the chroma block samples to be predicted) well. Moreover, different filter shapes can adapt well to different video block contents, resulting in more accurate predictions. To solve this problem, the filter shape / number of filter taps can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. A set of filter shape candidates can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different components (U / V) can have different filter switching controls. For example, as shown in the following table, a set of filter shape candidates (idx = 0 to 5) is predefined, and filter shape (1, 2) represents a 2-tap luma filter, while filter shape (1, 2, 4) represents a 3-tap luma filter, as Figure 11As shown, etc. The filter shape selection of the U / V component can be switched at the PH or CU / CTU level. Note that N taps can represent N taps with or without the offset β as described above.

[0307] Different chroma types / color formats can have different predefined filter shapes / taps. For example, as Figure 12 shown, the predefined filter shapes are used for 420 type 0: (1, 2, 4, 5), 420 type 2: (0, 1, 2, 4, 7), 422: (1, 4), 444: (0, 1, 2, 3, 4, 5).

[0308] The unavailable luma / chroma samples for deriving the MLR model can be filled from the available reconstructed samples. For example, if a 6-tap (0, 1, 2, 3, 4, 5) filter as Figure 12 is used, for a CU located at the left picture boundary, the left column including (0, 3) is unavailable (outside the picture boundary), so (0, 3) is filled as a repetition from (1, 4) to apply the 6-tap filter. Note that the filling process is applied to both the training data (top / left adjacent reconstructed luma / chroma samples) and the test data (luma / chroma samples in the CU).

[0309] According to one or more embodiments of the present disclosure, the unavailable luma / chroma samples for deriving the MLR model can be skipped and not used. Then, the unavailable luma / chroma samples do not require the filling process.

[0310] CCLM / MMLM with LDL decomposition

[0311] CCCM needs to process LDL decomposition to calculate the model parameters of the CCCM model, avoiding the use of square root operations and only requiring integer operations. In this section, CCLM / MMLM with LDL decomposition is proposed. As mentioned above, LDL decomposition can also be used in ELM / FLM / GLM.

[0312] Please note that the methods / examples in this section can be combined / reused with the above methods, including but not limited to the methods related to classification, filter shape, matrix derivation (with special processing), application area, and syntax. In addition, the methods / examples listed in this section can also be applied together with the above methods / examples to have better performance under certain complexity trade-offs.

[0313] In the present disclosure, reference samples / training templates / reconstructed adjacent regions generally refer to the luma samples used to derive the MLR model parameters, which are then applied to the internal luma samples in a CU to predict the chroma samples in the CU.

[0314] CCLM / MMLM with extended range

[0315] One or more reference sample points can be used for CCLM / MMLM prediction, i.e., as Figure 10B shown, the reference region can be the same as the reference region in CCCM. Based on the previous decoded information at the TB / CB / slice / picture / sequence level, different reference regions can be used for CCLM / MMLM prediction.

[0316] Although the training data with multiple reference regions can be well-suited for the calculation of model parameters, in some cases, the training data does not capture the complete characteristics of the test data, which may lead to overfitting and may not be able to predict the test data well (i.e., the chrominance block sample points to be predicted). In addition, different reference regions can be well-adapted to different video block contents, resulting in more accurate predictions. To solve this problem, the reference shape / number of the reference regions can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. A set of reference region candidates can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different components (U / V) can have different reference region switching controls. For example, a set of reference region candidates (idx = 0 to 4) is predefined as shown in the following table. The reference region selection for the U / V components can be switched at the PH or CU / CTU level. Different chrominance types / color formats can have different predefined reference regions.

[0317] The unavailable luma / chroma sample points used to derive the MLR model can be filled from the available reconstructed sample points, and the filling process is applied to both the training data (the top / left adjacent reconstructed luma / chroma sample points) and the test data (the luma / chroma sample points in the CU).

[0318] According to one or more embodiments of the present disclosure, the unavailable luma / chroma sample points used to derive the MLR model can be skipped and not used. Then, the unavailable luma / chroma sample points do not require filling processing.

[0319] FLM / GLM / ELM / CCCM with minimum sample limit

[0320] FLM needs to process the downsampled luminance reference values and calculate model parameters, which imposes a burden on the decoder processing cycle, especially for small blocks. In this section, FLM with a minimum sample limit is proposed. For example, FLM is only used for a number of samples greater than a predetermined number, such as 64, 128. One or more different limits can be used for this purpose. For example, FLM is only used for samples greater than a predefined number (such as 256) in a single model, and FLM is only used for samples greater than a predefined number (such as 128) in a multi-model.

[0321] According to one or more embodiments of the present disclosure, the number of predefined minimum samples for a single model can be greater than or equal to the number of predefined minimum samples for a multi-model. For example, FLM / GLM / ELM / CCCM is only used for samples greater than or equal to a predefined number (such as 128) in a single model, and FLM / GLM / ELM / CCCM is only used for samples greater than or equal to a predefined number (such as 256) in a multi-model.

[0322] According to one or more embodiments of the present invention, the number of predefined minimum samples for FLM / GLM / ELM can be greater than or equal to the number of predefined minimum samples for CCCM. For example, CCCM is only used in a single model for samples greater than or equal to a predefined number (such as 0), and CCCM is only used in a multi-model for samples greater than or equal to a predefined number (such as 128). FLM is only used in a single model for samples greater than or equal to a predefined number (such as 128), and FLM is only used in a multi-model for samples greater than or equal to a predefined number (such as 256).

[0323] Note that the methods / examples in this section can be combined / reused with the above methods, including but not limited to methods related to classification, filter shape, matrix derivation (with special processing), application area, and syntax. In addition, the methods / examples listed in this section can also be applied to the above methods / examples (more taps) to have better performance under certain complexity trade-offs.

[0324] Combined multiple modes of FLM / GLM / ELM / CCCM / CCLM

[0325] According to one or more embodiments of the present disclosure, two models in multiple modes of FLM / GLM / ELM / CCCM / CCLM can be further combined to bring additional codec efficiency. For example, first, the parameters of CCCM(c i ) and GLM(a, b) are derived separately, then the weight (w i ) between CCCM and GLM is derived through linear regression, and finally, the weighted CCCM and GLM are used to predict chrominance samples based on the reconstructed luminance samples. GLMpredChromaVal = a * lumaVal + b CCCMpredChromaVal = c 0 *C + c 1 *N + c 2 *S + c 3 *E + c 4 *W + c 5 *P + c 6 *B FinalpredChromaVal = w 0 *GLMpredChromaVal + w 1 *CCCMpredChromaVal

[0326] According to one or more embodiments of the present disclosure, there is a flag signaled or switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level to indicate whether to use the combined mode.

[0327] According to one or more embodiments of the present disclosure, the mode flag can be derived at the decoder to save bit overhead instead of explicitly signaling the selected mode flag. (1) Determine M combined candidates for the current CU (2) Divide the available L-shaped template area into N regions, denoted as R 0 , R 1 , … R N-1 (Divide the training data into N groups, N-fold training) (3) Independently apply the M combined candidates to a partial available template area (which can be a single or multiple regions among R 0 , R 1 , … R N-1 )(4) Derive M sets of filter coefficients (according to M filter shapes), denoted as F (5) Apply the derived sets of F 0 , F 1 , … F M-1 (6) Apply the derived sets of F 0 , F 1 , … F M-1 to the other part of the available template area, different from the available template part in (3) (6) Accumulate the error by SAD, SSD or SATD, denoted as E 0 , E 1 , … EM-1 (7) Sort and select K minimum errors, denoted as E’ 0 , E’ 1 , … E’ K-1 , corresponding to K combined filters / K sets of filter coefficients (8) Signal and select 1 out of K combined filters to be applied to the current CU for chrominance prediction; if K is 1, no signaling is required (the filter shape with the minimum error is the applied combined filter)

[0328] For example, (1) Pre - define 3 candidate filters for the current CU, e.g., CCCM, GLM, and the combined CCCM and GLM (2) Divide the available L - shaped template region (6 chrominance rows / columns for CCCM, note that in the CCCM design, each chrominance sample involves 6 luminance samples for down - sampling) into 2 regions, denoted as R 0 、R 1 For example, even rows / columns: R 0 , odd rows / columns; R 1 Figure 20A Shows an example where in the template region, the even - row region R 0 is used to train / derive 3 sets of filter coefficients, while the odd - row region R 1 is used to validate / compare and sort the costs of the 3 sets of filter coefficients. (3) Independently apply the 3 candidate filters to a partial available template region (a single R 0 ) (4) Derive 3 sets of filter coefficients (according to 4 filter shapes), denoted as F 0 、F 1 、F 2 (5) Apply the derived F 0 、F 1 、F 2 sets of filter coefficients to other partial available template regions (R 1 ), different from the partial available template in (3) (6) Accumulate the errors through SAD, SSD or SATD, denoted as E 0 、E 1 、E 2 (7) Sort and select 1 minimum error, denoted as E’ 0 , corresponding to 1 filter shape / 1 set of filter coefficients (8) If K is 1, then no signal transmission is required (the filter with the minimum error is the applied filter shape).

[0329] Note that the methods / examples in this section can be combined / reused with the methods mentioned in all sections, including but not limited to classification, filter shape, matrix derivation (with special processing), application area, and syntax. Additionally, the methods / examples listed in this section can also be applied to all sections to achieve better performance under a certain complexity trade-off.

[0330] CCCM with non-downsampled and downsampled luma reference values

[0331] In CCCM, only the downsampled luma reference values can be used to calculate the model parameters and apply the CCCM model. In this section, non-downsampled luma reference values are also used in CCCM to calculate the model parameters and apply the CCCM model, including using non-downsampled and downsampled luma reference values at different or the same positions. As mentioned above, one or more filter shapes can be used for this purpose.

[0332] In one example, a convolutional 8-tap filter can include a 6-tap rectangular spatial component, downsampled luma samples, and a bias term. The input to the 6-tap spatial component of the filter includes the central (b) non-downsampled luma sample that is co-located with the chroma sample to be predicted and its non-downsampled lower-left or southwest (d), lower-right or southeast (f), lower or south (e), left or west (a), and right or east (c) adjacent samples, as shown in Shape 1 of Figure 16 and the central (c) downsampled luma sample that is co-located with the chroma sample to be predicted.

[0333] The bias term B represents a scalar offset between the input and the output (similar to the offset term in CCLM) and is set to the intermediate chroma value (512 for 10-bit content).

[0334] The output of the filter is calculated as the convolution between the filter coefficients c i and the input values, and is clipped to the valid chroma sample range: predChromaVal = c 0 a + c 1 b + c 2 c + c 3 d + c 4 e + c 5 f + c 6 C + c 7 B

[0335] Note that the methods / examples in this section can be combined / reused with the above methods, including but not limited to classification, filter shape, matrix derivation (with special handling), application area, grammar. In addition, the methods / examples listed in this section can also be applied to the above-mentioned methods (more taps) to achieve better performance with a certain complexity trade-off.

[0336] In the present disclosure, a reference sample / training template / reconstructed neighboring region generally refers to a luminance sample used to derive MLR model parameters, which are then applied to internal luminance samples in a CU to predict chrominance samples in the CU.

[0337] Those skilled in the art should understand that the above examples are for illustrative purposes only, and other implementations are possible without departing from the present disclosure.

[0338] Figure 21 The workflow of method 2100 for decoding video data according to one or more aspects of the present disclosure is shown. Method 2100 can be executed by a decoder.

[0339] At step 2110, a video block can be obtained from the bitstream. For example, a block-based video decoder can receive and process the bitstream to obtain the video block, as described with reference to Figure 3 above.

[0340] At step 2120, the decoder can obtain the internal luminance sample values of the video block, the external luminance sample values and the external chrominance sample values of the external region of the video block. For example, the external region of the video block can include one or more rows and / or columns adjacent to the video block. The rows and / or columns can be included in the reconstructed region.

[0341] At step 2130, a set of weighting coefficients corresponding to a filter shape can be determined by using values based on the external luminance sample values and the external chrominance sample values. For example, the filter shape and a set of weighting coefficients can correspond to a CCCM filter and can be configured to predict chrominance sample values based on multiple corresponding luminance sample values. The values used to calculate a set of weighting coefficients can include non-downsampled values of the external luminance sample values, or non-downsampled values of the external luminance sample values together with downsampled values of at least one external luminance sample value among the external luminance sample values.

[0342] In one example, non-downsampled luminance sample values and downsampled luminance sample values can be combined to accommodate different attributes of an image. For example, a luminance sample co-located with a chrominance sample can be referred to herein as a central luminance sample, and both the downsampled value and the non-downsampled value associated with the central luminance sample can be used to calculate a set of weighted coefficients for a CCCM filter. The downsampled value associated with the central luminance sample can be obtained by applying any suitable filter to the neighboring luminance sample values of the central luminance sample.

[0343] In another example, the information of non-downsampled luminance sample values can be more helpful for predicting the corresponding chrominance sample values, and, with or without an offset (B) and / or a non-linear term (P), the weighted coefficients of the CCCM filter can be trained at the decoder using the non-downsampled luminance sample values. For example, any non-downsampled luminance sample values of neighboring luminance samples can be used to derive the weighted coefficients of the CCCM filter. The neighboring luminance samples can be located around the central luminance sample for the chrominance sample.

[0344] At step 2140, a set of weighted coefficients calculated using a filter shape and the internal luminance sample values of a video block can be used to predict the internal chrominance sample values of the video block.

[0345] At step 2150, the video block can be decoded using the predicted internal chrominance sample values to obtain the decoded video block. For example, the decoded video block can be presented on a screen or used for further processing, such as editing.

[0346] Figure 23 The workflow of a method 2300 for encoding video data according to one or more aspects of the present disclosure is shown. Method 2300 can correspond to method 2100 and can be executed by an encoder. Steps 2320, 2330, and 2340 of method 2300 can correspond to steps 2120, 2130, and 2140 of method 2100. At step 2310, the encoder can obtain a video block as described with reference to Figure 1 At step 2350, a bitstream including the encoded video block can be generated at the encoder by using the predicted internal chrominance sample values. The compressed bitstream can be transmitted to a decoder via any suitable medium (such as a wireless or wired network, or a Blu-ray Disc).

[0347] Figure 22 The workflow of a method 2200 for decoding video data according to one or more aspects of the present disclosure is shown. Method 2200 can be executed by a decoder.

[0348] At step 2210, video blocks can be obtained from the bitstream. For example, a block-based video decoder can receive and process the bitstream to obtain video blocks as described with reference to Figure 3 above.

[0349] At step 2220, the internal luma sample values of the video block, the external luma sample values and the external chroma sample values of the first external region of the video block can be obtained at the decoder. The first external region can be included in the reconstruction region and can be a part of the reconstruction region.

[0350] At step 2230, multiple sets of weighted coefficients corresponding to multiple filters can be obtained based on the external luma sample values and the external chroma sample values of the first external region. For example, each of the multiple filters can be independently trained on the first external region to obtain a corresponding set of weighted coefficients for the multiple filters. For example, each of the multiple filters can be any one or any combination of FLM, CCCM, GLM, ELM, CCLM.

[0351] At step 2240, at least one of the multiple filters can be applied to predict the internal chroma sample values of the video block. Determining which one or more of the multiple filters to use to predict the internal chroma samples can be based on explicit signaling sent by the encoder or inferred by the decoder without additional signaling overhead.

[0352] In one example, the decoder can receive signaling in the bitstream that indicates the multiple filters to be used to predict chroma sample values based on multiple corresponding luma sample values. In this case, the multiple filters with the respective sets of weighted coefficients calculated at step 2230 can be used to predict the internal chroma sample values of the video block.

[0353] In another example, the decoder can use the respective sets of weighted coefficients calculated at step 2230 to evaluate the performance of the multiple filters and select at least one of the multiple filters with the highest performance for use. For example, a second external region can be obtained at the decoder to be used as a data set for testing or evaluation. The second external region for evaluation can be different from the first external region used to derive a set of weighted coefficients. One filter with the minimum prediction error among the multiple filters can be selected as the filter to be applied. Optionally, K filters with the K minimum prediction errors among the multiple filters can be selected as the filters to be applied. The weights of the selected K filters (e.g., w i ) can be trained in a manner similar to that for deriving the multiple sets of weighted coefficients of the multiple filters (e.g., by using the first external region or other external regions).

[0354] In some examples, the first external region and the second external region can have any shape. For example, one or more columns and / or rows in the reconstruction region, or partial columns and / or rows, and have any number of samples. For example, the first external region and / or the second external region with a small number of samples can reduce the computational burden on the decoder.

[0355] At step 2250, the video block can be decoded using the predicted intra chroma sample values to obtain the decoded video block. For example, the decoded video block can be presented on a screen or used for further processing, such as editing.

[0356] Figure 24 The workflow of a method 2400 for encoding video data according to one or more aspects of the present disclosure is shown. Method 2400 can correspond to method 2200 and can be performed by an encoder. Steps 2420, 2430, 2440 of method 2400 can correspond to steps 2220, 2230, 2240 of method 2200. At step 2410, the encoder can obtain a video block as described with reference to Figure 1 At step 2450, a bitstream including the encoded video block can be generated at the encoder by using the predicted intra chroma sample values. For example, the compressed bitstream can be transmitted to the decoder through any suitable medium (such as a wireless or wired network, or a Blu-ray Disc).

[0357] In one example, the encoder can use the sets of weighting coefficients calculated at step 2430 to evaluate the performance of multiple filters, and select K filters out of the multiple filters to be used based on the evaluated performance. For example, by determining N filters with the minimum prediction error and selecting K filters to be combined from the N filters. The weights (such as wi) of the selected K filters can be trained in a similar manner as deriving the sets of weighting coefficients of the multiple filters, for example, by using the first external region or other external regions. In this case, the encoder can generate signaling indicating the selected K filters and send the signaling to the decoder through the bitstream generated at step 2450. In other cases, the encoder may not generate signaling indicating which one or more filters to use in cross-component prediction. Instead, one or more filters to be used can be derived at the decoder, for example, by calculating the perspective prediction error of each of the multiple filters with a predefined filter shape and sorting the prediction errors to select the best one or more filters with the minimum prediction error.

[0358] Those skilled in the art should understand that, without departing from the present disclosure, one or more filters to be used for cross-component prediction among multiple filters can be selected based on other key factors (such as, complexity, trade-off between prediction error and complexity, or compatibility).

[0359] Figure 25 An exemplary computing system 2500 is shown in accordance with one or more aspects of the present disclosure. The computing system 2500 may include at least one processor 2510. The computing system 2500 may also include at least one storage device 2520. The storage device 2520 may store computer-executable instructions that, when executed, cause the processor 2510 to perform the steps of the above-described method. The processor 2510 may be a general-purpose processor or may also be implemented as a combination of computing devices, such as, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. The storage device 2520 may store input data, output data, data generated by the processor 2510, and / or instructions executed by the processor 2510.

[0360] It should be understood that the storage device 2520 may store computer-executable instructions that, when executed, cause the processor 2510 to perform any operation according to an embodiment of the present disclosure.

[0361] In one embodiment, a non-transitory computer-readable storage medium is also provided, which includes a plurality of programs, for example, in the storage device 2520, executable by a processor 2510 in a computing environment, for performing the above-described method and / or storing a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. In one example, the plurality of programs may be executed by a processor 2510 in a computing environment to (for example, receive from Figure 1 a video encoder in) a bitstream or data stream including encoded video information (such as, video blocks representing encoded video frames and / or one or more associated syntax elements, etc.), and may also be executed by a processor 2510 in a computing environment to perform the above-described decoding method according to the received bitstream or data stream. In another example, the plurality of programs may be executed by a processor 2510 in a computing environment to perform the above-described encoding method to encode video information (such as, video blocks representing video frames and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and may also be executed by a processor 251 in a computing environment to send the bitstream or data stream (such as, to Figure 3 a video decoder in). Optionally, a bitstream or data stream may be stored in the non-transitory computer-readable storage medium, and the bitstream or data stream includes data encoded by an encoder (such as,Figure 1 The video encoder in [reference] uses the encoded video information generated by, for example, the above-described encoding method (such as video blocks representing encoded video frames and / or one or more associated syntax elements, etc.) for use by a decoder (such as, Figure 3 the video decoder in [reference]) when decoding video data. The non-transitory computer-readable storage medium can be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, and the like.

[0362] In one embodiment, there is provided a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. In one embodiment, there is provided a bitstream including encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method.

[0363] In one embodiment, the computing environment can be implemented using one or more ASICs, DSPs, digital signal processor devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0364] It should be understood that all operations in the above-described method are merely exemplary, and the present disclosure is not limited to any operation in the method or the sequence order of these operations, and should cover all other equivalents under the same or similar concepts.

[0365] It should also be understood that all modules in the above-described method can be implemented in various ways. These modules can be implemented as hardware, software, or a combination thereof. In addition, any one of these modules can be further functionally divided into sub-modules or combined together.

[0366] It should be noted that the terms "first", "second", etc. used in the specification, claims, and drawings of the present disclosure are used to distinguish objects, rather than to describe any specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate conditions, so that the embodiments of the present disclosure described herein can be implemented in an order other than the order shown in the drawings or described in the present disclosure.

[0367] The foregoing description is provided to enable a person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents of the elements of the various aspects described in the present disclosure that are known or will be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims.

Claims

1. A method for decoding video data, comprising: obtaining a video block from a bitstream; obtaining internal luminance sample values of the video block, external luminance sample values of an external region of the video block, and external chrominance sample values; determining a set of weighting coefficients corresponding to a filter shape by using a value based on the external luminance sample values and the external chrominance sample values, wherein the filter shape and the set of weighting coefficients are configured to predict chrominance sample values based on a plurality of corresponding luminance sample values, and the value includes a non-downsampled value of the external luminance sample values, or the non-downsampled value of the external luminance sample values together with a downsampled value of at least one external luminance sample value among the external luminance sample values; predicting internal chrominance sample values of the video block based on the internal luminance sample values by using the filter shape and the set of weighting coefficients; and obtaining a predicted video block by using the predicted internal chrominance sample values.

2. The method according to claim 1, wherein, at least one external luminance sample value among the external luminance sample values includes a central luminance sample value co-located with an external chrominance sample value.

3. The method according to claim 2, wherein, the non-downsampled value of the external luminance samples includes at least one of the following: a non-downsampled luminance sample value of the central luminance sample value, a non-downsampled luminance sample value above the central luminance sample value, a non-downsampled luminance sample below the central luminance sample, a non-downsampled luminance sample value to the left of the central luminance sample value, or a non-downsampled luminance sample value to the right of the central luminance sample value.

4. The method according to claim 2, wherein, the non-downsampled value of the external luminance samples includes at least one of the following: a non-downsampled luminance sample value of the central luminance sample value, a non-downsampled luminance sample value at the lower left corner of the central luminance sample value, a non-downsampled luminance sample value at the lower right corner of the central luminance sample value, a non-downsampled luminance sample value below the central luminance sample value, a non-downsampled luminance sample value to the left of the central luminance sample value, or a non-downsampled luminance sample value to the right of the central luminance sample value.

5. The method according to claim 1, further comprising: determining whether the value includes the non-downsampled value of the external luminance sample values or the non-downsampled value of the external luminance sample values together with a downsampled value of at least one external luminance sample value among the external luminance sample values based on the filter shape.

6. The method according to claim 5, further comprising: if the filter shape is square, determining that the value includes the non-downsampled value of the external luminance sample values.

7. A method for decoding video data, comprising: obtaining a video block from a bitstream; obtaining internal luminance sample values of the video block, external luminance sample values of a first external region of the video block, and external chrominance sample values; determining, based on the external luma sample values ​​and the external chroma sample values ​​of the first external region, a plurality of sets of weighting coefficients corresponding to a plurality of filters, wherein each filter of the plurality of filters is configured to predict a chroma sample value based on a plurality of corresponding luma sample values; Based on the intra luma sample values, predicting intra chroma sample values ​​of the video block using at least one filter of the plurality of filters; and Using the predicted intra chroma sample values, a predicted video block is obtained.

8. The method according to claim 7, further comprising: include: Signaling is received in the bitstream, the signaling indicating whether to use a combination mode to predict chrominance sample values ​​based on a plurality of corresponding luma sample values.

9. The method according to claim 8, in, The signaling is at SPS, DPS, VPS, SEI, APS, PPS, PH, SH, area, CTU, CU, sub-block or sample level.

10. The method according to claim 7, further comprising: include: Signaling is received in the bitstream, the signaling indicating that chrominance sample values ​​are to be predicted based on a plurality of corresponding luma sample values ​​using a plurality of filters.

11. The method according to claim 7, further comprising: include: Obtaining an external luminance sample value and an external chrominance sample value of a second external area of ​​the video block; evaluating a prediction error for each of the plurality of filters based on the outer luma sample values ​​and the outer chroma sample values ​​of the second outer region; as well as The at least one filter among the plurality of filters is determined based on the prediction error.

12. The method according to claim 11, in, A filter having a lowest prediction error among the plurality of filters is determined as the at least one filter.

13. A method for encoding video data, include: Get the video block; Obtaining an internal luminance sample value of the video block, an external luminance sample value and an external chrominance sample value of an external area of ​​the video block; determining a set of weighting coefficients corresponding to a filter shape by using numerical values ​​based on the outer luma sample values ​​and the outer chroma sample values, wherein the filter shape and the set of weighting coefficients are configured to predict chroma sample values ​​based on a plurality of corresponding luma sample values, and the numerical values ​​include non-subsampled values ​​of the outer luma sample values, or non-subsampled values ​​of the outer luma sample values ​​together with a subsampled value of at least one of the outer luma sample values; predicting intra chroma sample values ​​of the video block based on the intra luma sample values ​​using the filter shape and the set of weighting coefficients; and By using the predicted intra chroma sample values, a bitstream including an encoded video block is generated.

14. The method according to claim 13, in, At least one of the outer luma sample values ​​comprises a central luma sample value co-located with an outer chroma sample value.

15. The method according to claim 14, in, The non-downsampled value of the external luminance sample includes at least one of the following: the non-downsampled luminance sample value of the central luminance sample value, the non-downsampled luminance sample value above the central luminance sample value, the non-downsampled luminance sample below the central luminance sample value, the non-downsampled luminance sample value to the left of the central luminance sample value, or the non-downsampled luminance sample value to the right of the central luminance sample value.

16. The method according to claim 14, wherein, The non-downsampled value of the external luminance sample includes at least one of the following: the non-downsampled luminance sample value of the central luminance sample value, the non-downsampled luminance sample value at the lower left corner of the central luminance sample value, the non-downsampled luminance sample value at the lower right corner of the central luminance sample value, the non-downsampled luminance sample value below the central luminance sample value, the non-downsampled luminance sample value to the left of the central luminance sample value, or the non-downsampled luminance sample value to the right of the central luminance sample value.

17. The method according to claim 13, further comprising: Based on the filter shape, determining whether the value includes the non-downsampled value of the external luminance sample value or the non-downsampled value of the external luminance sample value together with the downsampled value of at least one external luminance sample value among the external luminance sample values.

18. The method according to claim 17, further comprising: If the filter shape is square, determining that the value includes the non-downsampled value of the external luminance sample value.

19. A method for encoding video data, comprising: Obtaining a video block; Obtaining the internal luminance sample value of the video block, the external luminance sample value and the external chrominance sample value of the first external region of the video block; Based on the external luminance sample value and the external chrominance sample value of the first external region, determining multiple sets of weighting coefficients corresponding to multiple filters, wherein each filter among the multiple filters is configured to predict a chrominance sample value based on multiple corresponding luminance sample values; Based on the internal luminance sample value, using at least one of the multiple filters to predict the internal chrominance sample value of the video block; and Using the predicted internal chrominance sample value to generate a bitstream including the encoded video block.

20. The method according to claim 19, further comprising: Signaling in the bitstream indicating whether to use a combined mode to predict a chrominance sample value based on multiple corresponding luminance sample values.

21. The method according to claim 20, wherein, The signaling is at the SPS, DPS, VPS, SEI, APS, PPS, PH, SH, region, CTU, CU, sub-block or sample level.

22. The method according to claim 19, further comprising: Obtaining the external luminance sample value and the external chrominance sample value of the second external region of the video block; Based on the external luminance sample value and the external chrominance sample value of the second external region, evaluating the prediction error for each of the multiple filters; and Based on the prediction error, the at least one filter among the plurality of filters is determined.

23. The method according to claim 19, further comprising: signaling in the bitstream, the signaling indicating K numbers of filters among the plurality of filters for predicting chrominance sample values based on a plurality of corresponding luminance sample values.

24. The method according to claim 23, wherein, the filters among the plurality of filters having the K lowest prediction errors are determined as the K numbers of filters.

25. A computer system, comprising: one or more processors; and one or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform the operations of the method according to any one of claims 1 to 24.

26. A computer program product storing computer-executable instructions that, when executed, cause one or more processors to perform the operations of the method according to any one of claims 1 to 24.

27. A computer-readable storage medium storing instructions that, when executed by a computing device having one or more processors, cause the one or more processors to perform the decoding method according to any one of claims 1 to 12, and store a bitstream to be decoded by the decoding method according to any one of claims 1 to 12.

28. A computer-readable storage medium storing instructions that, when executed by a computing device having one or more processors, cause the one or more processors to perform the encoding method according to any one of claims 13 to 24, and store a bitstream generated by the encoding method according to any one of claims 13 to 24.

29. A computer-readable medium storing a bitstream, wherein, the bitstream will be decoded by performing the operations of the method according to any one of claims 1 - 12.

30. A computer-readable medium storing a bitstream, wherein, the bitstream is obtained by performing the operations of the method according to any one of claims 13 - 24.