Method and device for video processing and medium

The cross-component residual model (CCRM) addresses the problem of insufficient encoding and decoding efficiency in existing video encoding and decoding technologies, achieving higher encoding and decoding gains.

CN121176017APending Publication Date: 2025-12-19DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480034460.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-23
Filing Date
2024-05-22
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are insufficient in terms of encoding and decoding efficiency and gain, and need to be further improved.

Method used

The Cross-Component Residual Model (CCRM) is used to improve the efficiency of video encoding and decoding by determining the conversion relationship between video units and bitstreams and performing conversion based on CCRM.

Benefits of technology

It improves the efficiency of video encoding and decoding, and achieves higher encoding and decoding gain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121176017A_ABST
    Figure CN121176017A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises: for a conversion between a video unit of a video and a bitstream of the video, determining one or more cross-component residual models (CCRMs) for the video unit; and performing a conversion based on the one or more CCRMs.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure generally relate to video processing technology, and more particularly, to cross-component models for residual coding. BACKGROUND

[0002] Nowadays, digital video capability is being applied to various aspects of people's life. For video coding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there are some undesirable problems in conventional video coding. Therefore, it is generally desirable to further improve the coding gain of conventional video coding technology. SUMMARY

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method comprises: determining one or more cross-component residual models (CCRM) for a video unit of a video for conversion between the video unit and a bitstream of the video; and performing the conversion based on the one or more CCRM. In this way, it can improve the coding efficiency and achieve higher coding gain.

[0005] In a second aspect, an apparatus for video processing is proposed. The apparatus comprises a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.

[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method comprises: determining one or more cross-component residual models (CCRM) for a video unit of a video; and generating a bitstream of the video unit based on the one or more CCRM.

[0008] In a fifth aspect, a method for storing a bitstream of a video is presented. The method includes determining one or more cross-component residual models (CCRM) for a video unit of the video, generating a bitstream of the video unit based on the one or more CCRM, and storing the bitstream in a non-transitory computer-readable recording medium.

[0009] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other objects, features and advantages of the example embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which like reference characters refer to like elements throughout. In the example embodiments of the present disclosure, the same reference numbers in different drawings identify the same components.

[0011] Figure 1 A block diagram showing an example video coding system is shown in accordance with some embodiments of the present disclosure; Figure 2 A block diagram showing a first example video encoder is shown in accordance with some embodiments of the present disclosure; Figure 3 A block diagram showing an example video decoder is shown in accordance with some embodiments of the present disclosure; Figure 4 A diagram showing the effect of the slope adjustment parameter "u" is shown, where the model created with the current CCLM is shown on the left and the updated model as proposed is shown on the right; Figure 5 Neighboring blocks (L, A, BL, AR, AL) used in derivation of the general MPM list are shown; Figure 6 Neighboring reconstructed samples for DIMD chroma mode are shown; Figure 7 An intra template matching search region used is shown; Figure 8 Use of IntraTMP block vector for IBC block is shown; Figure 9 A partitioning method for angular mode is shown; Figure 10 An extended MRL candidate list is shown; Figure 11 A diagram of a template region is shown; Figure 12 A spatial domain part of a convolution filter is shown; Figure 13Reference region for derivation of filter coefficients (with its padding) is shown; Figure 14 Four Sobel-based gradient patterns for GLM are shown; Figure 15 Spatial GPM candidates are shown; Figure 16 GPM templates are shown; Figure 17 GPM blending is shown; Figure 18 Possible positions of candidate regions are shown; Figure 19 Positions of neighboring spatial candidates are shown; Figure 20 Transform selection process for planar mode with direction is shown; Figure 21 Luma block for derivation of direct block vector is shown; Figure 22 Three types of reconstruction regions defined including thirteen columns or rows of reconstructed pixels are shown; Figure 23 Three types of filter shapes defined with fifteen inputs and one output are shown; Figure 24 Examples of prediction for different positions in the current block are shown; Figure 25 Proposed method on decoder is shown; Figure 26 Luma samples L0,..,L5 relative to chroma sample C (shown in half-pel luma grid) are shown; Figure 27 A flowchart of a method for video processing according to an embodiment of the disclosure is shown; and Figure 28 A block diagram of a computing device in which various embodiments of the disclosure can be implemented is shown.

[0012] Throughout the drawings, identical or similar reference numerals can designate identical or similar elements throughout the several views. DETAILED DESCRIPTION

[0013] The principles of the disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is merely intended to be illustrative and help the person skilled in the art to understand and implement the disclosure, without implying any limitation on the scope of the disclosure. The disclosure described herein can be implemented in various ways in addition to the ways described below.

[0014] In the following description and claims, unless otherwise specified, all technical and scientific terms have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0015] References in the present disclosure to "one embodiment", "an embodiment", "example embodiments", etc., indicate that the embodiment described can include a particular feature, structure, or characteristic, but every embodiment can not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that it is within the knowledge of those skilled in the art to effect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0016] It should be understood that although the terms "first" and "second" etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed terms.

[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including", when used herein, specify the presence of stated features, elements and / or components etc. but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.

[0018] Example Environment Figure 1 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure. As shown, video coding system 100 can include a source device 110 and a destination device 120. Source device 110 can also be referred to as a video encoding device, and destination device 120 can also be referred to as a video decoding device. In operation, source device 110 can be configured to generate encoded video data, and destination device 120 can be configured to decode the encoded video data generated by source device 110. Source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0019] Video source 112 can include a source such as a video capture device, for example. Examples of video capture devices include, but are not limited to, an interface to receive video data from a video content provider, a computer graphics system to generate video data, and / or a combination thereof.

[0020] Video data can comprise one or more pictures. Video encoder 114 encodes video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form an encoded representation of the video data. The bitstream can include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 can include a modulator / demodulator and / or a transmitter. The encoded video data can be transmitted directly to destination device 120 via I / O interface 116 by network 130A. The encoded video data can also be stored onto a storage medium / server 130B for access by destination device 120.

[0021] Destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122. I / O interface 126 can include a receiver and / or a modem. I / O interface 126 can acquire encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 can decode the encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 which is configured to interface with an external display device.

[0022] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other existing and / or further standards.

[0023] Figure 2 FIG. 2 is a block diagram illustrating an example of a video encoder 200 that can be Figure 1 the video encoder 114 in the system 100 shown.

[0024] Video encoder 200 can be configured to implement any or all of the techniques of this disclosure. In Figure 2 In examples, video encoder 200 includes a number of functional components. The techniques described in this disclosure can be shared between the components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0025] In some embodiments, video encoder 200 can include partition unit 201, prediction unit 202, which can include mode select unit 203, motion estimation unit 204, motion compensation unit 205, and intra-prediction unit 206, residual generation unit 207, transform unit 208, quantization unit 209, inverse quantization unit 210, inverse transform unit 211, reconstruction unit 212, buffer 213, and entropy encoding unit 214.

[0026] In other examples, video encoder 200 can include more, less, or different functional components. In one example, prediction unit 202 can include an intra block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0027] Furthermore, although some components, such as motion estimation unit 204 and motion compensation unit 205, can be integrated, for purposes of explanation, these components are shown separately in Figure 2 examples.

[0028] Partition unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.

[0029] Mode select unit 203 can select one of a plurality of encoding modes (intra- or inter- coding) based on, for example, error results, and provide the resulting intra- or inter- coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, mode select unit 203 can select a combined inter-intra prediction (CIIP) mode in which prediction is based on both inter- and intra-prediction signals. In the case of inter-prediction, mode select unit 203 can also select a resolution for motion vectors (e.g., sub-pixel precision or integer pixel precision) for the block.

[0030] To perform inter-prediction for a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture in which the current video block is located.

[0031] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" can refer to a portion of a picture composed of macroblocks all of which are based on macroblocks within the same picture. Further, as used herein, a "P slice" and a "B slice" can refer, in some aspects, to portions of a picture composed of macroblocks that are independent of macroblocks in the same picture.

[0032] In some examples, motion estimation unit 204 can perform uni-prediction on a current video block, and motion estimation unit 204 can search a reference picture in List 0 or List 1 for a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference picture in List 0 or List 1 containing the reference video block, and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.

[0033] Alternatively, in other examples, motion estimation unit 204 can perform bi-prediction on a current video block. Motion estimation unit 204 can search a reference picture in List 0 for one reference video block for the current video block, and can also search a reference picture in List 1 for another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices indicating multiple reference pictures in List 0 and List 1 containing the multiple reference video blocks, and multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and the multiple motion vectors for the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information for the current video block.

[0034] In some examples, motion estimation unit 204 can output a full set of motion information for use in decoding processing by a decoder. Alternatively, in some embodiments, motion estimation unit 204 can reference motion information of another video block to signal motion information of the current video block. For example, motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0035] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0036] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0037] As discussed above, the video encoder 200 can signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0038] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0039] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0040] In other examples, such as in skip mode, there can be no residual data for the current video block for the current video block, and the residual generation unit 207 can not perform the subtraction operation.

[0041] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0042] After the transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, the quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0043] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0044] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0045] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0046] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.

[0047] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0048] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.

[0049] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy encoded video data, and motion compensation unit 302 can determine motion information from the entropy decoded video data, including motion vectors, motion vector precision, reference picture list index, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge modes. AMVP is used, including deriving a number of most probable candidates based on data from neighboring PBs and reference pictures. The motion information typically includes a horizontal motion vector displacement value and a vertical motion vector displacement value, one or two reference picture indices, and in the case of a prediction region in a B slice, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" can refer to deriving motion information from a spatially or temporally neighboring block.

[0050] Motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax elements.

[0051] Motion compensation unit 302 can use the interpolation filter used by video encoder 200 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from the received syntax information, and motion compensation unit 302 can use the interpolation filter to generate the prediction block.

[0052] Motion compensation unit 302 can use at least some of the syntax information to determine the size of the blocks used to encode frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice can be an entire picture, or can also be a region of a picture.

[0053] Intra prediction unit 303 can use intra prediction modes, e.g., received in the bitstream, to form a prediction block from spatial neighboring blocks. Dequantization unit 304 dequantizes quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0054] The reconstruction unit 306 can obtain the decoded block, e.g., by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If needed, a deblocking filter can also be applied to filter the decoded block in order to remove blocking artifacts. The decoded video blocks are then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and which also produces decoded video for presentation on a display device.

[0055] Some example embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for convenience only and not to be construed as limiting the embodiments disclosed in that section to that section only. Furthermore, although some embodiments are described with reference to a multi-functional video codec or other specific video codec, the disclosed techniques are applicable to other video coding technologies. Moreover, although some embodiments describe video encoding steps in detail, it will be appreciated that corresponding decoding steps will be implemented by a decoder. Furthermore, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compressed format to another compressed format or at a different compressed bit rate.

[0056] 1. BRIEF OVERVIEW The present disclosure relates to video coding techniques. In particular, it is about chroma prediction in image / video coding. It can be applied to existing video coding standards such as HEVC, VVC, etc. It can also be applicable to future video coding standards or video codecs.

[0057] 2. INTRODUCTION Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. The ITU-T produced H.261 and H.263 standards, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) standards and the H.265 / HEVC standard. From H.262, video coding standards are based on the hybrid video coding structure, where temporal prediction plus transform coding are utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. The JVET meeting is held once every quarter, and the new video coding standard was officially named as Versatile Video Coding (VVC) in the April 2018 JVET meeting, and the first version of VVC test model (VTM) was released at that time. The working draft of VVC and the test model VTM are updated after each meeting. The VVC project achieved the technical completion (FDIS) in the July 2020 meeting.

[0058] 2.1 Intra prediction In intra prediction, the minimum chroma intra prediction unit (SCIPU) constraint in VVC is removed. In addition, the VPDU constraint for reducing CCLM prediction delay is also removed.

[0059] 2.1.1 Multi-model LM (MMLM) The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, the reconstructed neighboring samples are classified into two categories using a threshold that is the average of the luma reconstructed neighboring samples. A linear model for each category is derived using the least mean square (LMS) method. For the CCLM mode, a linear model is also derived using the LMS method. A slope adjustment is applied to the cross-component linear model (CCLM) and multi-model LM prediction. The adjustment is a linear function that maps luma values to chroma values with a tilt relative to a center point determined by the average luma value of the reference samples.

[0060] 2.1.1.1 Slope adjustment for CCLM CCLM maps luma values to chroma values using a model with 2 parameters. The slope parameter “a” and the bias parameter “b” define the following mapping:

[0061] The adjustment “u” to the slope parameter is signaled to update the model to the following form:

[0062] where

[0063] By this selection, the mapping function is tilted or rotated around the point with luminance value y r The average value of the reference luma samples used in the model creation is used as y r in order to provide a meaningful modification for the model. The following picture illustrates this process. Figure 4 A plot showing the effect of the slope adjustment parameter "u". Left: model created with the current CCLM. Right: updated model as proposed.

[0064] Figure 4 A plot showing the effect of the slope adjustment parameter "u". Left: model created with the current CCLM. Right: updated model as proposed.

[0065] Implementation The slope adjustment parameter is provided as an integer between -4 and 4 (inclusive) and is signaled in the bitstream. The unit of the slope adjustment parameter is 1 / 8 of a chroma sample value per luma sample value (for 10-bit content).

[0066] The adjustment applies to CCLM models using reference samples both above and to the left of the block ("LM_CHROMA IDX" and "MMLM_CHROMA IDX"), but not to "single-side" modes. This selection is based on a trade-off between coding efficiency and complexity.

[0067] When applying slope adjustment to multi-mode CCLM models, both models can be adjusted, so up to two slope updates are signaled for a single chroma block.

[0068] Encoder scheme The proposed encoder method performs a SATD-based search for the best value of slope update for Cr and a similar SATD-based search for Cb. If either results in a non-zero slope adjustment parameter, the combined slope adjustment pair (SATD-based update for Cr, SATD-based update for Cb) is included in the list of RD checks for the TU.

[0069] 2.1.2 Gradient PDPC In VVC, for some scenarios, PDPC can not be applied due to the unavailability of the secondary reference samples. In these cases, gradient-based PDPC extended from horizontal / vertical mode is applied. The PDPC weight (wT / wL) and nScale parameters used to determine the decay of the PDPC weight with respect to the distance from the left / above boundary are set equal to the corresponding parameters in the horizontal / vertical mode. When the secondary reference samples are located at fractional sample positions, bilinear interpolation is applied.

[0070] 2.1.3 Secondary MPM A secondary MPM list is introduced. The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list includes 16 entries. A general MPM list with 22 entries is first constructed, then the first 6 entries in the general MPM list are included into the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the planar mode. The remaining entries consist of the intra modes of the left (L), above (A), below-left (BL), above-right (AR), and above-left (AL) neighboring blocks, the band direction modes with offsets added to the first two available band directions of the neighboring blocks, and the default mode.

[0071] If the CU block is vertically oriented, the order of the neighboring blocks is A, L, BL, AR, AL; otherwise, L, A, BL, AR, AL. Figure 5 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown.

[0072] The PMPM flag is first parsed, if equal to 1, the PMPM index is parsed to determine which entry of the PMPM list is selected, otherwise the SPMPM flag is parsed to determine whether the SMPM index is parsed or the remaining modes are parsed.

[0073] 2.1.4 Reference sample interpolation and smoothing for intra prediction 4-tap cubic interpolation is replaced by a 6-tap cubic interpolation filter for deriving the prediction samples from the reference samples.

[0074] For reference sample filtering, a 6-tap Gaussian filter is applied for larger blocks (W >= 32 and H >= 32), otherwise the existing VVC 4-tap Gaussian interpolation filter is applied. The extended intra reference samples are derived using the 4-tap interpolation filter instead of nearest-neighbor rounding.

[0075] 2.1.5 Decoder-side intra mode derivation (DIMD) When DIMD is applied, two intra modes are derived from the reconstructed neighboring samples and these two predictors are combined with the planar mode predictor with weights derived from the gradients. The division operations in the weight derivation are performed using the same LUT based integerization scheme used by CCLM. For example, the division operation in the orientation calculation

[0076] is calculated by the following LUT based scheme:

[0077] where DivSigTable

[16] = { 0, 7, 6, 5,5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}.

[0078] The derived intra mode is included in the main list of intra most probable modes (MPM), so the DIMD process is performed before the MPM list is constructed. The main derived intra mode of a DIMD block is stored together with the block and is used for the MPM list construction of neighboring blocks.

[0079] 2.1.5.1 DIMD chroma mode The DIMD chroma mode uses the DIMD derivation method to derive the chroma intra prediction mode of the current block based on the neighboring reconstructed Y, Cb and Cr samples in the second neighboring row and column. Specifically, the horizontal and vertical gradients are calculated for each co-located reconstructed luma sample and the reconstructed Cb and Cr samples of the current chroma block to construct the HoG. Then the intra prediction mode with the largest histogram amplitude value is used to perform the chroma intra prediction of the current chroma block. Figure 6 The neighboring reconstructed samples used for the DIMD chroma mode are shown.

[0080] When the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode with the second largest histogram amplitude value is used as the DIMD chroma mode. A CU level flag is signaled to indicate whether the proposed DIMD chroma mode is applied.

[0081] 2.1.6 Fusion of chroma intra prediction modes The DM mode and the four default modes can be fused with the MMLM LT mode as follows:

[0082] where is the prediction value obtained by applying the non-LM mode, is the prediction value obtained by applying the MMLM LT mode, and is the final prediction value of the current chroma block. The two weights and are determined by the intra prediction mode of the neighboring chroma blocks, and is set equal to 2. Specifically, when both the above and left neighboring blocks are coded with LM mode, = {1, 3}; when both the above and left neighboring blocks are coded with non-LM mode, = {3, 1}; otherwise, = {2, 2}.

[0083] For the syntax design, if non-LM mode is selected, a flag is signaled to indicate whether fusion is applied or not. This method is only applied to I slices.

[0084] 2.1.7 Intra Template Matching Intra Template Matching Prediction (Intra TMP) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the most similar template to the current template in the reconstructed part of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode and the same prediction operation is performed at the decoder side.

[0085] The prediction signal is generated by matching the L-shaped causal neighbors of the current block with another block in Figure 7 , which consists of the following parts: R1: the current CTU R2: the top-left CTU R3: the above CTU R4: the left CTU The sum of absolute differences (SAD) is used as the cost function.

[0086] Within each region, the decoder searches for the template with the smallest SAD with respect to the current template and uses its corresponding block as the prediction block.

[0087] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is:

[0088] where “ ” is a constant that controls the gain / complexity trade-off. In practice, “ " is equal to 5. Figure 7 The intra template matching search region used is shown.

[0089] To speed up the template matching process, the search range of all search regions is down-sampled by a factor of 2. This results in a reduction of the template matching search by a factor of 4. After finding the best match, a refinement process is performed. The refinement is done by a second template matching search around the best match with a reduced range. The reduced range is defined as min(BlkW, BlkH) / 2.

[0090] The intra template matching tool is enabled for CUs with width and height size smaller or equal to 64. This maximum CU size for intra template matching is configurable.

[0091] When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level by a dedicated flag.

[0092] 2.1.7.1 IntraTMP derived block vector candidates for IBC In this method, the block vectors (BVs) derived from intra template matching prediction (IntraTMP) are used for intra block copy (IBC). The stored IntraTMP BVs of neighboring blocks as well as the IBC BVs are used as spatial BV candidates in the IBC candidate list construction.

[0093] The IntraTMP block vectors are stored in the IBC block vector buffer and the current IBC block can use both the IBC BVs and the IntraTMP BVs of neighboring blocks as BV candidates for the IBC BV candidate list as shown in Figure 8 .

[0094] The IntraTMP block vectors are added to the IBC block vector candidate list as spatial candidates.

[0095] 2.1.8 Fusion for template based intra mode derivation (TIMD) For each intra prediction mode in the MPMs, the SATD between the predicted samples of the template and the reconstructed samples is calculated. The first two intra prediction modes with the smallest SATD are selected as the TIMD modes. These two TIMD modes are fused with a weight after applying the PDPC process and this weighted intra prediction is used to code the current CU. The position dependent intra prediction combination (PDPC) is included in the derivation of the TIMD modes.

[0096] The cost of the two selected modes is compared with a threshold, a cost factor of 2 is applied in the test as follows:

[0097] If the condition is true, the fusion is applied, otherwise only mode1 is used.

[0098] The weight of the mode is computed from its SATD cost as follows:

[0099] The division operation is done using the same LUT-based integerization scheme used by CCLM.

[0100] 2.1.9 Intra prediction merge This intra prediction method derives the prediction samples as a weighted combination of multiple predictors generated from different reference lines. In this process, multiple intra prediction values are generated and then fused by weighted averaging. The process of deriving the predictors to be used in the fusion process is described as follows: For angular intra prediction modes including TIMD and DIMD, the proposed method derives the intra prediction by weighting the intra predictions obtained from multiple reference lines represented as where is the intra prediction from the default reference line and is the prediction from the line above the default reference line. The weights are set to and .

[0101] For TIMD mode with mixing, is used for the first mode ( ) and is used for the second mode ( ).

[0102] For DIMD mode with mixing, the number of predictors selected for the weighted averaging is increased from 3 to 6.

[0103] When the angular intra mode has a non-integer slope (required reference sample interpolation) and the block size is greater than 16, the intra prediction merge method is applied to the luma block which is used with MRL and is not applied to ISP coded blocks. In the method investigated in sub-test a, PDPC is applied to the intra prediction mode using the reference line closest to the current block.

[0104] 2.1.10 Combination of CIIP with TIMD and TM Merge In CIIP mode, prediction samples are generated by weighting the inter-prediction signal using CIIP-TM Merge candidate prediction and the intra-prediction signal using the intra-prediction mode derived using TIMD. This method is only applied to codec blocks with an area of ​​1024 or less.

[0105] The TIMD derivation method was used to derive intra-prediction modes in CIIP. Specifically, the intra-prediction mode with the smallest SATD value in the TIMD mode list was selected and mapped to one of 67 regular intra-prediction modes.

[0106] Furthermore, it is proposed that if the derived intra-prediction mode is an angle mode, the weights (wIntra, wInter) for the two tests should be modified. For near-horizontal mode (2 <= angle mode index < 34), the current block is vertically partitioned; for near-vertical mode (34 <= angle mode index <= 66), the current block is horizontally partitioned.

[0107] For different sub-blocks (wIntra, wInter) such as Figure 9 As shown.

[0108] Table 1. Weights used for modification of angular patterns

[0109] Using CIIP-TM, a CIIP-TM Merge candidate list is constructed for the CIIP-TM pattern. Merge candidates are refined through template matching. CIIP-TM Merge candidates are also reordered as regular Merge candidates using the ARMC method. The maximum number of CIIP-TM Merge candidates is 2.

[0110] 2.1.11 Extended Multi-Reference Row (MRL) List The MRL list in VVC is expanded to include more reference lines for intra-frame prediction. The expanded reference line list consists of line indices {1, 3, 5, 7, 12}. For Template-Based Intra-Frame Mode Derivation (TIMD), only the first two reference line candidates (i.e., {1, 3}) are used, instead of the complete MRL candidate list. Figure 10 An expanded list of MRL candidates is shown.

[0111] 2.1.12 Template-based multi-reference row intra-frame prediction A template-based multi-reference line intra prediction (TMRL) mode combines reference lines and prediction modes together and uses a template matching method to construct a list of candidate combinations. An index of the coded candidate combination list is signaled to indicate which reference line and prediction mode are used when coding the current block. The regular multi-reference line (MRL) for non-TIMD parts is replaced by the TMRL mode.

[0112] The TMRL mode extends the reference line candidate list and the intra prediction mode candidate list. The extended reference line candidate list is {1, 3, 5, 7, 12}. The restriction on the top CTU row remains unchanged. The size of the intra prediction mode candidate list is 10. The construction of the intra prediction mode candidate list is similar to MPM, except that the planar mode is excluded from the intra prediction mode candidate list, the DC mode is added after the 5 neighboring PU modes and the DIMD mode if it is not included, and the angular modes with the delta angles from to are added.

[0113] The TMRL candidate is constructed as follows. There are 5x10=50 combinations of the extended reference lines and allowed intra prediction modes for a block. Since the extended reference lines start from reference line 1, the area covered by reference line 0 is used for template matching. The SAD cost on the template area (see Figure 11 ) is calculated between the prediction (generated by the 50 combinations) and the reconstruction. The 20 combinations with the smallest SAD cost are selected in ascending order to form the TMRL candidate list.

[0114] For TMR signaling, instead of coding the reference line and intra mode directly, the index of the TMRL candidate list is coded to indicate which combination of the reference line and prediction mode is used to code the current block.

[0115] 2.1.13 Convolutional Cross-Component Intra Prediction Model In this method, a convolutional cross-component model (CCCM) is applied to predict the chroma samples from the reconstructed luma samples, in the spirit similar to what is done by the current CCLM mode. As with CCLM, when chroma downsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid. Similar to CCLM, the top, left, or both top and left reference samples are used as the template for model derivation.

[0116] Furthermore, similar to CCLM, there is the option to use a single model or a multi-model variant of CCCM. The multi-model variant uses two models, one model is derived for samples above the average luma reference value and the other model is derived for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples.

[0117] 2.1.13.1 Convolutional filter The convolutional 7-tap filter consists of a 5-tap plus sign-shaped spatial component, a non-linear term and a bias term. The input of the spatial 5-tap component of the filter consists of the center (C) luma sample co-located with the chroma sample to be predicted and its above / north (N), below / south (S), left / west (W) and right / east (E) neighbors, as follows. Figure 12 The spatial part of the convolutional filter is shown.

[0118] The non-linear term P is expressed as the 2nd power of the center luma sample C and is scaled to the sample value range of the content:

[0119] I.e. for 10-bit content, it is calculated as:

[0120] The bias term B represents a scalar offset between the input and the output (similar to the offset term in CCLM) and is set to the middle chroma value (512 for 10-bit content).

[0121] The output of the filter is calculated as the convolution of the filter coefficients c i and the input values, and is clipped to the range of valid chroma samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B 2.1.13.2 Calculation of filter coefficients The filter coefficients c i are calculated by minimizing the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region. Figure 13 The reference region consisting of the 6 rows of chroma samples above and to the left of the PU is shown. The reference region extends one PU width to the right and one PU height below the PU boundary. The region is adjusted to only include available samples. The extension of the region shown in blue is required to support the "edge samples" of the plus-shaped spatial filter and is padded when not available.

[0122] MSE minimization is performed by computing the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is LDL-decomposed, and the final filter coefficients are computed using back-substitution. This process roughly follows the computation of the ALF filter coefficients in the ECM, however, LDL-decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations.

[0123] The autocorrelation matrix is computed using the reconstructed values of the luma and chroma samples. These samples are full range (e.g., between 0 and 1023 for 10-bit content), resulting in relatively large values in the autocorrelation matrix. This requires high bit-depth operations during the model parameter computation. It is proposed to remove a fixed offset from the luma and chroma samples in each PU for each model. This reduces the amplitude of the values used in the model creation and allows to reduce the precision required for fixed-point arithmetic. As a result, it is proposed to use 16-bit decimal precision instead of the 22-bit precision of the original CCCM implementation.

[0124] For simplicity, the reference sample values immediately outside the top-left corner of the PU are used as offsets (offsetLuma, offsetCb, and offsetCr). The sample values used in both the model creation and the final prediction (i.e., the luma and chroma in the reference region and the luma in the current PU) are reduced by these fixed values as follows: C' = C - offsetLuma N' = N - offsetLuma S' = S - offsetLuma E' = E - offsetLuma W' = W - offsetLuma P' = nonLinear(C') B = midValue = 1 « (bitDepth - 1) And the chroma values are predicted using the following equation, where offsetChroma is equal to offsetCr for the Cr component and to offsetCb for the Cb component: predChromaVal = c0C' + c1N' + c2S' + c3E' + c4W' + c5P' + c6B + offsetChroma To avoid any additional sample-level operations, the luma offset is removed during luma reference sample interpolation. This can be done, for example, by replacing the rounding term used in luma reference sample interpolation with an updated offset that includes both the rounding term and the offsetLuma. The chroma offset can be removed by directly subtracting the chroma offset from the reference chroma samples. As an alternative, the effect of the chroma offset can be removed from the cross-component vector, resulting in the same outcome. To add the chroma offset back to the output of the convolutional prediction operation, the chroma offset is added to the bias term of the convolutional model.

[0125] The process of CCCM model parameter calculation requires division operations. Division operations are not always considered friendly to implementations. Division operations are replaced by multiplication (with a scaling factor) and shift operations, where the scaling factor and the number of shifts are calculated based on the denominator, similar to the method used in the calculation of CCLM parameters.

[0126] 2.1.13.3 Gradient Linear Model For YUV 4:2:0 color format, the Gradient Linear Model (GLM) method can be used to predict chroma samples from luma sample gradients. Two modes are supported: two-parameter GLM mode and three-parameter GLM mode.

[0127] Compared to CCLM, two-parameter GLM utilizes luma sample gradients to derive the linear model, instead of down-sampled luma values. Specifically, when two-parameter GLM is applied, the input of the CCLM process (i.e., down-sampled luma samples ) is replaced by luma sample gradients . Other parts of CCLM (e.g., parameter derivation, prediction sample linear transformation) remain unchanged.

[0128]

[0129] In three-parameter GLM, chroma samples can be predicted based on both luma sample gradients and down-sampled luma values with different parameters. The model parameters of three-parameter GLM are derived from 6 rows and columns of neighboring samples by the LDL-decomposition based MSE minimization method as used in CCCM.

[0130]

[0131] For signaling, when CCLM mode is enabled for the current CU, one flag is signaled to indicate whether GLM is enabled for both Cb and Cr components; if GLM is enabled, another flag is signaled to indicate which of the two GLM modes is selected, and further one syntax element is signaled to select one of the 4 gradient filters for gradient calculation.

[0132] Enable four gradient filters for GLM, such as Figure 14 As shown.

[0133] 2.1.13.4 Bitstream Signaling The use of this mode is signaled via a PU-level flag encoded and decoded by CABAC. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered a sub-mode of CCLM. That is, the CCCM flag is signaled only when the intra-frame prediction mode is LM_CHROMA.

[0134] 2.1.14 Spatial Geometric Partitioning Model (SGPM) SGPM is an intra-frame mode of inter-frame coding / decoding tools similar to GPM, where two prediction parts are generated from the intra-frame prediction process. In this mode, a candidate list is constructed, where each entry contains a segmentation partition and two intra-frame prediction modes, such as... Figure 15 As shown, 26 segmentation modes and 3 intra-frame prediction modes were used to form a combination. The length of the candidate list was set to 16. The selected candidate indices were transmitted via signaling.

[0135] Reorder the list using a template ( Figure 16 The SAD between the template's prediction and reconstruction was used for sorting. The template size was fixed at 1.

[0136] For each segmentation pattern, the same intra-to-inter-frame GPM list is used to derive the IPM list for each segment. The IPM list size is set to 3. In the list, the TIMD-derived pattern is replaced by two derived patterns with horizontal and vertical orientations.

[0137] The SGPM pattern is applied with a restricted block size, which is: 4 <= width <= 64, 4 <= height <= 64, and width < height. 8. Height < Width 8. Width Height >= 32.

[0138] Adaptive blending has also been used in spatial GPM, where Figure 17 The mixing depth τ shown is derived as follows:

[0139] 2.1.15 Nonlocal cross-component prediction Cross-component prediction (CCP) including CCLM, CCCM and their variants is adopted by ECM to exploit cross-component correlation. With CCLM or CCCM, the training samples are always adjacent to the current block. However, the cross-component relationship of the current block can be more relevant to the cross-component relationship of a non-local region.

[0140] A method of non-local cross-component prediction is proposed to enhance CCP by exploiting more advantages from non-local regions.

[0141] Method #1: A non-adjacent cross-component prediction (NA-CCP) mode is proposed. With NA-CCP mode, samples in a region that is not adjacent to the current block can be used to derive the CCCM model for the current block. A candidate region list with 6 candidates is constructed by sequentially checking potential 8x8 regions. If a checked region is available, it is put into the candidate region list. The top-left positions of potential 8x8 regions are pre-determined as {(-xStep, 0), (0, -yStep), (xStep, -yStep), (-xStep, yStep), (-xStep, -yStep), (-2 xStep, 0), (0, -2 yStep), (-2 xStep, 2 yStep), (2 xStep, -2 yStep), (-2 xStep, yStep), (xStep, -2 yStep), (-2 xStep, -yStep), (-xStep, -2 yStep), (-2 xStep, -2 yStep), (-xStep / 2, 0), (0, -yStep / 2), (xStep / 2, -yStep / 2), (-xStep / 2, yStep / 2), (-xStep / 2, -yStep / 2)}, where xStep = Max(width, 16), yStep = Max(height, 16). Figure 18 Some possible positions of the candidate regions are shown.

[0142] A flag is signaled to indicate whether NA-CCP is applied to chroma blocks. If NA-CCP is applied, an index is signaled to indicate which candidate in the candidate region list is used to derive the CCCM model.

[0143] Method #2: A history-based cross-component prediction (H-CCP) mode is proposed. With H-CCP, a H-CCLM table and a H-CCCM table are maintained similar to HMVP table. After decoding a block coded with CCLM or CCCM, the corresponding table is updated. In the implementation of H-CCP, the size of H-CCLM table or H-CCCM table is 6. If the current block is coded with CCLM or CCCM mode, a flag is signaled to indicate whether H-CCP is applied or not. If H-CCP is used, further an index is signaled to indicate which candidate model in H-CCLM table or H-CCCM table is selected.

[0144] 2.1.16 Cross-component Merge mode for intra coding of chroma Cross-component prediction (CCP) including cross-component linear model (CCLM), convolution cross-component model (CCCM) and gradient linear model (GLM) is adopted by ECM to exploit cross-component correlation. Cross-component Merge (CCMerge) mode is proposed as a new CCP mode. The cross-component model parameters of the current chroma block coded with CCMerge can be inherited from the neighboring blocks coded with CCP. With CCMerge, CCP can be more efficient and has less signaling overhead.

[0145] In CCMerge, the final cross-component model parameters of the current chroma block can be inherited from its spatial neighboring and non-neighboring neighbors or default model. A list is created which includes CCP models from spatial neighboring and non-neighboring neighbors coded with CCLM, MMLM, CCCM, GLM, chroma fusion and CCMerge mode. After including neighboring CCP models, default models are further included to fill the remaining empty positions in the list. To avoid including redundant CCP models in the list, a de-duplication operation is applied. More details are described as follows. Figure 19 The positions of neighboring spatial candidates are shown.

[0146] Spatial neighboring neighboring candidates The positions of spatial neighboring candidates are shown as Figure 19 . Spatial candidates are included in the following order: B1 -> A1 -> B0 -> A0 -> B2.

[0147] Spatial non-neighboring neighboring candidates After checking all spatial neighboring neighbors, spatial non-neighboring neighboring candidates are considered. In the current ECM design, in inter Merge mode, two groups of spatial non-neighboring neighboring candidates are obtained. In the proposed method, the positions and inclusion order of spatial non-neighboring neighboring candidates from the first group are used.

[0148] CCLM candidate with default scaling parameters After including spatial neighboring and non-neighboring candidates, if the list is not full, CCLM candidates with default scaling parameters are considered. The default scaling parameters are {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8}, and the offset parameters are derived according to the selected default scaling parameters, the average neighboring reconstructed luma sample values (Yavg) and the average neighboring reconstructed Cb / Cr sample values (Cavg).

[0149] 2.1.16.1 Merge model candidates When merging a CCLM candidate, only the scaling parameters are inherited. The offset parameters are derived by using the inherited scaling parameters, Yavg and Cavg.

[0150] When merging a MMLM candidate, the scaling parameters and the classification thresholds are inherited. The offset parameters in each class are derived according to the inherited classification thresholds and Yavg and Cavg in each class. If there is no available neighboring reconstructed sample in a certain class, the offset parameter is directly inherited from the candidate.

[0151] When merging a CCCM candidate, all the convolution parameters, offsets (i.e., offsetLuma, offsetCb and offsetCr) and the classification thresholds are inherited.

[0152] When merging a GLM candidate, if the GLM candidate is a 3 -parameter GLM mode, all the gradient pattern indices and the model parameters are inherited; otherwise, if the GLM candidate is a 2 -parameter GLM mode, the offset parameters are derived by using the inherited scaling parameters, Yavg and Cavg.

[0153] When merging a chroma fusion candidate, the derived MMLM parameters are inherited and used as the merged MMLM candidate.

[0154] For a CCMerge block, if its merged candidate mode is CCLM, MMLM, CCCM or GLM, the merged candidate mode is stored as the propagated mode of the current chroma block; otherwise, if its merged candidate mode is chroma fusion, the propagated mode is set to MMLM. When merging a CCMerge candidate, how to inherit or derive the CCP parameters depends on the propagated mode of the CCMerge candidate, as described in the above five paragraphs.

[0155] 2.1.16.2 Signaling After the cclm mode flag syntax element, an additional flag is signaled to indicate whether CCMerge is used or not. If CCMerge is used, additionally, a candidate index is signaled. The candidate index signaled is shared for Cb / Cr color components. Currently, the maximum allowed number of candidates is set to the default value 6. If the maximum allowed number of candidates is modified to 1, the candidate index does not need to be signaled. Each bin of the candidate index is context coded with a separate context.

[0156] 2.1.17 Planar mode with direction Two additional planar modes, where only horizontal interpolation or only vertical interpolation is used to obtain the prediction samples.

[0157] For the planar horizontal mode, only horizontal linear interpolation is performed based on the left reference sample and the top-right reference sample to predict the current sample as:

[0158] For the planar vertical mode, only vertical linear interpolation is performed based on the top reference sample and the bottom-left reference sample to predict the current sample as:

[0159] The transform kernel selection for the planar horizontal and planar vertical modes is shown as Figure 20 If the intra prediction mode of the current block is the planar vertical mode, the horizontal intra prediction mode is used to derive the transform kernels in the MTS set and the LFNST set. Furthermore, if the intra prediction mode of the current block is the planar horizontal mode, the vertical intra prediction mode is used to derive the transform kernels in the MTS set and the LFNST set.

[0160] 2.1.18 Direct block vector for chroma blocks Direct block vector is used for chroma blocks in dual tree slices. When the chroma dual tree is activated, one flag is signaled to indicate whether the chroma block is coded using IBC mode or not. If Figure 21 one of the five positions shown in the luminance block is coded in IBC or intra TMP mode, its block vector is scaled and used as the block vector for the chroma block. Template matching is used to perform the block vector scaling.

[0161] 2.1.19 Extrapolation filter based intra prediction mode (EFI mode) The proposed extrapolation filter based intra prediction is processed in two steps. First, extrapolation filter coefficients are obtained from the neighboring reconstructed pixels of the current block with a predetermined template. Second, extrapolation generates prediction values from top-left to bottom-right position by position within the current block.

[0162] 2.1.19.1 Search for mean, minimum and maximum Similar to CCCM mode, the mean should be removed when feeding the input to the EIP filter. The value of the DC mode of the current block is used as the mean for EIP prediction. The minimum and maximum are searched from the reconstructed pixels in the reconstruction area with thirteen columns and thirteen rows.

[0163] 2.1.19.2 Calculation of filter coefficients Three types of reconstruction areas and three filter shapes are proposed, as shown in Figure 22 The three types of reconstruction areas defined include thirteen columns or rows of reconstructed pixels. When the current block is predicted using the proposed EIP mode, the decoder decodes the relevant syntax elements to determine the selected reconstruction area type and filter shape for the current block.

[0164] Figure 22 The three types of reconstruction areas defined including thirteen columns or rows of reconstructed pixels are shown.

[0165] Figure 23 The three types of filter shapes defined with fifteen inputs and generating one output are shown.

[0166] The selected filter is slid in the selected reconstruction area with a step of one pixel to collect the input samples and output samples of the EIP. The autocorrelation matrix and cross-correlation vector are constructed while removing the mean from the input samples and output samples. Then, the EIP coefficients are obtained by the same method as in CCCM.

[0167] 2.1.19.3 Prediction of the current block The EIP mode predicts the current block position by position, as shown in Figure 24 .

[0168] For the positions located in the top-left of the current block, the input of the EIP filter is the reconstructed samples.

[0169] For the positions located on the boundary of the current block, part of the input of the EIP filter is the reference samples, and part of the input of the EIP filter is the previously predicted samples.

[0170] For other positions in the current block, the input of the EIP filter is the previously predicted samples.

[0171] To reduce the prediction error, the searched minimum and maximum are applied to limit the output range of each predicted value,

[0172] It is the predicted value at (x, y) in the current block. These are the minimum and maximum values ​​searched from the thirteen reconstructed columns and rows. It is the derived EIP filter. coefficient, These are the reconstructed or predicted values ​​used for the current location. It is a value calculated using the DC prediction model.

[0173] 2.2 Cross-component residual model (CCRM) for inter-frame prediction 2.2.1 Introduction It is proposed that when blocks use inter-frame prediction or intra-block copy (IBC), a cross-component residual model (CCRM) is applied to predict chrominance samples from reconstructed luminance samples. Figure 25 The decoder side of the method is shown. A cross-component filter is derived using the predicted signals for both luminance and chrominance. The derived filter is applied to the reconstructed luminance signal to produce the final chrominance prediction.

[0174] 2.2.2 Calculation of Convolution Filter and Filter Coefficients The proposed 8-tap filter consists of 6 spatial brightness samples, a nonlinear term, and a bias term. For example... Figure 26 As shown, the spatial luminance samples (L0, ..., L5) are obtained from the luminance grid. The six luminance samples closest to the chromaticity position C are selected without downsampling. The predicted chromaticity values ​​are obtained as follows:

[0175] Where nonlinear is the nonlinear operator of CCCM, and B is the bias.

[0176] The filter coefficients were derived using the division-free Gaussian elimination method of ECM, and the necessary offset was applied to the samples before the filter derivation.

[0177] When a block has fewer than 64 chroma samples, intra-frame reference samples are used as additional input samples in filter derivation. A CCCM design with up to 6 rows and columns of intra-frame reference samples is used.

[0178] A block with 256 or more chromaticity samples is divided into sub-blocks with a maximum of 256 chromaticity samples each. Sub-blocks containing zero luminance residuals are skipped.

[0179] 2.2.3 Bitstream Signaling The use of this mode is signaled by a TU level flag that is CABAC coded. A new CABAC context is included to support this. The CCRM flag is only signaled when the luma Cbf of the TU is not zero and the predMode of the CU is either MODE INTER or MODE IBC.

[0180] 3 Problem There are several problems in the existing video coding technology, which will be further improved for higher coding gain.

[0181] 1. Several aspects of CCRM coded video units (such as filter terms, model type, applied block type) can be further improved.

[0182] 4 Detailed solutions The following detailed solutions should be considered as examples to explain the general concepts. The solutions should not be interpreted in a narrow way. Furthermore, the solutions can be combined in any way.

[0183] The term "video unit" or "coding unit" can mean picture, slice, tile, coding tree block (CTB), coding tree unit (CTU), coding block (CB), CU, PU, TU, PB, TB.

[0184] The term "block" can mean coding tree block (CTB), coding tree unit (CTU), coding block (CB), CU, PU, TU, PB, TB.

[0185] The term "motion vector" or "block vector" can refer to a vector of horizontal and vertical displacement between the position of a reference block and the position of a current block. The reference block can be a video unit in a reference picture in the RPL list. The reference block can also be a video unit in the current picture.

[0186] The term "LM" can refer to any linear regression based method, such as CCLM, MMLM, CCCM, GL-CCCM, non-downsampled CCCM, GLM, GLM with luma values, etc. It can also be referred to as the term "cross component prediction (CCP)".

[0187] The term "CCLM" can refer to a single model LM mode, which can be single model CCLM, single model CCCM, single model GL-CCCM, non-downsampled single model CCCM, single model GLM, single model GLM with luma values, multi model CCLM, MMLM, multi model CCCM, multi model GL-CCCM, non-downsampled multi model CCCM, multi model GLM, multi model GLM with luma values, etc.

[0188] The term "MMLM" can refer to a multi-model LM mode, which can be multi-model CCLM, MMLM, multi-model CCCM, multi-model GL-CCCM, non-down-sampled multi-model CCCM, multi-model GLM, multi-model GLM with luma values, etc.

[0189] The term "CCCM" can refer to a regular CCCM mode, or GL-CCCM mode, or non-down-sampled CCCM, CCRM, etc.

[0190] The term "GL-CCCM" can refer to a CCCM mode that takes into account the gradient and location of the involved samples.

[0191] The term "non-down-sampled CCCM" can refer to a CCCM mode that takes into account non-down-sampled luma samples.

[0192] The term "CCRM" can refer to a residual coding or derivation based on a cross-component model. It can also imply an inter / IBC prediction based on a CCCM model, such as inter / IBC CCCM. It can also imply an intra prediction based on a CCCM model, such as intra CCCM.

[0193] In this document, cross-component prediction (CCP) can refer to any cross-component prediction method, such as any kind of CCLM / CCCM / GLM / GL-CCCM.

[0194] Note that the following terms are not limited to the specific terms defined in the existing standards. Any changes to the coding tools are applicable.

[0195] 1) The residual (and / or prediction) of a chroma block can be derived based on a cross-component model.

[0196] a. For example, the cross-component model can be a specific extrapolation filter (e.g., EIP, etc.).

[0197] b. For example, the cross-component model can be a specific interpolation filter (e.g., GLM, etc.).

[0198] c. For example, the cross-component model can be a specific convolution filter (CCCM, GL-CCCM, non-down-sampled CCCM, CCRM, inter CCCM, intra CCCM, etc.).

[0199] d. For example, the cross-component model can be a specific linear filter (e.g., CCLM, MMLM, etc.).

[0200] 2) The cross-component model (e.g., CCRM) used for residual coding can not contain a non-linear term.

[0201] a. For example, the cross-component model for residual coding can contain linear terms and / or bias terms, but not non-linear terms.

[0202] 3) CCRM can be used for intra or IBC blocks.

[0203] a. For example, it can be used for intra or IBC blocks in intra (such as I) slices.

[0204] b. For example, it can be used for intra or IBC blocks in inter (such as B or P) slices.

[0205] c. For example, in addition, it can be used for single tree.

[0206] d. For example, in addition, it can be used for dual tree.

[0207] e. For example, in single tree I slices, both luma and chroma are coded by IBC (or IntraTMP), CCRM can be generated based on the reconstructed luma and chroma samples within the reference block retrieved / guided by the block vector, and a residual model is applied to estimate the reconstructed values of chroma samples in the current block.

[0208] f. For example, in dual tree, luma is coded by IBC (or IntraTMP) while chroma is coded by intra, CCRM can be generated based on the reconstructed luma samples within the reference luma block retrieved / guided by the block vector, and the reconstructed chroma samples co-located (e.g., in the same position) with the luma block, and a residual model is applied to estimate the reconstructed values of chroma samples in the current block.

[0209] 4) CCRM can be used for DBV coded chroma blocks.

[0210] a. For example, based on the block vector of a DBV coded chroma block, a reference chroma block and its co-located luma block can be identified. These samples can be used as training samples for CCRM model computation.

[0211] b. For example, the derived CCRM model is applied to the reconstructed luma signal of the DBV chroma block to produce the final chroma prediction.

[0212] 5) CCRM model can be generated based on the correlation between luma reconstructed values and chroma reconstructed values from neighboring samples adjacent / non-adjacent to the current block.

[0213] a. For example, alternatively, CCRM model can be generated based on the correlation between luma reconstructed values and chroma reconstructed values in a reference block in the reference picture.

[0214] b. For example, alternatively, the CCRM model can be generated based on the correlation between the luma and chroma reconstructed values in the reference block in the current picture.

[0215] 6) For example, the CCCM for intra prediction and the CCCM (e.g., CCRM) for inter prediction can share the same logic.

[0216] a. For example, both of them can follow the same logic to obtain the training samples.

[0217] b. For example, both of them can follow the same logic to determine the training region.

[0218] 7) The CCRM model can be generated based on non-downsampled luma samples.

[0219] a. For example, the CCRM model coefficients can be solved based on the non-downsampled luma samples of the reference region as training samples.

[0220] b. For example, the CCRM model can be applied to a chroma block, where the chroma prediction of the current chroma block is generated based on the non-downsampled luma samples of the co-located luma block.

[0221] 8) More than one CCRM model (e.g., MM-CCRM) for a block can be generated.

[0222] a. For example, the training samples of the CCRM can be divided into more than one category (e.g., two categories), and each set of samples can contribute to a unique model. In this way, multiple models can be generated, each with its own filter coefficients. Each derived filter is applied to its corresponding set of luma reconstructed signals to produce the final predicted values of the current chroma samples belonging to the corresponding category.

[0223] i. For example, according to the multi-model CCRM (e.g., MM-CCRM) mode, the training sample pairs (e.g., these training samples are in the reference frame) in terms of luma and chroma samples of the reference block can be divided into more than one category.

[0224] ii. For example, alternatively, according to the multi-model CCRM (e.g., MM-CCRM) mode, the training sample pairs (e.g., these training samples are in the reference frame) in terms of luma and chroma samples of the neighboring samples adjacent / non-adjacent to the reference block can be divided into more than one category.

[0225] iii. For example, alternatively, pairs of training samples (e.g., these training samples are in the current frame) of luminance and chrominance samples pairs of neighboring samples adjacent / non-adjacent to the current video unit can be classified into more than one class according to a multi-model CCRM (e.g., MM-CCRM) mode.

[0226] iv. For example, in addition, luminance samples in the current video unit are classified into more than one group by following the same criteria (e.g., by a threshold), and for the luminance samples belonging to each class, a corresponding model can be applied to generate model estimated chrominance samples belonging to that group.

[0227] b. For example, multiple groups of training samples can be utilized to derive multiple models.

[0228] i. In one example, there are two groups, where the distance between the training samples and the current samples is different.

[0229] c. For example, the threshold (e.g., class threshold) to separate the samples into different classes can depend on the values of the samples within or neighboring the training region.

[0230] i. For example, the training region can be a reference block of the current video unit (e.g., these training samples are in the reference frame).

[0231] 1. For example, the reference block can be derived based on a block vector.

[0232] 2. For example, the reference block can be derived based on a motion vector.

[0233] ii. For example, the threshold can be derived based on samples adjacent / non-adjacent to a reference block of the current video unit (e.g., these training samples are in the reference frame).

[0234] iii. For example, the threshold can be derived based on samples adjacent / non-adjacent to the current video unit (e.g., these training samples are in the current frame).

[0235] iv. For example, the threshold can be derived based on an average / median / intermediate operation on more than one sample within or neighboring the training region.

[0236] v. For example, the class threshold can be derived based on non-downsampled luminance sample values.

[0237] 1. Alternatively, the class threshold can be derived based on downsampled luminance sample values.

[0238] a. For example, a K-tap (such as K=6) downsampling filter can be used to downsize K surrounding luminance samples to one downsampled luminance sample value.

[0239] vi. For example, the category threshold can be derived based on an offset removal scheme.

[0240] 1. For example, the offset can be derived based on luma samples located at a fixed position (such as top-left or center, etc.) in the reference video unit.

[0241] 2. For example, the offset value for category threshold derivation and CCRM model calculation can be the same.

[0242] vii. For example, the category threshold can be derived based on sub-block level.

[0243] viii. For example, the category threshold can be derived based on CU / PU / TU level.

[0244] ix. For example, the category threshold can be calculated based on (down-sampled or non-down-sampled) luma prediction samples.

[0245] x. For example, the category threshold can be derived based on luma residual sample values.

[0246] 1. For example, for a second video unit (e.g., sub-block) that does not have non-zero residual, the prediction samples of such video unit can not be counted in the calculation process of the category threshold for the first video unit.

[0247] a. For example, the second video unit can be a subset of the first video unit.

[0248] b. For example, the second video unit can be equal to the first video unit.

[0249] d. For example, MM-CCRM can be applied based on sub-block level.

[0250] i. For example, the sub-block size can be pre-defined.

[0251] 1. For example, the pre-defined sub-block size can be 16x16, or 32x32, etc.

[0252] 2. For example, a pre-defined rule can be used to determine the sub-block size for MM-CCRM of a particular video unit.

[0253] a. For example, the sub-block size can be adapted to the block dimension (width and / or height) of the current video block.

[0254] b. For example, for the sub-block of a video unit coded with MM-CCRM, a minimum number of chroma samples can be guaranteed.

[0255] ii. For example, if a video unit is larger than a predefined sub-block size, the video unit can be divided into more than one sub-block and MM-CCRM is performed.

[0256] iii. For example, at least one sub-block of a video unit can have more than one CCRM model.

[0257] iv. For example, each sub-block (and its associated training region) can have its own class threshold.

[0258] 1. For example, the class threshold for a particular sub-block can be computed based on the training sample values belonging to that sub-block.

[0259] a. For example, the luma training samples in the reference block can be used to compute the class threshold.

[0260] v. For example, all sub-blocks (and their associated training regions) can share the same class threshold.

[0261] 1. For example, one class threshold can be computed and used for all sub-blocks.

[0262] 2. For example, the class threshold for all sub-blocks in the current video unit can be computed based on the training sample values of the current video unit.

[0263] 3. For example, the class threshold for all applicable sub-blocks in the current video unit can be computed based on the training sample values of the current video unit.

[0264] a. For example, sub-blocks that do not contain non-zero residuals can not be counted.

[0265] vi. For example, each sub-block of a video unit can have its own training samples, and the training samples of a particular sub-block can be divided into more than one class.

[0266] 1. For example, the training samples in the reference video unit of the reference picture can be classified based on sub-blocks.

[0267] vii. For example, the training samples from the current picture can be classified into more than one group, but can not be divided into sub-blocks.

[0268] e. For example, MM-CCRM can be applied on a TU-level (or PU / CU-level).

[0269] i. For example, MM-CCRM can be applied on a TU / CU / PU basis (e.g., for the application of MM-CCRM, the TU / CU / PU can not be divided into sub-blocks).

[0270] ii. For example, whether to use a TU / PU / CU based multi-model CCRM (i.e., MM-CCRM) can be determined at the TU / PU / CU level.

[0271] 1. For example, a video unit (e.g., TU / PU / CU) can choose to use a subblock based single model CCRM or a TU / PU / CU based MM-CCRM.

[0272] a. For example, the decision can be made at the TU / PU / CU level.

[0273] f. For example, whether and / or how to apply MM-CCRM (and / or CCRM) can be derived based on coding information at both the encoder side and the decoder side (e.g., without the need for signaling).

[0274] i. In one example, it can be derived on-the-fly, e.g., using information of previously coded / reconstructed samples.

[0275] ii. For example, the determination of whether to use a subblock based CCRM or a TU / CU / PU level CCRM can be implicitly derived based on coding information (e.g., without the need for signaling).

[0276] iii. For example, the determination of whether to use a M1xM2 subblock based CCRM or a N1xN2 subblock based CCRM can be implicitly derived based on coding information (e.g., without the need for signaling).

[0277] 1. For example, M1 = 16 or 8 or 32 or TU / CU / PU.

[0278] 2. For example, M2 = 16 or 8 or 32 or TU / CU / PU.

[0279] 3. For example, N1 = 16 or 8 or 32 or TU / CU / PU.

[0280] 4. For example, N2 = 16 or 8 or 32 or TU / CU / PU.

[0281] 5. For example, M1!= N1 and / or M2!= N2.

[0282] iv. For example, the determination of whether to use a subblock based MM-CCRM or a TU / CU / PU level MM-CCRM can be implicitly derived based on coding information (e.g., without the need for signaling).

[0283] v. For example, the determination whether to use MM-CCRM based on M1xM2 sub-blocks or MM-CCRM based on N1xN2 sub-blocks can be implicitly derived based on coding information (e.g., without signaling).

[0284] 1. For example, M1 = 16 or 8 or 32 or TU / CU / PU.

[0285] 2. For example, M2 = 16 or 8 or 32 or TU / CU / PU.

[0286] 3. For example, N1 = 16 or 8 or 32 or TU / CU / PU.

[0287] 4. For example, N2 = 16 or 8 or 32 or TU / CU / PU.

[0288] 5. For example, M1!= N1 and / or M2!= N2.

[0289] vi. For example, the determination whether to use single model CCRM or MM-CCRM can be implicitly derived based on coding information (e.g., without signaling).

[0290] vii. For example, the determination can be according to a method based on decoder-derived cost.

[0291] 1. For example, the decoder-derived cost can be calculated based on minimizing SAD / SATD / SSE / MSE between model-estimated sample values and true reconstructed sample values, where the sample can refer to at least one of the training samples.

[0292] 2. For example, the method with lower cost can be selected as the final method applied to the current video unit.

[0293] viii. For example, the determination can be based on information of reference pictures.

[0294] 1. For example, the determination can be based on POC distance of the current picture and its reference pictures.

[0295] 2. For example, the determination can be based on reference index.

[0296] g. For example, whether and / or how to apply MM-CCRM (and / or CCRM) can be signaled in the bitstream.

[0297] i. For example, syntax elements (e.g., flags, indices, etc.) can be signaled based on the condition whether the current block is CCRM-coded.

[0298] 1. For example, if the video unit is CCRM coded, a syntax element (e.g., flag, index, etc.) can be further signaled to indicate whether it is MM-CCRM or not.

[0299] ii. For example, a syntax element (e.g., flag, index, etc.) can be signaled to indicate whether it is subblock-based MM-CCRM or TU / CP / PU-based MM-CCRM.

[0300] iii. For example, a syntax element (e.g., flag, index, etc.) can be signaled to indicate whether it is subblock-based CCRM or TU / CP / PU-based CCRM.

[0301] iv. For example, a syntax element can be signaled based on block dimension (width and / or height).

[0302] h. For example, a block restriction can be applied to indicate the allowance of MM-CCRM mode.

[0303] i. In one example, assuming the width and height of a chroma CU / PU / TU are denoted as W and H, MM-CCRM can be allowed when at least one of the following conditions is satisfied: 1. W H>T0 or W H>= T0 (e.g., T0 = 16 or 32 or 64 or 128) 2. W>T1, or, W>= T1 3. H>T2, or, H>= T2 4. Min (W,H)>T3, or, Min (W,H)>= T3 5. Max (W,H)<T4, or, Max (W,H)<= T4 6. W<T5 H, or, W<= T5 H 7. W>T6 H, or, W>= T6 H 8. H<T7 W, or, H<= T7 W 9. H>T8 W, or, H>= T8 W 10. W H<T9, or W H<= T9 ii. In one example, the MM-CCRM can be disabled for blocks enabled for a particular tool (e.g., affine motion compensation enabled).

[0304] i. For example, CCRM-coded video units can always use multi-model CCRM.

[0305] i. Alternatively, CCRM-coded video units can use single-model CCRM or multi-model CCRM.

[0306] 9) Chroma Cb and Cr can share one CCRM.

[0307] a. Alternatively, chroma Cb and Cr can build their own CCRM.

[0308] 10) For CCRM model filter design, sample values and / or gradients and / or position information can be considered.

[0309] a. For example, at least one K-tap filter can be used for CCRM model, which consists of K1 sample terms, K2 gradient terms, K3 positioning / position terms, K4 non-linear terms, K5 bias terms, etc.

[0310] i. For example, K1 = 0 or 1 or 2 or 5 or 6 ii. For example, K2 = 0 or 1 or 2 or 4 iii. For example, K3 = 0 or 1 or 2 or 4 iv. For example, K4 = 0 or 1 or 2 or 4 v. For example, K5 = 0 or 1 vi. For example, K = K1 + K2 + K3 + K4 + K5 vii. For example, sample terms can be calculated based on luma sample values.

[0311] viii. For example, gradient terms can be calculated based on more than one sample neighboring a particular luma sample.

[0312] ix. For example, positioning / position terms can be calculated based on horizontal and / or vertical coordinates of a particular luma sample, where the coordinates can be relative to the top-left position of a particular reference region.

[0313] x. For example, non-linear terms can be square of a particular value (e.g., bit-depth dependent intermediate value such as 512 or 256, or a particular luma value).

[0314] xi. For example, non-linear terms can be square of a gradient value based on a particular gradient term.

[0315] xii. For example, the offset can be subtracted from the term of the K-tap filter.

[0316] 1. For example, the offset can be derived based on a predefined rule, such as the value of the top-left training sample in the training region, or the average / median value of more than one sample in the training region.

[0317] xiii. For example, the coefficients of the K-tap filter can be solved by a Gaussian elimination solver.

[0318] xiv. For example, the coefficients of the K-tap filter can be solved by an LDL decomposition method.

[0319] i. For example, the coefficients of the K-tap filter can be solved by linear regression.

[0320] ii. For example, the coefficients of the K-tap filter can be solved by linear equations.

[0321] b. For example, more than one filter can be used, and the final prediction can be derived based on fusing the filtered outputs of the multiple filters together.

[0322] i. For example, the weights of fusing the multiple filtered values can be solved by a Gaussian elimination solver.

[0323] ii. For example, the weights of fusing the multiple filtered values can be solved by an LDL decomposition method.

[0324] 11) For example, for a video unit coded by CCRM, more than one filter can be allowed, and which filter is finally selected can be signaled or partitioned.

[0325] a. For example, a syntax element can be signaled to indicate which filter (e.g., CCLM or CCCM) is used for the CCRM mode.

[0326] b. For example, indicating which filter (e.g., CCLM or CCCM) is used for the CCRM mode can be determined based on the template cost from both the encoder and the decoder.

[0327] c. For example, indicating which filter (e.g., CCLM or CCCM) is used for the CCRM mode can be determined based on the decoder-derived cost from both the encoder and the decoder.

[0328] 12) The filter output can be clipped to a certain value.

[0329] a. For example, it can be clipped based on the reconstructed value in the training region.

[0330] i. For example, the training region can be derived based on the block vector (or motion vector).

[0331] ii. For example, the training region can be adjacent to the current block.

[0332] iii. For example, the training region can be the reference region of the current block.

[0333] iv. For example, the filter output can be clipped within the minimum and maximum of the reconstructed (or predicted) luma sample values in the training region.

[0334] b. For example, it can be clipped based on the reconstructed (or predicted) values in the co-located luma block of the current chroma block.

[0335] i. For example, it can be clipped within the minimum and maximum of the current block luma reconstructed (or predicted) values.

[0336] c. For example, if the value is outside the valid range, it can be ignored / dropped / not used.

[0337] 13) The CCRM parameters can be stored in a cache and used for coding of future blocks.

[0338] a. For example, the CCRM parameters for a video unit (e.g., CU, PU, color component, Cb, Cr, etc.) can include the model type, model coefficients, whether it is a single model or multiple models, the threshold to separate samples into multiple models, etc.

[0339] b. For example, it can be stored in a local cache for coding of future blocks in the current picture.

[0340] c. For example, it can be stored in a temporal / picture / frame cache for coding of future blocks in future decoded pictures.

[0341] i. For example, the CCRM parameters of the current frame / picture can be stored, which can be referenced for the CCP process of future frames / pictures.

[0342] ii. For example, it can be stored in association with the motion and mode information of the video unit.

[0343] 14) A video block can inherit the model parameters from a previous CCRM coded block.

[0344] a. For example, a video block can be coded by a CCP inheritance mode.

[0345] b. For example, a video block can be coded by a CCP Merge (e.g., CCMerge) mode.

[0346] c. For example, model parameters of previous CCRM coded blocks can be stored in a cache (e.g., local cache, picture cache, temporal cache, history based LUT, etc.) d. In one example, parameters can refer to filter information, linear or non-linear parameters of the model, model index, etc.

[0347] 15) The final prediction of a block can be generated based on multiple prediction candidates from different CCRMs.

[0348] a. For example, more than one CCRM prediction can be fused together.

[0349] b. For example, weights / coefficients of different fusion terms can be solved based on a Gaussian elimination method.

[0350] c. For example, weights / coefficients of different fusion terms can be solved based on an LDL decomposition method.

[0351] d. For example, bias terms can be involved for fusion.

[0352] e. For example, non-linear terms can be involved for fusion.

[0353] 16) The allowance of CCRM mode can depend on at least one of the following: a. Prediction mode of the video unit (e.g., MODE INTRA, MODE INTER, MODE IBC, MODE PLT, etc.) b. Transform type of the video unit (e.g., ACT, color transform, etc.) c. Number of non-zero coefficients of the video unit d. Partition tree type (e.g., single tree, dual tree) e. Slice type (e.g., I, B, P slice) f. Color format (e.g., whether 4:0:0) g. Availability of chroma components h. For example, CCRM can not be allowed for ACT and / or 4:0:0 color format.

[0354] 17) The disclosed CCRM mode can be based on one of the following filters: a. CCLM and / or its variants b. MMLM and / or its variants c. CCCM and / or its variants (e.g., GL-CCCM, non-down-sampled CCCM, BVG-CCCM, inter-CCCM, intra-CCCM, etc.) d. GLM and / or its variants e. any cross-component prediction that uses information in one channel / component to predict information in another channel / component f. any filter-based prediction, where filter coefficients are solved based on the correlation between prediction and / or reconstructed information 18) Block restrictions can be applied to limit the application of certain types of CCP modes.

[0355] a. For example, CCP modes can only be allowed to be used if the block size satisfies a pre-defined rule.

[0356] b. For example, a syntax element can only be signaled if a CCP mode is applicable.

[0357] c. For example, if a CCP mode is not allowed to be used, a syntax element can be inferred to a certain value that indicates that no such CCP mode is used for such a block.

[0358] d. For example, at least one of the following block restrictions can be applied to CCRM modes (assuming W denotes the block width and H denotes the block height): i. W < T1, or, W <= T1 ii. H < T2, or, H <= T2 iii. Min (W, H) > T3, or, Min (W, H) >= T3 iv. Max (W, H) < T4, or, Max (W, H) <= T4 v. W < T5 H, or, W <= T5 H vi. W > T6 H, or, W >= T6 H vii. H < T7 W, or, H <= T7 W viii. H > T8 W, or, H >= T8 W ix. W H < T9, or W H <= T9 x. For example, T1, T2, … T9 can be pre-defined integer constants.

[0359] e. For example, CCRM modes can only be allowed to be used for small blocks.

[0360] i. For example, it can be allowed for blocks smaller than 4x4, or 8x8, or 16x16, or 32x32.

[0361] ii. For example, it can be allowed for blocks with a number of samples smaller than 32, or 64, or 128.

[0362] iii. For example, it can be allowed for blocks with a number of samples smaller than 32, or 64, or 128.

[0363] iv. For example, it can not be allowed for 2xN blocks, where N can be larger than 4 or 8 or 16.

[0364] v. For example, it can not be allowed for Nx2 blocks, where N can be larger than 4 or 8 or 16.

[0365] 19) The disclosed method can be used in single tree.

[0366] 20) The disclosed method can be used in dual tree.

[0367] 21) The disclosed method can be used in inter (such as B or P) slices.

[0368] 22) The disclosed method can be used in intra (such as I) slices.

[0369] 23) The "block vector" in the disclosed method can be a "motion vector".

[0370] 24) The training / reference sample in the disclosed method can refer to a predicted sample and / or a reconstructed sample in the training / reference region.

[0371] 25) Whether and / or how to apply the disclosed method above can be signaled at sequence level / group of pictures level / picture level / slice level / tile group level, such as in sequence header / picture header / S PS / VPS / DPS / DCI / PPS / APS / slice header / tile group header.

[0372] 26) Whether and / or how to apply the disclosed method above can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / tile / subpicture / other kind of region containing more than one sample or pixel.

[0373] 27) Whether and / or how to apply the disclosed method above can depend on coded information, such as block size, color format, single / dual tree partitioning, color component, slice / picture type.

[0374] Figure 27A flowchart illustrating a method 2700 for video processing according to embodiments of the present disclosure is shown. The method 2700 is implemented during conversion between a video unit of a video and a bitstream of the video.

[0375] At block 2710, for conversion between a video unit of a video and a bitstream of the video, one or more cross-component residual models (CCRM) used for the video unit are determined. In some embodiments, the one or more CCRMs include a multi-mode CCRM (MM-CCRM).

[0376] At block 2720, the conversion is performed based on the one or more CCRMs. In some embodiments, the conversion includes encoding the video unit into the bitstream. Alternatively, the conversion includes decoding the video unit from the bitstream.

[0377] In some embodiments, the training sample pairs of the one or more CCRMs are partitioned into multiple categories, and the sample pairs of each category are applied to a unique model. In some other embodiments, according to the MM-CCRM mode, the training sample pairs of the luma and chroma sample pairs of a reference block are partitioned into multiple categories.

[0378] In some embodiments, according to the MM-CCRM mode, the training sample pairs of the luma and chroma sample pairs of neighboring samples adjacent to the reference block are partitioned into multiple categories. Alternatively or additionally, according to the MM-CCRM mode, the training sample pairs of the luma and chroma sample pairs of neighboring samples non-adjacent to the reference block are partitioned into multiple categories.

[0379] In some embodiments, according to the MM-CCRM mode, the training sample pairs of the luma and chroma sample pairs of neighboring samples adjacent to the video unit are partitioned into multiple categories. Alternatively or additionally, according to the MM-CCRM mode, the training sample pairs of the luma and chroma sample pairs of neighboring samples non-adjacent to the video unit are partitioned into multiple categories.

[0380] In some embodiments, the luma samples in the video unit are partitioned into multiple groups according to the same criteria, and for the luma samples belonging to each category, a corresponding model is applied to generate the model estimated chroma samples belonging to the group. In some embodiments, multiple models are derived using multiple groups of training samples. In some embodiments, there are two groups of training samples, where the distance between the training samples and the current sample is different.

[0381] In some embodiments, the threshold used to separate the samples into different categories depends on the value of one or more samples within the training region. Alternatively, the threshold used to separate the samples into different categories depends on the value of one or more samples adjacent to the training region. In some embodiments, the threshold is a category threshold.

[0382] In some embodiments, the training samples are in the reference frame. In some embodiments, the threshold is derived based on training samples neighboring the reference block of the video unit. Alternatively or additionally, the threshold is derived based on training samples not neighboring the reference block of the video unit.

[0383] In some embodiments, the training samples are in the reference frame. In some embodiments, the threshold is derived based on training samples neighboring the reference block of the video unit. Alternatively or additionally, the threshold is derived based on training samples not neighboring the reference block of the video unit.

[0384] In some embodiments, the training samples are in the current frame. In some embodiments, the threshold is derived based on at least one of the following: an average operation, a median operation, or an intermediate operation on a plurality of samples within or neighboring the training region.

[0385] In some embodiments, the category threshold is derived based on a non-downsampled luma sample value. In some embodiments, the category threshold is derived based on a downsampled luma sample value. For example, a K-tap downsampling filter is used to downsize K surrounding luma samples to one downsampled luma sample value, where K is an integer. In some embodiments, K is 6.

[0386] In some embodiments, the category threshold is derived based on an offset removal scheme. In some embodiments, the offset is derived based on a luma sample located at a fixed position in the reference video unit. In some embodiments, the fixed position is one of the following: top-left or center. In some embodiments, the offset value is the same for both the category threshold derivation and the CCRM model calculation.

[0387] In some embodiments, the category threshold is derived at a sub-block level. In some embodiments, the category threshold is derived at one of the following: a coding unit (CU) level, a prediction unit (PU) level, or a transform unit (TU) level. In some embodiments, the category threshold is calculated based on a luma prediction sample. In some embodiments, the luma prediction sample is downsampled. Alternatively, the luma prediction sample is non-downsampled.

[0388] In some embodiments, the category threshold is derived based on a luma residual sample value. In some embodiments, for a second video unit that does not have a non-zero residual, the prediction sample of the second video unit is not counted in the calculation of the category threshold for a first video unit. In some embodiments, the second video unit is a subset of the first video unit. Alternatively, the second video unit is equal to the first video unit, and / or the second video unit is a sub-block.

[0389] In some embodiments, MM-CCRM is applied on a sub-block level. In some embodiments, the sub-block size is predefined. In some embodiments, the predefined sub-block block size is 16x16 or 32x32.

[0390] In some embodiments, a predefined rule is used to determine the sub-block size of MM-CCRM for a target video unit. In some embodiments, the sub-block block size is adapted to the block dimension of the current video block. In some embodiments, the block dimension includes at least one of width or height. In some embodiments, a minimum number of chroma samples is guaranteed for the sub-blocks of the MM-CCRM coded video unit.

[0391] In some embodiments, if a video unit is larger than a predefined sub-block size, the video unit is divided into multiple sub-blocks and MM-CCRM is applied. In some embodiments, at least one sub-block of the video unit has multiple CCRM models.

[0392] In some embodiments, each sub-block and / or the associated training region of each sub-block has a class threshold. In some embodiments, the class threshold of a target sub-block is calculated based on the training sample values belonging to the target sub-block. In some embodiments, the luma training samples in the reference block are used to calculate the class threshold.

[0393] In some embodiments, all sub-blocks and / or the associated training regions of all sub-blocks share the same class threshold. In some embodiments, one class threshold is calculated and used for all sub-blocks. In some embodiments, the class threshold of all sub-blocks in a video unit is calculated based on the training sample values of the video unit. In some embodiments, the class threshold of all applicable sub-blocks in a video unit is calculated based on the training sample values of the video unit.

[0394] In some embodiments, sub-blocks that do not contain non-zero residuals are not counted. In some embodiments, each sub-block of a video unit has its own training samples and the training samples of a target sub-block are divided into multiple classes.

[0395] In some embodiments, the training samples in the reference video unit of the reference picture are classified based on sub-blocks. In some other embodiments, the training samples of the current picture are classified into multiple groups but not divided into sub-blocks.

[0396] In some embodiments, the MM-CCRM is applied on one of: TU basis, CU basis, or PU basis. In some embodiments, the MM-CCRM is applied on one of: TU basis, CU basis, or PU basis. In some embodiments, for application of the MM-CCRM, the TU is not divided into sub-blocks. Alternatively or additionally, for application of the MM-CCRM, the CU is not divided into sub-blocks. Alternatively or additionally, for application of the MM-CCRM, the PU is not divided into sub-blocks.

[0397] In some embodiments, whether to use one of: TU-based multi-model CCRM, PU-based multi-model CCRM, or CU-based multi-model CCRM is determined at one of: TU level, PU level, or CU level. In some embodiments, the video unit decides to use sub-block-based single model CCRM or one of: TU-based MM-CCRM, PU-based MM-CCRM, or CU-based MM-CCRM. In some embodiments, the decision is made at one of: TU level, PU level, or CU level.

[0398] In some embodiments, whether and / or how to apply at least one of: MM-CCRM or CCRM is derived based on coding information at both encoder side and decoder side. In some other embodiments, whether and / or how to apply at least one of: MM-CCRM or CCRM is derived on-the-fly.

[0399] In some embodiments, whether and / or how to apply at least one of: MM-CCRM or CCRM is derived using information of previously coded or reconstructed samples. In some embodiments, the determination of whether to use sub-block-based CCRM or one of: TU-level CCRM, CU-level CCRM, or PU-level CCRM is implicitly derived based on coding information.

[0400] In some embodiments, the determination of whether to use M1xM2 sub-block-based CCRM or N1xN2 sub-block-based CCRM is implicitly derived based on coding information. In some embodiments, M1 = 16 or 8 or 32 or TU or CU or PU. Alternatively or additionally, M2 = 16 or 8 or 32 or TU or CU or PU. Alternatively or additionally, N1 = 16 or 8 or 32 or TU or CU or PU. Alternatively or additionally, N2 = 16 or 8 or 32 or TU or CU or PU. In some embodiments, M1 is not equal to N1 and / or M2 is not equal to N2.

[0401] In some embodiments, the determination of whether to use MM-CCRM based on sub-blocks or one of the following is derived based on coding information: TU-level MM-CCRM, CU-level MM-CCRM, or PU-level MM-CCRM. In some embodiments, the determination of whether to use MM-CCRM based on M1xM2 sub-blocks or MM-CCRM based on N1xN2 sub-blocks is derived implicitly based on coding information.

[0402] In some embodiments, M1 = 16 or 8 or 32 or TU or CU or PU. Alternatively or additionally, M2 = 16 or 8 or 32 or TU or CU or PU. Alternatively or additionally, N1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N2 = 16 or 8 or 32 or TU or CU or PU. Alternatively or additionally, M1 is not equal to N1 and / or M2 is not equal to N2.

[0403] In some embodiments, the determination of whether to use single model CCRM or MM-CCRM is derived based on coding information. In some other embodiments, the determination of whether and / or how to apply at least one of MM-CCRM or CCRM is according to a scheme based on decoder-derived cost.

[0404] In some embodiments, the decoder-derived cost is computed based on minimizing one of the following: sum of absolute difference (SAD), sum of absolute transformed difference (SATD), sum of squared error (SSE), or mean squared error (MSE) between model-estimated sample values of samples and true reconstructed sample values of the samples, wherein the samples include at least one of the training samples. In some embodiments, the scheme based on decoder-derived cost with lower cost is selected as the final scheme applied to the video unit.

[0405] In some embodiments, the determination of whether and / or how to apply at least one of MM-CCRM or CCRM is based on information of reference pictures. In some other embodiments, the determination is based on picture order count (POC) distance of the current picture and reference pictures of the current picture. In some further embodiments, the determination is based on reference index.

[0406] In some embodiments, whether and / or how to apply at least one of MM-CCRM or CCRM is signaled in the bitstream. In some embodiments, a syntax element is signaled based on whether the video unit is CCRM coded. In some embodiments, if the video unit is CCRM coded, a syntax element is further signaled to indicate whether it is MM-CCRM.

[0407] In some embodiments, a syntax element is signaled to indicate whether it is a subblock-based MM-CCRM or one of the following: a TU-based MM-CCRM, a CP-based MM-CCRM, or a PU-based MM-CCRM. In some other embodiments, a syntax element is signaled to indicate whether it is a subblock-based CCRM or one of the following: a TU-based CCRM, a CP-based CCRM, or a PU-based CCRM.

[0408] In some embodiments, the syntax element is signaled based on a condition regarding a block dimension. In some embodiments, the block dimension includes at least one of a width or a height.

[0409] In some embodiments, a block restriction is applied to indicate the allowance of MM-CCRM mode. In some embodiments, if the width and height of one of a chroma CU, a chroma PU, or a chroma TU is denoted as W and H, MM-CCRM is allowed when at least one of the following conditions is satisfied: H > T0 or W H >= T0; W > T1, or, W >= T1; H > T2, or, H >= T2; Min (W, H) > T3, or, Min (W, H) >= T3; Max (W, H) < T4, or, Max (W, H) <= T4; W < T5 H, or, W <= T5 H; W > T6 H, or, W >= T6 H; H < T7 W, or, H <= T7 W; H > T8 W, or, H >= T8 W; or W H < T9, or W H <= T9, and T0, T1, T2, T3, T4, T5, T6, T7, T8, and T9 are threshold parameters.

[0410] In some embodiments, MM-CCRM is disabled for blocks for which a tool is enabled. For example, if affine motion compensation is enabled, MM-CCRM is disabled.

[0411] In some embodiments, the CCRM-coded video unit uses multi-model CCRM. Alternatively, the CCRM-coded video unit uses single-model CCRM or multi-model CCRM.

[0412] In some embodiments, which filter is used for CCRM mode is indicated, which filter is used for CCRM mode is determined based on decoder-derived cost from both the encoder and the decoder. In some embodiments, the video unit inherits parameters of CCRM from a previous CCRM-coded block. In some embodiments, the parameters include at least one of the following: filter information, one or more linear parameters of the model, one or more non-linear parameters of the model, or a model index.

[0413] In some embodiments, the determination of one or more CCRM is used in at least one of the following: single tree or dual tree. In some other embodiments, the determination of one or more CCRM is used in inter- frame slices. In some embodiments, the inter-frame slices are B slices or P slices.

[0414] In some embodiments, the determination of one or more CCRM is used in intra-frame slices. In some embodiments, the intra-frame slices are I slices.

[0415] In some embodiments, the training samples or the reference samples are predicted samples in the training region or the reference region. In some other embodiments, the training samples or the reference samples are reconstructed samples in the training region or the reference region.

[0416] In some embodiments, an indication of whether and / or how one or more CCRM is determined for a video unit is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or tile group level.

[0417] In some embodiments, an indication of whether and / or how one or more CCRM is determined for a video unit is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependent parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, or tile group header.

[0418] In some embodiments, an indication of whether and / or how one or more CCRM is determined for a video unit is included in one of the following: prediction block (PB), transform block (TB), coding block (CB), prediction unit (PU), transform unit (TU), coding unit (CU), virtual pipeline data unit (VPDU), coding tree unit (CTU), CTU row, slice, tile, sub-picture, or a region containing more than one sample or pixel.

[0419] In some embodiments, the method [0AA-NUM]00 further includes determining, based on coded information of the video unit, whether to determine one or more CCRMs for the video unit and / or how to determine one or more CCRMs for the video unit, the coded information including at least one of: block size, color format, single tree partitioning and / or dual tree partitioning, color component, slice type, or picture type.

[0420] According to further embodiments of the disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes determining one or more cross-component residual models (CCRM) for a video unit of the video; and generating a bitstream of the video unit based on the one or more CCRM.

[0421] According to still further embodiments of the disclosure, a method for storing a bitstream of a video is provided. The method includes determining one or more cross-component residual models (CCRM) for a video unit of the video; generating a bitstream of the video unit based on the one or more CCRM; and storing the bitstream in a non-transitory computer-readable recording medium.

[0422] Implementations of the disclosure can be described according to the following clauses, which features can be combined in any reasonable manner.

[0423] Clause 1. A method for video processing, comprising: determining, for a conversion between a video unit of a video and a bitstream of the video, one or more cross-component residual models (CCRM) for the video unit; and performing the conversion based on the one or more CCRM.

[0424] Clause 2. The method of clause 1, wherein the one or more CCRM comprises a multi-mode CCRM (MM-CCRM).

[0425] Clause 3. The method of clause 1, wherein training samples of the one or more CCRM are divided into multiple categories, and the samples of each category are applied to a unique model.

[0426] Clause 4. The method of clause 3, wherein, according to a MM-CCRM mode, training sample pairs of luma and chroma samples of a reference block are divided into multiple categories.

[0427] Item 5. The method of item 3, wherein the training sample pairs of luma and chroma sample pairs of neighboring samples adjacent to a reference block are classified into a plurality of categories according to the MM-CCRM mode, and / or wherein the training sample pairs of luma and chroma sample pairs of neighboring samples non-adjacent to the reference block are classified into a plurality of categories according to the MM-CCRM mode.

[0428] Item 6. The method of item 3, wherein the training sample pairs of luma and chroma sample pairs of neighboring samples adjacent to the video unit are classified into a plurality of categories according to the MM-CCRM mode, and / or wherein the training sample pairs of luma and chroma sample pairs of neighboring samples non-adjacent to the video unit are classified into a plurality of categories according to the MM-CCRM mode.

[0429] Item 7. The method of item 3, wherein the luma samples in the video unit are classified into a plurality of groups according to the same criteria, and for luma samples belonging to each category, a corresponding model is applied to generate model estimated chroma samples belonging to the group.

[0430] Item 8. The method of item 1, wherein multiple sets of training samples are utilized to derive multiple models.

[0431] Item 9. The method of item 9, wherein there are two sets of training samples, wherein the distance between the training samples and the current sample is different.

[0432] Item 10. The method of item 1, wherein the threshold used to separate samples into different categories depends on the value of one or more samples within the training region, or wherein the threshold used to separate samples into different categories depends on the value of one or more samples neighboring the training region.

[0433] Item 11. The method of item 10, wherein the threshold is a category threshold.

[0434] Item 12. The method of item 10, wherein the training samples are in a reference frame.

[0435] Item 13. The method of item 10, wherein the threshold is derived based on training samples neighboring a reference block of the video unit, and / or wherein the threshold is derived based on training samples non-adjacent to the reference block of the video unit.

[0436] Item 14. The method of item 13, wherein the training samples are in a reference frame.

[0437] Item 15. The method of item 10, wherein the threshold is derived based on training samples that are adjacent to the video unit, and / or wherein the threshold is derived based on training samples that are not adjacent to the video unit.

[0438] Item 16. The method of item 15, wherein the training samples are in a current frame.

[0439] Item 17. The method of item 10, wherein the threshold is derived based on at least one of: an average operation, a median operation, or an intermediate operation on a plurality of samples within or adjacent to a training region.

[0440] Item 18. The method of item 10, wherein a category threshold is derived based on a non-downsampled luma sample value.

[0441] Item 19. The method of item 10, wherein a category threshold is derived based on a downsampled luma sample value.

[0442] Item 20. The method of item 19, wherein a K-tap downsampling filter is used to downsize K surrounding luma samples to one downsampled luma sample value, wherein K is an integer.

[0443] Item 21. The method of item 20, wherein K is 6.

[0444] Item 22. The method of item 10, wherein a category threshold is derived based on an offset removal scheme.

[0445] Item 23. The method of item 22, wherein an offset is derived based on a luma sample located at a fixed position in a reference video unit.

[0446] Item 24. The method of item 23, wherein the fixed position is one of: top-left or center.

[0447] Item 25. The method of item 22, wherein the offset value is the same for both category threshold derivation and CCRM model calculation.

[0448] Item 26. The method of item 10, wherein a category threshold is derived based on a subblock level.

[0449] Item 27. The method of item 10, wherein a category threshold is derived based on one of: a coding unit (CU) level, a prediction unit (PU) level, or a transform unit (TU) level.

[0450] Item 28. The method of item 10, wherein a category threshold is calculated based on a luma prediction sample.

[0451] Item 29. The method of item 28, wherein the luma prediction samples are down-sampled, or wherein the luma prediction samples are not down-sampled.

[0452] Item 30. The method of item 10, wherein a category threshold is derived based on luma residual sample values.

[0453] Item 31. The method of item 30, wherein for a second video unit having no non-zero residual, prediction samples of the second video unit are not counted in a process of calculating the category threshold for a first video unit.

[0454] Item 32. The method of item 31, wherein the second video unit is a subset of the first video unit, or wherein the second video unit is equal to the first video unit, and / or wherein the second video unit is a sub-block.

[0455] Item 33. The method of item 1, wherein MM-CCRM is applied on a sub-block level.

[0456] Item 34. The method of item 33, wherein a sub-block size is predefined.

[0457] Item 35. The method of item 34, wherein the predefined sub-block size is 16x16 or 32x32.

[0458] Item 36. The method of item 34, wherein a predefined rule is used to determine the sub-block size for MM-CCRM of a target video unit.

[0459] Item 37. The method of item 36, wherein the sub-block size is adapted to block dimensions of a current video block.

[0460] Item 38. The method of item 37, wherein the block dimensions include at least one of width or height.

[0461] Item 39. The method of item 36, wherein for a sub-block of a MM-CCRM coded video unit, a minimum number of chroma samples is ensured.

[0462] Item 40. The method of item 33, wherein if the video unit is larger than a predefined sub-block size, the video unit is divided into a plurality of sub-blocks, and the MM-CCRM is applied.

[0463] Item 41. The method of item 33, wherein at least one sub-block of the video unit has a plurality of CCRM models.

[0464] Item 42. The method of item 33, wherein each sub-block and / or each sub-block's associated training region has a class threshold.

[0465] Item 43. The method of item 42, wherein the class threshold for a target sub-block is computed based on training sample values belonging to the target sub-block.

[0466] Item 44. The method of item 43, wherein luma training samples in a reference block are used to compute the class threshold.

[0467] Item 45. The method of item 33, wherein all sub-blocks and / or all sub-blocks' associated training regions share the same class threshold.

[0468] Item 46. The method of item 45, wherein one class threshold is computed and used for all sub-blocks.

[0469] Item 47. The method of item 45, wherein the class threshold for all sub-blocks in the video unit is computed based on training sample values of the video unit.

[0470] Item 48. The method of item 45, wherein the class threshold for all applicable sub-blocks in the video unit is computed based on training sample values of the video unit.

[0471] Item 49. The method of item 48, wherein sub-blocks that do not contain non-zero residual are not counted.

[0472] Item 50. The method of item 33, wherein each sub-block of the video unit has its own training samples and the training samples of a target sub-block are partitioned into multiple classes.

[0473] Item 51. The method of item 50, wherein training samples in a reference video unit of a reference picture are classified based on sub-blocks.

[0474] Item 52. The method of item 33, wherein training samples of a current picture are classified into multiple groups but not partitioned into sub-blocks.

[0475] Item 53. The method of item 1, wherein MM-CCRM is applied on one of the following: TU level, PU level or CU level.

[0476] Item 54. The method of item 53, wherein the MM-CCRM is applied on one of the following: TU basis, CU basis or PU basis.

[0477] Item 55. The method of item 54, wherein, for application of the MM-CCRM, TUs are not divided into sub-blocks, and / or wherein, for application of the MM-CCRM, CUs are not divided into sub-blocks, and / or wherein, for application of the MM-CCRM, PUCs are not divided into sub-blocks.

[0478] Item 56. The method of item 53, wherein whether to use one of a TU-based multi-model CCRM, a PU-based multi-model CCRM, or a CU-based multi-model CCRM is determined at one of: a TU level, a PU level, or a CU level.

[0479] Item 57. The method of item 56, wherein the video unit decides to use a sub-block based single model CCRM or one of: a TU-based MM-CCRM, a PU-based MM-CCRM, or a CU-based MM-CCRM.

[0480] Item 58. The method of item 57, wherein the decision is made at one of: a TU level, a PU level, or a CU level.

[0481] Item 59. The method of item 1, wherein whether and / or how to apply at least one of a MM-CCRM or a CCRM is inferred based on coding information at both an encoder side and a decoder side.

[0482] Item 60. The method of item 59, wherein whether and / or how to apply at least one of a MM-CCRM or a CCRM is inferred on-the-fly.

[0483] Item 61. The method of item 60, wherein whether and / or how to apply at least one of a MM-CCRM or a CCRM is inferred using information of previously coded or reconstructed samples.

[0484] Item 62. The method of item 59, wherein a determination of whether to use a sub-block based CCRM or one of: a TU level CCRM, a CU level CCRM, or a PU level CCRM is implicitly inferred based on coding information.

[0485] Item 63. The method of item 59, wherein a determination of whether to use a MlxM2 sub-block based CCRM or a NlxN2 sub-block based CCRM is implicitly inferred based on coding information.

[0486] Item 64. The method of item 63, wherein M1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein M2 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N2 = 16 or 8 or 32 or TU or CU or PU, and / or wherein M1 is not equal to N1 and / or M2 is not equal to N2.

[0487] Item 65. The method of item 63, wherein M1 is not equal to N1 and / or M2 is not equal to N2.

[0488] Item 66. The method of item 60, wherein the determination of whether to use a subblock-based MM-CCRM or one of the following is derived based on coding information: TU-level MM-CCRM, CU-level MM-CCRM, or PU-level MM-CCRM.

[0489] Item 67. The method of item 60, wherein the determination of whether to use a MM-CCRM based on M1 x M2 subblocks or a MM-CCRM based on N1 x N2 subblocks is implicitly derived based on coding information.

[0490] Item 68. The method of item 67, wherein M1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein M2 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N2 = 16 or 8 or 32 or TU or CU or PU, and / or wherein M1 is not equal to N1 and / or M2 is not equal to N2.

[0491] Item 69. The method of item 60, wherein the determination of whether to use a single model CCRM or a MM-CCRM is derived based on coding information.

[0492] Item 70. The method of item 60, wherein the determination of whether and / or how to apply at least one of a MM-CCRM or a CCRM is according to a scheme based on decoder-derived cost.

[0493] Item 71. The method of item 70, wherein the decoder-derived cost is computed based on minimizing one of the following: sum of absolute differences (SAD) between model-estimated sample values of samples and true reconstructed sample values of the samples, sum of absolute transformed differences (SATD), sum of squared errors (SSE), or mean squared error (MSE), wherein the samples include at least one of the training samples.

[0494] Item 72. The method of item 70, wherein the scheme with lower cost based on decoder-derived cost is selected as the final scheme applied to the video unit.

[0495] Item 73. The method of item 60, wherein the determination of whether and / or how to apply at least one of MM-CCRM or CCRM is based on information of reference pictures.

[0496] Item 74. The method of item 73, wherein the determination is based on picture order count (POC) distance of a current picture and reference pictures of the current picture.

[0497] Item 75. The method of item 73, wherein the determination is based on reference index.

[0498] Item 76. The method of item 1, wherein whether and / or how to apply at least one of MM-CCRM or CCRM is signaled in the bitstream.

[0499] Item 77. The method of item 76, wherein a syntax element is signaled based on whether the video unit is CCRM coded.

[0500] Item 78. The method of item 77, wherein if the video unit is CCRM coded, a syntax element is further signaled to indicate whether it is MM-CCRM.

[0501] Item 79. The method of item 76, wherein a syntax element is signaled to indicate whether it is subblock-based MM-CCRM or one of the following: TU-based MM-CCRM, CP-based MM-CCRM, or PU-based MM-CCRM.

[0502] Item 80. The method of item 76, wherein a syntax element is signaled to indicate whether it is subblock-based CCRM or one of the following: TU-based CCRM, CP-based CCRM, or PU-based CCRM.

[0503] Item 81. The method of item 76, wherein a syntax element is signaled based on a condition on block dimension.

[0504] Item 82. The method of item 81, wherein the block dimension comprises at least one of width or height.

[0505] Item 83. The method of item 1, wherein a block restriction is applied to indicate allowance of MM-CCRM mode.

[0506] Item 84. The method of item 83, wherein the MM-CCRM is allowed if the width and height of one of the chroma CUs, chroma PUs or chroma TUs are denoted as W and H, the MM-CCRM is allowed when at least one of the following conditions is met: W H > T0 or W H >= T0; W > T1, or, W >= T1; H > T2, or, H >= T2; Min (W,H) > T3, or, Min (W,H) >= T3; Max (W,H) < T4, or, Max (W,H) <= T4; W < T5 H, or, W <= T5 H; W > T6 H, or, W >= T6 H; H < T7 W, or, H <= T7 W; H > T8 W, or, H >= T8 W; or W H < T9, or W H <= T9, and T0, T1, T2, T3, T4, T5, T6, T7, T8, and T9 are threshold parameters.

[0507] Item 85. The method of item 83, wherein the MM-CCRM is disabled for blocks for which tool is enabled.

[0508] Item 86. The method of item 85, wherein the MM-CCRM is disabled if affine motion compensation is enabled.

[0509] Item 87. The method of item 1, wherein CCRM coded video units use multi-model CCRM, or wherein the CCRM coded video units use single model CCRM or multi-model CCRM.

[0510] Item 88. The method of item 1, wherein which filter is used for a CCRM mode is indicated, wherein which filter is used for the CCRM mode is determined based on decoder derived cost from both the encoder and the decoder.

[0511] Item 89. The method of item 1, wherein the video units inherit parameters of the CCRM from a previous CCRM coded block.

[0512] Item 90. The method of item 89, wherein the parameters include at least one of: filter information, one or more linear parameters of a model, one or more non-linear parameters of the model, or a model index.

[0513] Item 91. The method of item 1, wherein the determination of the one or more CCRMs is used in at least one of: single tree or dual tree.

[0514] Item 92. The method of item 1, wherein the determination of the one or more CCRMs is used in an inter frame slice.

[0515] Item 93. The method of item 92, wherein the inter frame slice is a B slice or a P slice.

[0516] Item 94. The method of item 1, wherein the determination of the one or more CCRMs is used in an intra frame slice.

[0517] Item 95. The method of item 94, wherein the intra frame slice is an I slice.

[0518] Item 96. The method of item 1, wherein a training sample or a reference sample is a predicted sample in a training region or a reference region.

[0519] Item 97. The method of item 1, wherein a training sample or a reference sample is a reconstructed sample in a training region or a reference region.

[0520] Item 98. The method of any of items 1-97, wherein an indication of whether and / or how the one or more CCRMs are determined for the video unit is indicated at one of: a sequence level, a picture group level, a picture level, a slice level, or a tile group level.

[0521] Item 99. The method of any of items 1-97, wherein an indication of whether and / or how the one or more CCRMs are determined for the video unit is indicated in one of: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependent parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.

[0522] Item 100. The method of any of items 1-97, wherein an indication of whether and / or how the one or more CCRMs for the video unit are determined is included in one of: a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding unit (CU), a virtual pipeline data unit (VPDU), a coding tree unit (CTU), a CTU row, a slice, a tile, a subpicture, or a region containing more than one sample or pixel.

[0523] Item 101. The method of any of items 1-97, further comprising determining, based on coded information of the video unit, whether and / or how the one or more CCRMs for the video unit are determined, the coded information including at least one of: a block size, a color format, a single tree partitioning and / or a dual tree partitioning, a color component, a slice type, or a picture type.

[0524] Item 102. The method of any of items 1-97, wherein the converting includes encoding the video unit into the bitstream.

[0525] Item 103. The method of any of items 1-97, wherein the converting includes decoding the video unit from the bitstream.

[0526] Item 104. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of items 1-103.

[0527] Item 105. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method of any of items 1-103.

[0528] Item 106. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises determining one or more cross-component residual models (CCRM) for a video unit of the video; and generating the bitstream of the video unit based on the one or more CCRMs.

[0529] Item 107. A method for storing a bitstream of a video, comprising: determining one or more cross-component residual models (CCRM) for a video unit of the video; generating the bitstream of the video unit based on the one or more CCRM; and storing the bitstream in a non-transitory computer-readable recording medium.

[0530] Example device Figure 28 A block diagram of a computing device 2800 in which various embodiments of the present disclosure can be implemented is shown. The computing device 2800 can be implemented as, or included in, the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300).

[0531] It should be understood that, Figure 28 The computing device 2800 shown in FIG. 13 is for purposes of illustration and explanation only and is not intended to imply any limitation on the functionality and scope of embodiments of the present disclosure.

[0532] As Figure 28 As shown, the computing device 2800 includes a general-purpose computing device 2800. The computing device 2800 can include at least one or more processors or processing units 2810, a memory 2820, a storage unit 2830, one or more communication units 2840, one or more input devices 2850, and one or more output devices 2860.

[0533] In some embodiments, the computing device 2800 can be implemented as any user terminal or server terminal having computing capability. The server terminal can be a server provided by a service provider, a mainframe computing device, or the like. The user terminal may, for example, be any type of mobile terminal, fixed terminal, or portable terminal including a mobile telephone, a station, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 2800 can support any type of interface to the user (such as "wearable" circuitry, etc.).

[0534] Processing unit 2810 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 2800. Processing unit 2810 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0535] Computing device 2800 typically includes various computer storage media. Such media can be any media accessible by computing device 2800, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2830 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2800.

[0536] The computing device 2800 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 28 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0537] Communication unit 2840 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 2800 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 2800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0538] Input device 2850 can be one or more of various input devices such as a mouse, keyboard, trackball, voice input device, or the like. Output device 2860 can be one or more of various output devices such as a display, speakers, a printer, and the like. Computing device 2800 can also include communication interface 2840 taking the form of one or more communication buses, and one or more input / output (I / O) interfaces 2870, which can enable computing device 2800 to communicate with one or more external devices (not shown), such as storage devices and display devices, or one or more devices that enable a user to interact with computing device 2800.

[0539] In some embodiments, some or all of the components of computing device 2800 can also be arranged in a cloud computing architecture, rather than being integrated in a single device. In a cloud computing architecture, components can be provided remotely and work together to achieve the functionality described in this disclosure. In some embodiments, cloud computing provides computation, software, data access, and storage services that do not require end-user knowledge of the physical location or configuration of the system that delivers the services. In various embodiments, cloud computing uses suitable protocols and interfaces, such as APIs, to communicate with a network that serves the functionality described in this disclosure. For example, a cloud computing provider provides an application over a wide area network, such as the Internet, which can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, and corresponding data, can be stored on servers at remote locations. Computing resources in a cloud computing environment can be consolidated or distributed at locations remote from the user. Cloud computing infrastructure can provide services through a shared data center, although they appear to the user as a single point of access. Thus, a cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, the components and functionality described herein can be provided by a conventional server, or installed directly on a client device, either directly or in other ways.

[0540] In embodiments of the present disclosure, computing device 2800 can be used to implement video encoding / decoding. Memory 2820 can include one or more video codec modules 2825 having one or more program instructions. These modules are accessible and executable by processing unit 2810 to perform the functions of the various embodiments described herein.

[0541] In example embodiments in which video encoding is performed, input device 2850 can receive video data as input 2870 to be encoded. The video data can be processed, e.g., by video coding module 2825, to generate an encoded bitstream. The encoded bitstream can be provided as output 2880 via output device 2860.

[0542] In example embodiments in which video decoding is performed, input device 2850 can receive an encoded bitstream as input 2870. The encoded bitstream can be processed, e.g., by video coding module 2825, to generate decoded video data. The decoded video data can be provided as output 2880 via output device 2860.

[0543] While this disclosure has been particularly shown and described with reference to the preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of embodiments of the application is not intended to be limiting.

Claims

1. A method for video processing, comprising: determining one or more cross-component residual models (CCRM) for a video unit of a video for a conversion between the video unit and a bitstream of the video; and performing the conversion based on the one or more CCRM.

2. The method of claim 1, wherein the one or more CCRM comprises a multi-mode CCRM (MM-CCRM).

3. The method of claim 1, wherein training samples of the one or more CCRM are partitioned into multiple categories, and the samples of each category are applied to a unique model.

4. The method of claim 3, wherein training sample pairs of luma and chroma sample pairs of a reference block are partitioned into multiple categories according to MM-CCRM modes.

5. The method of claim 3, wherein training sample pairs of luma and chroma sample pairs of neighboring samples adjacent to a reference block are partitioned into multiple categories according to MM-CCRM modes, and / or wherein training sample pairs of luma and chroma sample pairs of neighboring samples non-adjacent to the reference block are partitioned into multiple categories according to MM-CCRM modes.

6. The method of claim 3, wherein training sample pairs of luma and chroma sample pairs of neighboring samples adjacent to the video unit are partitioned into multiple categories according to MM-CCRM modes, and / or wherein training sample pairs of luma and chroma sample pairs of neighboring samples non-adjacent to the video unit are partitioned into multiple categories according to MM-CCRM modes.

7. The method of claim 3, wherein luma samples in the video unit are partitioned into multiple groups according to the same criteria, and for luma samples belonging to each category, a corresponding model is applied to generate model estimated chroma samples belonging to the group.

8. The method of claim 1, wherein multiple sets of training samples are utilized to derive multiple models.

9. The method of claim 9, wherein there are two sets of training samples, wherein the distance between the training samples and the current sample is different.

10. The method of claim 1, wherein a threshold used to separate samples into different categories depends on the value of one or more samples within the training region, or wherein a threshold used to separate samples into different categories depends on the value of one or more samples adjacent to the training region.

11. The method of claim 10, wherein the threshold is a category threshold.

12. The method of claim 10, wherein the training samples are in a reference frame.

13. The method of claim 10, wherein the threshold is derived based on training samples adjacent to a reference block of the video unit, and / or wherein the threshold is derived based on training samples non-adjacent to the reference block of the video unit.

14. The method of claim 13, wherein the training samples are in a reference frame.

15. The method of claim 10, wherein the threshold is derived based on training samples adjacent to the video unit, and / or ​ ​ ​ ​ ​ wherein the threshold value is derived based on training samples that are not adjacent to the video unit.

16. The method of claim 15, wherein the training samples are in a current frame.

17. The method of claim 10, wherein the threshold value is derived based on at least one of: an average operation, a median operation, or a mid operation on a plurality of samples within or adjacent to a training region.

18. The method of claim 10, wherein a category threshold value is derived based on a non-downsampled luma sample value.

19. The method of claim 10, wherein a category threshold value is derived based on a downsampled luma sample value.

20. The method of claim 19, wherein a K-tap downsampling filter is used to downsize K surrounding luma samples to one downsampled luma sample value, wherein K is an integer.

21. The method of claim 20, wherein K is 6.

22. The method of claim 10, wherein a category threshold value is derived based on an offset removal scheme.

23. The method of claim 22, wherein the offset is derived based on a luma sample located at a fixed position in a reference video unit.

24. The method of claim 23, wherein the fixed position is one of: top-left or center.

25. The method of claim 22, wherein the offset value is the same for category threshold derivation and CCRM model calculation.

26. The method of claim 10, wherein a category threshold value is derived based on a subblock level.

27. The method of claim 10, wherein a category threshold value is derived based on one of: a coding unit (CU) level, a prediction unit (PU) level, or a transform unit (TU) level.

28. The method of claim 10, wherein a category threshold value is calculated based on a luma prediction sample.

29. The method of claim 28, wherein the luma prediction sample is downsampled, or wherein the luma prediction sample is non-downsampled.

30. The method of claim 10, wherein a category threshold value is derived based on a luma residual sample value.

31. The method of claim 30, wherein for a second video unit that does not have a non-zero residual, prediction samples of the second video unit are not counted in a calculation process for the category threshold value for a first video unit.

32. The method of claim 31, wherein the second video unit is a subset of the first video unit, or wherein the second video unit is equal to the first video unit, and / or wherein the second video unit is a subblock.

33. The method of claim 1, wherein MM-CCRM is applied based on a subblock level.

34. The method of claim 33, wherein a subblock size is predefined.

35. The method of claim 34, wherein the predefined subblock size is 16x16 or 32x32.

36. The method of claim 34, wherein a predefined rule is used to determine the subblock size for MM-CCRM for a target video unit.

37. The method of claim 36, wherein the sub-block block size is adapted to a block dimension of a current video block.

38. The method of claim 37, wherein the block dimension comprises at least one of a width or a height.

39. The method of claim 36, wherein a minimum number of chroma samples is guaranteed for sub-blocks of a video unit coded with MM-CCRM.

40. The method of claim 33, wherein if the video unit is larger than a predefined sub-block size, the video unit is divided into a plurality of sub-blocks and the MM-CCRM is applied.

41. The method of claim 33, wherein at least one sub-block of the video unit has a plurality of CCRM models.

42. The method of claim 33, wherein each sub-block and / or an associated training region of each sub-block has a class threshold.

43. The method of claim 42, wherein the class threshold of a target sub-block is calculated based on training sample values belonging to the target sub-block.

44. The method of claim 43, wherein luma training samples in a reference block are used to calculate the class threshold.

45. The method of claim 33, wherein all sub-blocks and / or associated training regions of all sub-blocks share a same class threshold.

46. The method of claim 45, wherein one class threshold is calculated and used for all sub-blocks.

47. The method of claim 45, wherein the class threshold of all sub-blocks in the video unit is calculated based on training sample values of the video unit.

48. The method of claim 45, wherein the class threshold of all applicable sub-blocks in the video unit is calculated based on training sample values of the video unit.

49. The method of claim 48, wherein sub-blocks that do not contain non-zero residuals are not counted.

50. The method of claim 33, wherein each sub-block of the video unit has its own training samples and training samples of a target sub-block are divided into a plurality of classes.

51. The method of claim 50, wherein training samples in a reference video unit of a reference picture are classified based on sub-blocks.

52. The method of claim 33, wherein training samples of a current picture are classified into a plurality of groups but not divided into sub-blocks.

53. The method of claim 1, wherein MM-CCRM is applied on one of: TU level, PU level, or CU level.

54. The method of claim 53, wherein the MM-CCRM is applied on one of: TU basis, CU basis, or PU basis.

55. The method of claim 54, wherein for the application of the MM-CCRM, TUs are not divided into sub-blocks, and / or wherein for the application of the MM-CCRM, CUs are not divided into sub-blocks, and / or wherein for the application of the MM-CCRM, PUs are not divided into sub-blocks.

56. The method of claim 53, wherein whether to use one of TU-based multi-model CCRM, PU-based multi-model CCRM, or CU-based multi-model CCRM is determined at one of: TU level, PU level, or CU level.

57. The method of claim 56, wherein the video unit decides to use sub-block based single model CCRM or one of: TU-based MM-CCRM, PU-based MM-CCRM, or CU-based MM-CCRM.

58. The method of claim 57, wherein the decision is made at one of: TU level, PU level, or CU level.

59. The method of claim 1, wherein whether and / or how to apply at least one of MM-CCRM or CCRM is inferred based on codec information at both encoder side and decoder side.

60. The method of claim 59, wherein whether and / or how to apply at least one of MM-CCRM or CCRM is inferred on-the-fly.

61. The method of claim 60, wherein whether and / or how to apply at least one of MM-CCRM or CCRM is inferred using information of previously coded or reconstructed samples.

62. The method of claim 59, wherein whether to use sub-block based CCRM or one of: TU-level CCRM, CU-level CCRM, or PU-level CCRM is implicitly inferred based on codec information.

63. The method of claim 59, wherein whether to use MlxM2 sub-block based CCRM or NlxN2 sub-block based CCRM is implicitly inferred based on codec information.

64. The method of claim 63, wherein M1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein M2 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N2 = 16 or 8 or 32 or TU or CU or PU.

65. The method of claim 63, wherein M1 is not equal to N1 and / or M2 is not equal to N2.

66. The method of claim 60, wherein whether to use sub-block based MM-CCRM or one of: TU-level MM-CCRM, CU-level MM-CCRM, or PU-level MM-CCRM is inferred based on codec information.

67. The method of claim 60, wherein whether to use MlxM2 sub-block based MM-CCRM or NlxN2 sub-block based MM-CCRM is implicitly inferred based on codec information.

68. The method of claim 67, wherein M1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein M2 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N2 = 16 or 8 or 32 or TU or CU or PU. where N2 = 16 or 8 or 32 or TU or CU or PU, and / or where M1 is not equal to N1 and / or M2 is not equal to N2.

69. The method of claim 60, wherein the determination of whether to use a single model CCRM or MM-CCRM is derived based on coding information.

70. The method of claim 60, wherein the determination of whether and / or how to apply at least one of MM-CCRM or CCRM is according to a scheme based on decoder-derived cost.

71. The method of claim 70, wherein the decoder-derived cost is calculated based on minimizing one of sum of absolute difference (SAD), sum of absolute transformed difference (SATD), sum of squared error (SSE), or mean squared error (MSE) between model-estimated sample values of samples and true reconstructed sample values of the samples, wherein the samples include at least one of the training samples.

72. The method of claim 70, wherein the scheme based on decoder-derived cost with lower cost is selected as the final scheme applied to the video unit.

73. The method of claim 60, wherein the determination of whether and / or how to apply at least one of MM-CCRM or CCRM is based on information of reference pictures.

74. The method of claim 73, wherein the determination is based on picture order count (POC) distance of a current picture and reference pictures of the current picture.

75. The method of claim 73, wherein the determination is based on reference index.

76. The method of claim 1, wherein whether and / or how to apply at least one of MM-CCRM or CCRM is signaled in the bitstream.

77. The method of claim 76, wherein a syntax element is signaled based on a condition of whether the video unit is CCRM coded.

78. The method of claim 77, wherein a syntax element is further signaled to indicate whether it is MM-CCRM if the video unit is CCRM coded.

79. The method of claim 76, wherein a syntax element is signaled to indicate whether it is sub-block based MM-CCRM or one of the following: TU-based MM-CCRM, CP-based MM-CCRM, or PU-based MM-CCRM.

80. The method of claim 76, wherein a syntax element is signaled to indicate whether it is sub-block based CCRM or one of the following: TU-based CCRM, CP-based CCRM, or PU-based CCRM.

81. The method of claim 76, wherein a syntax element is signaled based on a condition on block dimension.

82. The method of claim 81, wherein the block dimension includes at least one of width or height.

83. The method of claim 1, wherein block restriction is applied to indicate allowance of MM-CCRM mode.

84. The method of claim 83, wherein if width and height of one of chroma CU, chroma PU or chroma TU are denoted as W and H, the MM-CCRM is allowed when at least one of the following conditions is met: W H > T0or W H >= T0; W > T1, or, W >= T1; H > T2, or, H >= T2; Min (W,H) > T3, or, Min (W,H) >= T3; Max (W,H) < T4, or, Max (W,H) <= T4; W < T5 H, or W <= T5 H; W > T6 H, or W >= T6 H; H < T7 W, or H <= T7 W; H > T8 W, or H >= T8 W; or W H < T9, or W H <= T9, and TO, T1, T2, T3, T4, T5, T6, T7, T8, and T9 are threshold parameters.

85. The method of claim 83, wherein the MM-CCRM is disabled for blocks for which tool is enabled.

86. The method of claim 85, wherein the MM-CCRM is disabled if affine motion compensation is enabled.

87. The method of claim 1, wherein a CCRM coded video unit uses multi-model CCRM, or wherein the CCRM coded video unit uses single model CCRM or multi-model CCRM.

88. The method of claim 1, wherein which filter is used for a CCRM mode is indicated, wherein which filter is used for the CCRM mode is determined based on decoder derived cost from both encoder and decoder.

89. The method of claim 1, wherein the video unit inherits parameters of the CCRM from a previous CCRM coded block.

90. The method of claim 89, wherein the parameters include at least one of: filter information, one or more linear parameters of a model, one or more non-linear parameters of the model, or a model index.

91. The method of claim 1, wherein the determination of the one or more CCRM is used in at least one of: single tree or dual tree.

92. The method of claim 1, wherein the determination of the one or more CCRM is used in an inter frame slice.

93. The method of claim 92, wherein the inter frame slice is a B slice or a P slice.

94. The method of claim 1, wherein the determination of the one or more CCRM is used in an intra frame slice.

95. The method of claim 94, wherein the intra frame slice is an I slice.

96. The method of claim 1, wherein a training sample or a reference sample is a predicted sample in a training region or a reference region.

97. The method of claim 1, wherein a training sample or a reference sample is a reconstructed sample in a training region or a reference region.

98. The method of any of claims 1-97, wherein an indication of whether and / or how the one or more CCRM for the video unit is determined is indicated at one of: sequence level, picture group level, picture level, slice level, or tile group level.

99. The method of any of claims 1-97, wherein an indication of whether and / or how the one or more CCRMs for the video unit are determined is indicated in one of: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependent parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.

100. The method of any of claims 1-97, wherein an indication of whether and / or how the one or more CCRMs for the video unit are determined is included in one of: a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding unit (CU), a virtual pipeline data unit (VPDU), a coding tree unit (CTU), a CTU row, a slice, a tile, a sub-picture, or a region containing more than one sample or pixel.

101. The method of any of claims 1-97, further comprising: determining, based on coded information of the video unit, whether and / or how the one or more CCRMs for the video unit are determined, the coded information including at least one of: a block size, a color format, a single tree partitioning and / or a dual tree partitioning, a color component, a slice type, or a picture type.

102. The method of any of claims 1-101, wherein the converting comprises encoding the video unit into the bitstream.

103. The method of any of claims 1-101, wherein the converting comprises decoding the video unit from the bitstream.

104. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of claims 1-103.

105. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method of any of claims 1-103.

106. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining one or more cross-component residual models (CCRM) for a video unit of the video; and generating the bitstream of the video unit based on the one or more CCRMs.

107. A method for storing a bitstream of a video, comprising: determining one or more cross-component residual models (CCRM) for a video unit of the video; generating the bitstream of the video unit based on the one or more CCRMs; and storing the bitstream in a non-transitory computer-readable recording medium. ​