Method, apparatus, and medium for video processing
CCM-based filters enhance video coding quality by improving filtering during the filtering stage, addressing inefficiencies in existing video processing techniques.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2025-11-28
- Publication Date
- 2026-06-04
AI Technical Summary
Existing video coding techniques require improvements in coding quality to enhance the efficiency and effectiveness of video processing.
The implementation of cross-component model-based filters (CCM-based) during the filtering stage of video processing to improve the filtering quality and coding efficiency.
Enhances the coding quality by introducing CCM-based filters, leading to improved video processing outcomes.
Smart Images

Figure CN2025138709_04062026_PF_FP_ABST
Abstract
Description
METHOD, APPARATUS, AND MEDIUM FOR VIDEO PROCESSINGFIELDS
[0001] Embodiments of the present disclosure generally relate to computer technologies, and more particularly, to image / video coding.BACKGROUND
[0002] In nowadays, digital video capabilities are being applied in various aspects of peoples’ lives. Multiple types of video compression technologies, such as motion picture expert group (MPEG) -2, MPEG-4, international telecommunication union -telecommunication standardization sector (ITU-T) H. 263, ITU-T H. 264 / MPEG-4 Part 10 advanced video coding (AVC) , ITU-T H. 265 high efficiency video coding (HEVC) standard, versatile video coding (VVC) standard, have been proposed for video encoding / decoding. However, coding quality of video coding techniques is generally expected to be further improved.SUMMARY
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method comprises: obtaining, for a conversion between a current video block of a video and a bitstream of the video, reconstructed samples of the current video block; applying at least one cross-component model based (CCM-based) filter on the reconstructed samples; and performing the conversion based on the applying.
[0005] Based on the method in accordance with the first aspect of the present disclosure, at least one CCM-based filter is applied on reconstructed samples of the current video block. Compared with the conventional solution, the proposed method can advantageously introduce a CCM-based filter (s) into the filtering stage, and improve the filtering quality. Thereby, the coding quality can be improved.
[0006] In a second aspect, an apparatus for video processing is proposed. The apparatus comprises a processor and a non-transitory memory with instructions thereon. The instructions upon execution by the processor, cause the processor to perform a method in accordance with the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to perform a method in accordance with the first aspect of the present disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video which is generated by a method performed by an apparatus for video processing. The method comprises: obtaining reconstructed samples of a current video block of the video; applying at least one cross-component model based (CCM-based) filter on the reconstructed samples; and generating the bitstream based on the applying.
[0009] In a fifth aspect, a method for storing a bitstream of a video is proposed. The method comprises: obtaining reconstructed samples of a current video block of the video; applying at least one cross-component model based (CCM-based) filter on the reconstructed samples; generating the bitstream based on the applying; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Through the following detailed description with reference to the accompanying drawings, the above and other objectives, features, and advantages of example embodiments of the present disclosure will become more apparent. In the example embodiments of the present disclosure, the same reference numerals usually refer to the same components.
[0012] FIG. 1 illustrates a block diagram of an example video coding system in accordance with some embodiments of the present disclosure;
[0013] FIG. 2 illustrates a block diagram of an example video encoder in accordance with some embodiments of the present disclosure;
[0014] FIG. 3 illustrates a block diagram of an example video decoder in accordance with some embodiments of the present disclosure;
[0015] FIG. 4 illustrates an illustration of the effect of the slope adjustment parameter “u” . Left: model created with the current CCLM. Right: model updated as proposed;
[0016] FIG. 5 illustrates neighboring blocks (L, A, BL, AR, AL) used in the derivation of a general MPM list;
[0017] FIG. 6 illustrates non-adjacent spatial neighboring candidates for OBIC mode;
[0018] FIG. 7 illustrates neighboring reconstructed samples used for DIMD chroma mode;
[0019] FIG. 8 illustrates the modified search range for LIC;
[0020] FIG. 9 illustrates the intra template matching search area used;
[0021] FIG. 10 illustrates an example of IntraTMP-AR-BVP’s construction;
[0022] FIG. 11 illustrates the five positions in reference Block;
[0023] FIG. 12 illustrates the use of IntraTMP block vector for IBC block;
[0024] FIG. 13A and FIG. 13B illustrate the division method for angular modes, respectively;
[0025] FIG. 14 illustrates an extended MRL candidate list;
[0026] FIG. 15 illustrates the template area;
[0027] FIG. 16 illustrates a spatial part of the convolutional filter;
[0028] FIG. 17 illustrates a reference area used to derive the filter coefficients with its paddings;
[0029] FIG. 18 illustrates four Sobel based gradient patterns for GLM;
[0030] FIG. 19 illustrates non-downsampled luma samples;
[0031] FIG. 20 illustrates a reference area for BVG-CCCM;
[0032] FIG. 21 illustrates spatial samples used for GL-CCCM;
[0033] FIG. 22 illustrates various downsampling filters used in cross-component models;
[0034] FIG. 23 illustrates a filter on samples of MM-CCLM / MM-CCCM;
[0035] FIG. 24 illustrates the template adjacent to the current chroma CU;
[0036] FIG. 25 illustrates spatial GPM candidates;
[0037] FIG. 26 illustrates a GPM template;
[0038] FIG. 27 illustrates a GPM blending;
[0039] FIG. 28 illustrates a transform selection process for directional planar modes;
[0040] FIG. 29 illustrates luma blocks used to derive direct block vector;
[0041] FIG. 30 illustrates three EIP filter shapes;
[0042] FIG. 31 illustrates three types of reconstructed area for EIP filter;
[0043] FIG. 32 illustrates a L shaped neighborhood for a given predicted block;
[0044] FIG. 33 illustrates the collocated luma block and adjacent chroma block;
[0045] FIG. 34 illustrates the InterCCCM method on the decoder;
[0046] FIG. 35 illustrates luma samples L0, . ., L5 in relation to the chroma sample C;
[0047] FIG. 36 illustrates reference area for BVG-CCCM;
[0048] FIG. 37 illustrates a spatial part of the convolutional filter;
[0049] FIG. 38 illustrates locations used for block vector derivation from co-located luma block;
[0050] FIG. 39 illustrates a 25-tap long filter;
[0051] FIG. 40 illustrates an illustration of chroma SAO output samples applied to 4 taps in a 3x3 asymmetric cross shape;
[0052] FIG. 41 illustrates the diamond shape of ALF in ECM-5.0;
[0053] FIG. 42 illustrates the new filter shape for ALF;
[0054] FIG. 43A and FIG. 43B illustrate ALF filter shapes, respectively;
[0055] FIG. 44A and FIG. 44B illustrate an additional fixed filter;
[0056] FIG. 45 illustrates the decoder-side diagram of ALF-CCCM;
[0057] FIG. 46 illustrates an in-loop filtering in ECM 13.0;
[0058] FIG. 47 illustrates a BIF shape in ECM 13.0;
[0059] FIG. 48 illustrates a filtering stage of BIF-Chroma;
[0060] FIG. 49 illustrates a modified SAO process when the proposed CCSAO is applied;
[0061] FIG. 50 illustrates an illustration of the candidate positions used for the CCSAO classifier;
[0062] FIG. 51 illustrates the joint clipping after adding SAO / BIF / CCSAO offsets to the input sample;
[0063] FIG. 52 illustrates four 1-D directional patterns for CCSAO EO sample classification;
[0064] FIG. 53 illustrates a flowchart of a method for video processing in accordance with some embodiments of the present disclosure; and
[0065] FIG. 54 illustrates a block diagram of a computing device in which various embodiments of the present disclosure can be implemented.
[0066] Throughout the drawings, the same or similar reference numerals usually refer to the same or similar elements.DETAILED DESCRIPTION
[0067] Principle of the present disclosure will now be described with reference to some embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below.
[0068] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0069] References in the present disclosure to “one embodiment, ” “an embodiment, ” “an example embodiment, ” and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0070] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0071] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Example Environment
[0072] FIG. 1 is a block diagram that illustrates an example video coding system 100 that may utilize the techniques of this disclosure. As shown, the video coding system 100 may include a source device 110 and a destination device 120. The source device 110 can be also referred to as a video encoding device, and the destination device 120 can be also referred to as a video decoding device. In operation, the source device 110 can be configured to generate encoded video data and the destination device 120 can be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0073] The video source 112 may include a source such as a video capture device. Examples of the video capture device include, but are not limited to, an interface to receive video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0074] The video data may comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be transmitted directly to destination device 120 via the I / O interface 116 through the network 130A. The encoded video data may also be stored onto a storage medium / server 130B for access by destination device 120.
[0075] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120 which is configured to interface with an external display device.
[0076] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard and other current and / or further standards.
[0077] FIG. 2 is a block diagram illustrating an example of a video encoder 200, which may be an example of the video encoder 114 in the system 100 illustrated in FIG. 1, in accordance with some embodiments of the present disclosure.
[0078] The video encoder 200 may be configured to implement any or all of the techniques of this disclosure. In the example of FIG. 2, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0079] In some embodiments, the video encoder 200 may include a partition unit 201, a prediction unit 202 which may include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0080] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture where the current video block is located.
[0081] Furthermore, although some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be integrated, but are represented in the example of FIG. 2 separately for purposes of explanation.
[0082] The partition unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0083] The mode select unit 203 may select one of the coding modes, intra or inter, e.g., based on error results, and provide the resulting intra-coded or inter-coded block to a residual generation unit 207 to generate residual block data and to a reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode select unit 203 may select a combined inter and intra prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. The mode select unit 203 may also select a resolution for a motion vector (e.g., a sub-pixel or integer pixel precision) for the block in the case of inter-prediction.
[0084] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.
[0085] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations for a current video block, for example, depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an “I-slice” may refer to a portion of a picture composed of macroblocks, all of which are based upon macroblocks within the same picture. Further, as used herein, in some aspects, “P-slices” and “B-slices” may refer to portions of a picture composed of macroblocks that are not dependent on macroblocks in the same picture.
[0086] In some examples, the motion estimation unit 204 may perform uni-directional prediction for the current video block, and the motion estimation unit 204 may search reference pictures of list 0 or list 1 for a reference video block for the current video block. The motion estimation unit 204 may then generate a reference index that indicates the reference picture in list 0 or list 1 that contains the reference video block and a motion vector that indicates a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate the predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0087] Alternatively, in other examples, the motion estimation unit 204 may perform bi-directional prediction for the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate reference indexes that indicate the reference pictures in list 0 and list 1 containing the reference video blocks and motion vectors that indicate spatial displacements between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference indexes and the motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate the predicted video block of the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0088] In some examples, the motion estimation unit 204 may output a full set of motion information for decoding processing of a decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0089] In one example, the motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as the another video block.
[0090] In another example, the motion estimation unit 204 may identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD) . The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0091] As discussed above, video encoder 200 may predictively signal the motion vector. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0092] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0093] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by the minus sign) the predicted video block (s) of the current video block from the current video block. The residual data of the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.
[0094] In other examples, there may be no residual data for the current video block, for example in a skip mode, and the residual generation unit 207 may not perform the subtracting operation.
[0095] The transform unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to a residual video block associated with the current video block.
[0096] After the transform unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0097] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transforms to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.
[0098] After the reconstruction unit 212 reconstructs the video block, loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0099] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0100] FIG. 3 is a block diagram illustrating an example of a video decoder 300, which may be an example of the video decoder 124 in the system 100 illustrated in FIG. 1, in accordance with some embodiments of the present disclosure.
[0101] The video decoder 300 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 3, the video decoder 300 includes a plurality of functional components. The techniques described in this disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0102] In the example of FIG. 3, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306 and a buffer 307. The video decoder 300 may, in some examples, perform a decoding pass generally reciprocal to the encoding pass described with respect to video encoder 200.
[0103] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy coded video data (e.g., encoded blocks of video data) . The entropy decoding unit 301 may decode the entropy coded video data, and from the entropy decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may, for example, determine such information by performing the AMVP and merge mode. AMVP is used, including derivation of several most probable candidates based on data from adjacent PBs and the reference picture. Motion information typically includes the horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, a “merge mode” may refer to deriving the motion information from spatially or temporally neighboring blocks.
[0104] The motion compensation unit 302 may produce motion compensated blocks, possibly performing interpolation based on interpolation filters. Identifiers for interpolation filters to be used with sub-pixel precision may be included in the syntax elements.
[0105] The motion compensation unit 302 may use the interpolation filters as used by the video encoder 200 during encoding of the video block to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 according to the received syntax information and use the interpolation filters to produce predictive blocks.
[0106] The motion compensation unit 302 may use at least part of the syntax information to determine sizes of blocks used to encode frame (s) and / or slice (s) of the encoded video sequence, partition information that describes how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information to decode the encoded video sequence. As used herein, in some aspects, a “slice” may refer to a data structure that can be decoded independently from other slices of the same picture, in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice can either be an entire picture or a region of a picture.
[0107] The intra prediction unit 303 may use intra prediction modes for example received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes, i.e., de-quantizes, the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0108] The reconstruction unit 306 may obtain the decoded blocks, e.g., by summing the residual blocks with the corresponding prediction blocks generated by the motion compensation unit 302 or intra-prediction unit 303. If desired, a deblocking filter may also be applied to filter the decoded blocks in order to remove blockiness artifacts. The decoded video blocks are then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0109] Some example embodiments of the present disclosure will be described in detailed hereinafter. It should be understood that section headings are used in the present document to facilitate ease of understanding and do not limit the embodiments disclosed in a section to only that section. Furthermore, while certain embodiments are described with reference to Versatile Video Coding or other specific video codecs, the disclosed techniques are applicable to other video coding technologies also. Furthermore, while some embodiments describe video coding steps in detail, it will be understood that corresponding steps decoding that undo the coding will be implemented by a decoder. Furthermore, the term video processing encompasses video coding or compression, video decoding or decompression and video transcoding in which video pixels are represented from one compressed format into another compressed format or at a different compressed bitrate. 1 Brief Summary The present disclosure is related to video coding technologies. Specifically, it is about the usage of cross-component prediction in image / video coding. It may be applied to the existing video coding standard like HEVC, VVC, and etc. It may be also applicable to future video coding standards or video codec. 2 Introduction Video coding standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. The ITU-T produced H. 261 and H. 263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H. 262 / MPEG-2 Video and H. 264 / MPEG-4 Advanced Video Coding (AVC) and H. 265 / HEVC standards. Since H. 262, the video coding standards are based on the hybrid video coding structure wherein temporal prediction plus transform coding are utilized. To explore the future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. The JVET meeting is concurrently held once every quarter, and the new video coding standard was officially named as Versatile Video Coding (VVC) in the April 2018 JVET meeting, and the first version of VVC test model (VTM) was released at that time. The VVC working draft and test model VTM are then updated after every meeting. The VVC project achieved technical completion (FDIS) at the July 2020 meeting. 2.1 Intra prediction In intra prediction the smallest chroma intra prediction unit (SCIPU) constraint in VVC is removed. In addition, the VPDU constraint for reducing CCLM prediction latency is also removed. 2.1.1 Multi-model LM (MMLM) CCLM included in VVC is extended by adding three Multi-model LM (MMLM) modes. In each MMLM mode, the reconstructed neighboring samples are classified into two classes using a threshold which is the average of the luma reconstructed neighboring samples. The linear model of each class is derived using the Least-Mean-Square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model. A slope adjustment to is applied to cross-component linear model (CCLM) and to Multi-model LM prediction. The adjustment is tilting the linear function which maps luma values to chroma values with respect to a center point determined by the average luma value of the reference samples. 2.1.1.1 Slope adjustment of CCLM CCLM uses a model with 2 parameters to map luma values to chroma values. The slope parameter “a” and the bias parameter “b” define the mapping as follows: chromaVal = a *lumaVal + b An adjustment “u” to the slope parameter is signaled to update the model to the following form: chromaVal = a’ *lumaVal + b’ where a’= a + u b’= b -u *yr. With this selection the mapping function is tilted or rotated around the point with luminance value yr. The average of the reference luma samples used in the model creation as yr in order to provide a meaningful modification to the model. Picture below illustrates the process. FIG. 4 illustrates an illustration of the effect of the slope adjustment parameter “u” . Left: model created with the current CCLM. Right: model updated as proposed. Implementation Slope adjustment parameter is provided as an integer between -4 and 4, inclusive, and signaled in the bitstream. The unit of the slope adjustment parameter is 1 / 8th of a chroma sample value per one luma sample value (for 10-bit content) . Adjustment is available for the CCLM models that are using reference samples both above and left of the block ( “LM_CHROMA_IDX” and “MMLM_CHROMA_IDX” ) , but not for the “single side” modes. This selection is based on coding efficiency vs. complexity trade-off considerations. When slope adjustment is applied for a multimode CCLM model, both models can be adjusted and thus up to two slope updates are signaled for a single chroma block. Encoder approach The proposed encoder approach performs an SATD based search for the best value of the slope update for Cr and a similar SATD based search for Cb. If either one results as a non-zero slope adjustment parameter, the combined slope adjustment pair (SATD based update for Cr, SATD based update for Cb) is included in the list of RD checks for the TU. 2.1.2 Gradient PDPC In VVC, for a few scenarios, PDPC may not be applied due to the unavailability of the secondary reference samples. In these cases, a gradient based PDPC, extended from horizontal / vertical mode, is applied. The PDPC weights (wT / wL) and nScale parameter for determining the decay in PDPC weights with respect to the distance from left / top boundary are set equal to corresponding parameters in horizontal / vertical mode, respectively. When the secondary reference sample is at a fractional sample position, bilinear interpolation is applied. 2.1.3 Primary and Secondary MPM Secondary MPM lists is introduced. The existing primary MPM (PMPM) list consists of 6 entries and the secondary MPM (SMPM) list includes 16 entries. A general MPM list with 22 entries is constructed first, and then the first 6 entries in this general MPM list are included into the PMPM list, and the rest of entries form the SMPM list. The first entry in the general MPM list is the Planar mode. The remaining entries are composed of the intra modes of the left (L) , above (A) , below-left (BL) , above-right (AR) , and above-left (AL) neighbouring blocks, and DIMD modes which are sorted in ascending order of SAD cost. Up to 5 modes with the smallest SAD cost are added. The SAD cost is computed between the prediction and the reconstruction samples of the template. The sorted directional modes with added offset are added into the general MPM list, and then the default modes, until the general MPM list with 22 entries is constructed. If a CU block is vertically oriented, the order of neighbouring blocks is A, L, BL, AR, AL; otherwise, it is L, A, BL, AR, AL. FIG. 5 illustrates neighboring blocks (L, A, BL, AR, AL) used in the derivation of a general MPM list. MPM list is equally divided into four groups and the group index is parsed first. Then, a mode index is further parsed to indicate which mode in the selected group is used. 2.1.4 Reference sample interpolation and smoothing for intra-prediction The 4-tap cubic interpolation is replaced with a 6-tap cubic interpolation filter, for the derivation of predicted samples from the reference samples. For reference sample filtering, a 6-tap gaussian filter is applied for larger blocks (W >= 32 and H >=32) , existing VVC 4-tap gaussian interpolation filter is applied otherwise. The extended intra reference samples are derived using the 4-tap interpolation filter instead of the nearest neighbor rounding. 2.1.5 Decoder side intra mode derivation (DIMD) When DIMD is applied, up to five intra modes are derived from the reconstructed neighbor samples, and those five predictors are combined with the non-directional predictor (planar or block vector based predictor) with the weights derived from the histogram of gradients. The decision between for the non-directional modes is taken according to the template cost. Specifically, the block vectors of all adjacent and non-adjacent merge candidates (coded in IntraTMP or IBC) are compared to planar prediction on the reconstructed template. The template cost (SATD) is used to select the best predictor among them. The division operations in weight derivation are performed utilizing the same lookup table (LUT) based integerization scheme used by the CCLM. For example, the division operation in the orientation calculation Orient=Gy / Gx is computed by the following LUT-based scheme: x = Floor (Log2 (Gx) ) normDiff = ( (Gx<< 4) >> x) &15 x += (3 + (normDiff ! = 0) ? 1: 0) Orient = (Gy* (DivSigTable [normDiff ] | 8) + (1<< (x-1) ) ) >> x where DivSigTable
[0016] = {0, 7, 6, 5 , 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0 } . For a block of size W×H, the weight for each of the five derived modes is modified if the one the above or left histogram magnitudes is twice larger than the other one. In this case, the weights are location dependent and computed as follows: If the above histogram is twice the left, then: If the left histogram is twice the above, then: where wDimdi is the unmodified uniform weight of the DIMD selected, Δi is pre-defined and set to 10. Derived intra modes are included into the primary list of intra most probable modes (MPM) , so the DIMD process is performed before the MPM list is constructed. The primary derived intra mode of a DIMD block is stored with a block and is used for MPM list construction of the neighboring blocks. Finally, note the region of neighboring reconstructed samples used for computing the histogram of gradients is modified, depending on reconstructed samples availability. The region of decoded reference samples of current WxH luma CB is extended towards the above-right side if available, up to W additional columns. It is extended towards the bottom-left side if available, up to H additional rows. 2.1.5.1 Occurrence-based intra coding (OBIC) The occurrence-based intra coding (OBIC) derives the intra prediction modes of the current block based on the sample-wise occurrence of the intra modes in the spatial neighborhood of the block. For this, adjacent and non-adjacent spatial neighboring blocks are checked and the intra prediction modes of the blocks are collected into an occurrence histogram. Instead of Histogram of Gradients (HoGs) as in DIMD, the OBIC method uses the Histogram of Occurrences, which consists of the intra modes and their sample-wise occurrences. The occurrence values are calculated based on the number of samples that are coded in a certain intra prediction mode in that neighborhood. For example, if a uiWidth × uiHeight block is coded with an IPM mode, the occurrence of the mode in that block is calculated as: Histogram [IPM] += uiWidth × uiHeight; Where uiWidth and uiHeight are the width and height of a spatial neighboring block. The occurrences of the existing modes from the spatial neighborhood blocks are accumulated into the histogram. FIG. 6 shows the non-adjacent spatial neighboring blocks that are used in OBIC mode’s histogram generation. Up to five angular modes with the highest occurrence along with the planar mode or block vector-based prediction (same as in DIMD) are selected from the histogram and used for final prediction by blending the prediction of the selected modes. Some blocks, mentioned below, use more than one intra mode for prediction. In such cases, all the intra modes of such blocks are selected and used when creating the OBIC histogram: · DIMD: up to 5 angular modes; · TIMD: up to 2 modes; · SGPM: 2 modes; · OBIC: up to 5 angular modes. Moreover, the virtual intra prediction modes (VIPMs) of following blocks are considered only in inter slices when creating the histogram of OBIC mode: · MIP block; · IntraTMP block; · EIP block. The blending weights are calculated similar to the DIMD mode, but instead of using gradient values from the template, the occurrence values are used for OBIC. Moreover, the planar mode’s weight is also decided similar to DIMD mode. The OBIC mode is used as a sub-mode of DIMD tool and is applied to only luma blocks. Moreover, the mode is disabled for blocks that have less than 64 samples. 2.1.5.2 DIMD chroma mode The DIMD chroma mode uses the DIMD derivation method to derive the chroma intra prediction mode of the current block based on the neighboring reconstructed Y, Cb and Cr samples in the second neighboring row and column. Specifically, a horizontal gradient and a vertical gradient are calculated for each collocated reconstructed luma sample of the current chroma block, as well as the reconstructed Cb and Cr samples, to build a HoG. Then the intra prediction mode with the largest histogram amplitude values is used for performing chroma intra prediction of the current chroma block. FIG. 7 illustrates neighboring reconstructed samples used for DIMD chroma mode. When the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode with the second largest histogram amplitude value is used as the DIMD chroma mode. A CU level flag is signaled to indicate whether the proposed DIMD chroma mode is applied. Finally, the luma region of reconstructed samples used for computing the histogram of gradients for chroma DIMD mode is modified. For a WxH pair of chroma CBs to predict, to build the histogram of gradients associated to the collocated luma CB, the pairs of a vertical gradient and a horizontal gradient are extracted from the second and third lines in this luma CB instead of being extracted from the regular set of DIMD decoded reference samples around this luma CB. 2.1.6 Fusion of chroma intra prediction modes In ECM, two chroma intra prediction signals can be fused together. One of the two chroma intra prediction signals is predicted using one of the DM mode, DIMD chroma mode and the four default modes (non-LM mode) . The other chroma intra prediction signal is predicted using cross-component linear prediction modes (LM mode) . Two different methods are supported. In the first method, the LM mode can be either MM-CCLM or MM-CCCM, and the final predictor is derived as follows: predC (i, j) = (w0×pred0 (i, j) +w1×pred1 (i, j) + (1<< (shift-1) ) ) >>shift where pred0 (i, j) is the predictor obtained by applying the non-LM mode, pred1 (i, j) is the predictor obtained by applying the LM mode and predC (i, j) is the final predictor of the current chroma block. The two weights, w0 and w1 are determined by the intra prediction mode of adjacent chroma blocks and shift is set equal to 2. Specifically, when the above and left adjacent blocks are both coded with LM modes, {w0, w1} = {1, 3} ; when the above and left adjacent blocks are both coded with non-LM modes, {w0, w1} = {3, 1} ; otherwise, {w0, w1} = {2, 2} . Two template costs are calculated by fusing the angular chroma prediction with MM-CCLM or MM-CCCM, respectively, and the one of the two CCPs which provides a smaller template cost is utilized to derive pred1. In the second method, the LM mode can be either MMLM or CCLM mode, and the final predictor is derived as follows: predC (i, j) = α0×pred0 (i, j) + α1×recL′ (i, j) +α2×β where pred0 (i, j) is the predictor obtained by applying the non-LM mode, rec′L (i, j) is the set of downsampled reconstructed luma samples at co-located positions and predC (i, j) is the final predictor of the current chroma block. β is a fixed value and is set equal to 512 for 10-bit content. The three weights, α0, α1 and α2 are derived from the adjacent luma and chroma samples using the same LDL derivation method as in CCCM. For the syntax design, one index is signaled to indicate whether fusion is applied and which method is used. It is noted that for I slices, the non-LM mode can be DM mode, DIMD chroma mode and the four default modes. For non-I slices, only DIMD chroma mode is allowed to be fused with LM modes. 2.1.7 Intra template matching Intra template matching prediction (IntraTMP) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the most similar template to the current template in a reconstructed part of the current frame and uses the corresponding block as a prediction block. The encoder then signals the usage of this mode, and the same prediction operation is performed at the decoder side. The prediction signal is generated by matching the L-shaped, Top-only or Left-Only causal neighbor of the current block with another block in a predefined search area in FIG. 9. There are 6 predefined search areas, i.e., R1 to R6 in FIG. 9 which contain the reconstructed samples from the top and left CTUs as well as part of the reconstructed samples within the current CTU that are located above, left, bottom-left and top-right to the current block. IntraTMP employs an implicit merge mode, where merge candidates are considered without signaling a merge flag or index. Specifically, the reference positions pointed by the block vectors of all the adjacent and non-adjacent merge candidates (coded in IntraTMP or IBC mode) are used as additional candidates beyond the default search areas, up to 10 merge candidates are derived from the neighboring PUs and by prioritizing the candidates outside the IntraTMP search range. In addition, up to 20 auto-relocated block vector prediction (AR-BVP) candidates are constructed to get more reference positions. As shown in FIG. 10, a guiding block vector BV0 (i.e., an existing BVP already in the candidate list) associated with the current block B0 points at a reference block B1. If B1 has a BV denoted as BV1 pointing at a reference block B2, then BV0’ , given by BV0’ = BV0 + BV1, is defined as the AR-BVP. BV1 itself can also be directly used as AR-BVP to get more available candidates. When deriving AR-BVP, all five positions including top-left (e.g., LT in FIG. 11) , top-right (e.g., RT in FIG. 11) , center (e.g., Ctr in FIG. 11) , bottom-left (e.g., LB in FIG. 11) , and bottom-right (e.g., RB in FIG. 11) positions of the reference block are checked to find the reference block’s BVs. Both the merge candidates from the neighboring PUs and constructed AR-BVPs can be used as guiding BVs. The construction will be recursively processed until the number of AR-BVPs reaches 20 or no new AR-BVPs are constructed. Finally, up to 20 AR-BVPs can be constructed and there would be up to 30 merge candidates in total. If an AR-BVP candidate is selected for refinement, it would have a refinement range of 3x3. The same template matching cost is used to compare the merge positions and the defaults ones. For bi-directional IBC merge candidate, two candidates are retained corresponding to each reference frame. Similarly, for IntraTMP, two candidates are considered corresponding to the best candidate by template search and the coded candidate. For each block vector obtained by default search, merge or ARBVP, a refinement window of 3x3 is defined. For overlapping block vectors, a clustering process is defined. If two refinement windows overlap, the BV candidate with the lower template cost is selected and ranked in the list for further refinement. Subsequently, a new refinement window that comprises both individual refinement windows is determined for the winning candidate. Additionally, the following considered: - For block sizes 8x8 or smaller, a search region fully overlapped by a BVP candidate refinement window is skipped. - In each sparse region, the searching pattern is shifted by one sample per row in order to increase the candidate diversity. Sum of absolute differences (SAD) is used as a cost function. A given search order of the 6 regions is utilized, i.e., R4, R5, R6, R1, R2, and R3. Within each region, the decoder constructs a candidate list of up to “19” template matching block vectors that are ranked in ascending order according to the template cost (SAD) . The following modes are supported: 1-Single predictor: A single predictor is selected from the candidate list. 2-Fusion of multiple predictors: multiple predictors are blended multiple to derive the final prediction block. The blending weights are either computed from the template matching cost of each predictor, or with Wiener-filter based weight derivation method. 3-Sub-pel precision: When single predictor is used, sub-pel precisions are supported. A new candidate list is constructed by including the selected integer block vector and surrounding 1 / 2-pel and 1-4-pel sub-pel positions. The list is sorted based on the same cost function used for the integer bv search. After that, the first two candidates are allowed to be selected with one single flag being signaled from encoder to decoder. 4-linear filter model: A linear filter can be learned between the reference template and current template and be applied the linear model to reference block. This mode can be used for single predictor when sub-pel precision is not used. Additionally, IntraTMP with local illumination compensation is allowed. The following considerations are taken: 1-Usages of LIC and FLM (CCCM-like filtering) are mutually exclusive for a given CU. 2-Usages of LIC together with fusion in intra TMP is allowed. 3-Top-only and Left-only template usage for LIC model determination is allowed for screen content coding. For camera-captured coding, only the top-left template is employed. 4-Multi Mode Linear Model (MMLM) is supported similarly to IBC-LIC, for screen content coding. When LIC is used for a given CU, the Intra TMP search process employs MRSAD rather than SAD distortion function. Moreover, when LIC is used for a given CU in the Intra TMP search process, the position of the pixel at the top-left of each candidate reconstructed block is shifted by half the search sampling factor vertically and horizontally with respect to the associated position in the non-LIC case. As an exception, for a candidate reconstructed block with top-left pixel having shifted position, if the horizontal / vertical shift causes this candidate reconstructed block to go out of the bounds of the Intra TMP search region, the horizontal / vertical shift is canceled. The concept is illustrated in FIG. 8. FIG. 8 illustrates modified search range for LIC. Positions labeled with “x” are default positions whereas positions labed with “y” are the LIC positions. The new positions are shifted by half of the subsampling range. The dimensions of all regions (SearchRange_w, SearchRange_h) are set proportional to the block dimension (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = min (64, a*BlkW) SearchRange_h = min (64, a*BlkH) Where ‘a’ is a constant that controls the gain / complexity trade-off. In practice, ‘a’ is equal to 5. To speed-up the template matching process, the search range of all search regions is subsampled by a factor of 4. After finding the best match, a refinement process is performed. The refinement is done via a second template matching search around the best match with a reduced range. The Intra template matching tool is enabled for CUs with size less than or equal to 64 in width and height. This maximum CU size for Intra template matching is configurable. The Intra template matching prediction mode is signaled at CU level through a dedicated flag when DIMD is not used for current CU. 2.1.7.1 IntraTMP derived block vector candidates for IBC In this method block vector (BV) derived from the intra template matching prediction (IntraTMP) is used for intra block copy (IBC) . The stored IntraTMP BV of the neighbouring blocks along with IBC BV are used as spatial BV candidates in IBC candidate list construction. IntraTMP block vector is stored in the IBC block vector buffer and, the current IBC block can use both IBC BV and IntraTMP BV of neighbouring blocks as BV candidate for IBC BV candidate list as shown in FIG. 12. IntraTMP block vectors are added to IBC block vector candidate list as spatial candidates. IntraTMP block vectors are stored in quarter-pel resolution for coding of IBC block vectors and HMVP. 2.1.8 Fusion for template-based intra mode derivation (TIMD) For each intra prediction mode in MPMs, as well as the wide-angle modes if the above-right and / or bottom-left reference samples are available, SATD between the prediction and reconstruction samples of the template is calculated. First two intra prediction modes with the minimum SATD and one non-angular intra prediction mode (i.e. DC or Planar) with the lowest SATD cost are selected as the TIMD modes. These three TIMD modes are fused with the weights after applying PDPC process, and such weighted intra prediction is used to code the current CU. Position dependent intra prediction combination (PDPC) is included in the derivation of the TIMD modes. The conditions below are checked to determine whether the non-angular intra prediction mode is used in fusion: - the non-angular intra prediction mode is different from the two selected intra prediction modes. - costMode3 < 1.5*costMode1, where the costMode3 is the SATD cost of the non-angular intra prediction mode and costMode1 is the SATD cost of the first intra prediction mode. If both of the conditions are true, three intra prediction modes are used to generate the prediction. And the weights of each intra prediction mode are computed from SATD cost: Otherwise, the non-angular intra prediction mode is not used in prediction. And the costs of the two selected modes are compared with a threshold, in the test the cost factor of 2 is applied as follows: costMode2 < 2*costMode1. If this condition is true, the fusion is applied, otherwise the only mode1 is used. Weights of the modes are computed from their SATD costs as follows: weight1 = costMode2 / (costMode1+ costMode2) weight2 = 1 -weight1 The division operations are conducted using the same lookup table (LUT) based integerization scheme used by the CCLM. Besides, location-dependent sample-based fusion used in DIMD fusion process is used for the TIMD fusion but the location-dependent criterion applying to amplitudes of the selected predictors is replaced by a SATD cost-based criteria. The location-dependent criterion is determined from a ratio of the normalized SATD of the selected TIMD predictors computed in above and left template area. 2.1.9 Intra prediction fusion This intra prediction method derives predicted samples as a weighted combination of multiple predictors generated from different reference lines. In this process multiple intra predictors are generated and then fused by weighted averaging. The process of deriving the predictors to be used in the fusion process is described as follows: · For angular intra prediction modes including the single mode case of TIMD and DIMD, the proposed method derives intra prediction by weighting intra predictions obtained from multiple reference lines represented as pfusion=w0pline+w1pline+1, where pline is the intra prediction from the default reference line and pline+1 is the prediction from the line above the default reference line. The weights are set as w0=3 / 4 and w1=1 / 4. · For TIMD mode with blending, pline is used for the first mode (w0=1, w1=0) and pline+1 is used for the second mode (w0=0, w1=1) . · For DIMD mode with blending, the number of predictors selected for a weighted average is increased from 3 to 6. The angular intra prediction fusion method is applied to luma blocks when angular intra mode has non-integer slope (required reference samples interpolation) and the block size is greater than 16, it is used with MRL and not applied for ISP coded blocks. In the method studied in the sub-test a, PDPC is applied for the intra prediction mode using the closest to the current block reference line. The TIMD mode with blending method is applied when all the following conditions are satisfied: - both the first and second modes are angular prediction mode. - the current block is not ISP coded block. - all of the following conditions are false: ○ abs (predModeIntra1 –predModeIntra2) is greater than Threshold. The value of Threshold is set to 8 or 4 depending on block size. ○ (predModeIntra1 -EXT_HOR_IDX) * (predModeIntra2 -EXT_HOR_IDX) is less than 0. ○ (predModeIntra1 -EXT_VER_IDX) * (predModeIntra2 -EXT_VER_IDX) is less than 0. 2.1.10 Improvements of CIIP 2.1.10.1 Subblock CIIP A subblock-based merge candidate may be used to generate the inter signal of CIIP, where the same subblock-based merge candidate list used by affine and sbTMVP is utilized. When CIIP flag is true and CIIP-TM flag is false, a subblock-based CIIP flag is signalled. If subblock-based CIIP flag is true, an index indicating specific candidate in the subblock-based merge list is signalled, and TIMD is used to generate intra signal by default thus no CIIP-PDPC flag signalled any more. 2.1.10.2 Combination of CIIP with TIMD and TM merge In CIIP mode, the prediction samples are generated by weighting an inter prediction signal predicted using CIIP-TM merge candidate and an intra prediction signal predicted using TIMD derived intra prediction mode. The method is only applied to coding blocks with an area less than or equal to 1024. The TIMD derivation method is used to derive the intra prediction mode in CIIP. Specifically, the intra prediction mode with the smallest SATD values in the TIMD mode list is selected and mapped to one of the 67 regular intra prediction modes. In addition, it is also proposed to modify the weights (wIntra, wInter) for the two tests if the derived intra prediction mode is an angular mode. For near-horizontal modes (2 <= angular mode index < 34) , the current block is vertically divided; for near-vertical modes (34 <= angular mode index <= 66) , the current block is horizontally divided. FIG. 13A and FIG. 13B illustrates the division method for angular modes, respectively. The (wIntra, wInter) for different sub-blocks are shown below. Table 1. The modified weights used for angular modes. With CIIP-TM, a CIIP-TM merge candidate list is built for the CIIP-TM mode. The merge candidates are refined by template matching. The CIIP-TM merge candidates are also reordered by the ARMC method as regular merge candidates. The maximum number of CIIP-TM merge candidates is equal to two. 2.1.11 Extended multiple reference line (MRL) list MRL list in VVC is extended to include more reference lines for intra prediction. The extended reference line list consists of line indices {1, 3, 5, 7, 12} . For template-based intra mode derivation (TIMD) , instead of the full MRL candidate list, only the first two reference line candidates, i.e., {1, 3} , are used. FIG. 14 illustrates an extended MRL candidate list. 2.1.12 Template-based multiple reference line intra prediction Template-based multiple reference line intra prediction (TMRL) mode combines reference line and prediction mode together and uses a template matching method to construct a list of candidate combinations. An index to the candidate combination list is coded to indicate which reference line and prediction mode is used in coding the current block. The regular multiple reference line (MRL) for the non-TIMD part is replaced by TMRL mode. The TMRL mode extends reference line candidate list and the intra-prediction-mode candidate list. The extended reference line candidate list is {1, 3, 5, 7, 12} . The size of the intra-prediction-mode candidate list is 10. The construction of the intra-prediction-mode candidate list is similar to MPM except the PLANAR mode is excluded from the intra-prediction-mode candidate list, DC mode is added after 5 neighboring PUs’ modes and DIMD modes if its not included and the angular modes with delta angles from ±1 to ±4 (compared the existing angular modes in the intra-prediction-mode candidate list) are added. The precision of angular prediction is extended from 65 to 129. Additionally non-adjacent positions are added as candidates in constructing the intra candidate list. If the neighbouring or non-adjacent blocks are coded with SGPM or GPM modes, the intra modes of the blocks are replaced by the partitioning angles. The TMRL candidate is constructed as follows. There are 5x10=50 combinations of the extended reference line and the allowed intra-prediction modes for a block. Since the extended reference line starts from reference line 1, the area covered by reference line 0 is used for template matching. The SAD costs over the template area (see FIG. 15) are calculated between the predictions (generated by 50 combinations) and the reconstructions. The 20 combinations with the least SAD cost are selected in an ascending order to form the TMRL candidate list. For TMR signalling instead of coding the reference line and the intra mode directly, an index to the TMRL candidate list is coded to indicate which combination of reference line and prediction mode is used for coding the current block. 2.1.13 Convolutional cross-component intra prediction model In this method convolutional cross-component model (CCCM) is applied to predict chroma samples from reconstructed luma samples in a similar spirit as done by the current CCLM modes. As with CCLM, the reconstructed luma samples are down-sampled to match the lower resolution chroma grid when chroma sub-sampling is used. Similar to CCLM top, left or top and left reference samples are used as templates for model derivation. Also, similarly to CCLM, there is an option of using a single model or multi-model variant of CCCM. The multi-model variant uses two models, one model derived for samples above the average luma reference value and another model for the rest of the samples (following the spirit of the CCLM design) . Multi-model CCCM mode can be selected for PUs which have at least 128 reference samples available. 2.1.13.1 Convolutional filter The convolutional 7-tap filter consist of a 5-tap plus sign shape spatial component, a nonlinear term and a bias term. The input to the spatial 5-tap component of the filter consists of a center (C) luma sample which is collocated with the chroma sample to be predicted and its above / north (N) , below / south (S) , left / west (W) and right / east (E) neighbors as illustrated in FIG 13. The nonlinear term P is represented as power of two of the center luma sample C and scaled to the sample value range of the content: P = (C*C + midVal) >> bitDepth. That is, for 10-bit content it is calculated as: P = (C*C + 512) >> 10. The bias term B represents a scalar offset between the input and output (similarly to the offset term in CCLM) and is set to middle chroma value (512 for 10-bit content) . Output of the filter is calculated as a convolution between the filter coefficients ci and the input values and clipped to the range of valid chroma samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B. 2.1.13.2 Calculation of filter coefficients The filter coefficients ci are calculated by minimising MSE between predicted and reconstructed chroma samples in the reference area. FIG. 17 illustrates the reference area which consists of 2 or 6 lines of chroma samples above and left of the PU. Whether to use 6 lines or 2 lines of neighbouring samples to derive the CCCM model parameters in the single model CCCM is determined by a template cost. Similarly, for the multi-model CCCM mode, the two candidates use 6 lines neighbouring luma samples or luma samples collocated to the current chroma block to derive mean values which separate samples into two groups. The cost is derived by applying the candidate CCP (either 2 or 6 lines) on a template, calculating the sum of absolute difference (SAD) between CCP predicted samples and reconstructed samples in the template. Reference area extends one PU width to the right and one PU height below the PU boundaries. Area is adjusted to include only available samples. The extensions to the area shown in blue are needed to support the “side samples” of the plus shaped spatial filter and are padded when in unavailable areas. The MSE minimization is performed by calculating autocorrelation matrix for the luma input and a cross-correlation vector between the luma input and chroma output. Autocorrelation matrix is LDL decomposed and the final filter coefficients are calculated using back-substitution. The process follows roughly the calculation of the ALF filter coefficients in ECM, however LDL decomposition was chosen instead of Cholesky decomposition to avoid using square root operations. The autocorrelation matrix is calculated using the reconstructed values of luma and chroma samples. These samples are full range (e.g. between 0 and 1023 for 10-bit content) resulting in relatively large values in the autocorrelation matrix. This requires high bit depth operation during the model parameters calculation. It is proposed to remove fixed offsets from luma and chroma samples in each PU for each model. This is driving down the magnitudes of the values used in the model creation and allows reducing the precision needed for the fixed-point arithmetic. As a result, 16-bit decimal precision is proposed to be used instead of the 22-bit precision of the original CCCM implementation. Reference sample values just outside of the top-left corner of the PU are used as the offsets (offsetLuma, offsetCb and offsetCr) for simplicity. The samples values used in both model creation and final prediction (i.e., luma and chroma in the reference area, and luma in the current PU) are reduced by these fixed values, as follows: C'= C –offsetLuma N'= N –offsetLuma S'= S –offsetLuma E'= E –offsetLuma W'= W –offsetLuma P'= nonLinear (C') B = midValue = 1 << (bitDepth -1) and the chroma value is predicted using the following equation, where offsetChroma is equal to offsetCr and offsetCb for Cr and Cb components, respectively: predChromaVal = c0C'+ c1N'+ c2S'+ c3E'+ c4W'+ c5P'+ c6B + offsetChroma. In order to avoid any additional sample level operations, the luma offset is removed during the luma reference sample interpolation. This can be done, for example, by substituting the rounding term used in the luma reference sample interpolation with an updated offset including both the rounding term and the offsetLuma. The chroma offset can be removed by deducting the chroma offset directly from the reference chroma samples. As an alternative way, impact of the chroma offset can be removed from the cross-component vector giving identical result. In order to add the chroma offset back to the output of the convolutional prediction operation the chroma offset is added to the bias term of the convolutional model. The process of CCCM model parameter calculation requires division operations. Division operations are not always considered implementation friendly. The division operation are replaced with multiplication (with a scale factor) and shift operation, where scale factor and number of shifts are calculated based on denominator similar to the method used in calculation of CCLM parameters. 2.1.13.3 Gradient Linear Model For YUV 4: 2: 0 color format, a gradient linear model (GLM) method can be used to predict the chroma samples from luma sample gradients. Two modes are supported: a two-parameter GLM mode and a three-parameter GLM mode. Compared with the CCLM, instead of down-sampled luma values, the two-parameter GLM utilizes luma sample gradients to derive the linear model. Specifically, when the two-parameter GLM is applied, the input to the CCLM process, i.e., the down-sampled luma samples L, are replaced by luma sample gradients G. The other parts of the CCLM (e.g., parameter derivation, prediction sample linear transform) are kept unchanged. C=α·G+β In the three-parameter GLM, a chroma sample can be predicted based on both the luma sample gradients and down-sampled luma values with different parameters. The model parameters of the three-parameter GLM are derived from 6 rows and columns adjacent samples by the LDL decomposition based MSE minimization method as used in the CCCM. C=α0·G+α1·L+α2·β For signaling, when the CCLM mode is enabled to the current CU, one flag is signaled to indicate whether GLM is enabled for both Cb and Cr components; if the GLM is enabled, another flag is signaled to indicate which of the two GLM modes is selected and one syntax element is further signaled to select one of 4 gradient filters for the gradient calculation. · Four gradient filters are enabled for the GLM. FIG. 18 illustrates four Sobel based gradient patterns for GLM. 2.1.13.4 CCCM signalling Usage of the mode is signalled with a CABAC coded PU level flag. One new CABAC context was included to support this. When it comes to signalling, CCCM is considered a sub-mode of CCLM. That is, the CCCM flag is only signalled if intra prediction mode is LM_CHROMA. 2.1.13.5 CCCM using non-downsampled luma samples CCCM mode with 3x2 filter using non-downsampled luma samples is used, which consists of 6-tap spatial terms, four nonlinear terms and a bias term. The 6-tap spatial terms correspond to 6 neighboring luma samples (i.e., L0, L1, …, L5) around the chroma sample (i.e., C) to be predicted, the four non-linear terms are derived from the samples L0, L1, L2, and L3. FIG. 19 illustrates non-downsampled luma samples. where αi is the coefficient, β is the offset. Same to the existing CCCM design, up to 6 lines / columns of chroma samples above and left to the current CU are applied to derive the filter coefficients. The filter coefficients are derived based on the same LDL decomposition method used in CCCM. The proposed method is signaled as an additional CCCM model besides the existing one, when the CCCM is selected, one single flag is signaled and used for both two chroma components to indicate whether the default CCCM model or the proposed CCCM model is applied. Additionally, SPS signaling is introduced to indicate whether the CCCM using non-downsampled luma samples is enabled. 2.1.13.6 Block-vector guided CCCM (BVG-CCCM) When the co-located luma prediction is coded with IBC or IntraTMP in Intra slices, the BVG-CCCM mode can be used. In this mode, the block vectors of the co-located luma blocks, coded in IBC or intraTMP modes, are used to determine the reference area for calculating the CCCM parameters. The prediction is performed using uses the calculated model parameters and co-located luma samples. The BVG-CCCM mode uses an 11-tap filter for cross-component prediction as below: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P (C) + c6P (N) + c7P (S) + c8P (W) + c9P (E) + c10B. The input to the spatial 5-tap component of the filter consists of a center (C) luma sample which is collocated with the chroma sample to be predicted and its above / north (N) , below / south (S) , left / west (W) and right / east (E) neighbors as illustrated in FIG. 16. The nonlinear term P is represented as power of two of the corresponding luma sample and B is the bias term. Similar to Direct Block Vector (DBV) mode, five locations in the collocated luma block area are scanned and the associated block vectors are then used for determining the reference area for parameter calculation in BVG-CCCM method. 2.1.13.7 Gradient and Location based convolutional cross-component model (GL-CCCM) This method maps luma values into chroma values using a filter with inputs consisting of one spatial luma sample, two gradient values, two location information, a nonlinear term, and a bias term. The GL-CCCM method uses gradient and location information instead of the 4 spatial neighbor samples used in the CCCM filter. The GL-CCCM filter used for the prediction is: predChromaVal = c0C + c1Gy + c2Gx + c3Y + c4X + c5P + c6B Where Gy and Gx are the vertical and horizontal gradients, respectively, and are calculated as: Gy = (2N + NW + NE) - (2S + SW + SE) Gx = (2W + NW + SW) - (2E + NE + SE) Moreover, the Y and X are the spatial coordinates of the center luma sample. The rest of the parameters are the same as CCCM tool. The reference area for the parameter calculation is the same as CCCM method. FIG. 21 illustrates spatial samples used for GL-CCCM. The usage of the mode is signalled with a CABAC coded PU level flag. When it comes to signalling, GL-CCCM is considered a sub-mode of CCCM. That is, the GL-CCCM flag is only signalled if original CCCM flag is true. Similar to the CCCM, GL-CCCM tool has 6 modes for calculating the parameters: · Single-model GL-CCCM from above and left templates; · Single-model GL-CCCM from above template; · Single-model GL-CCCM from left template; · Multi-model GL-CCCM from above and left templates; · Multi-model GL-CCCM from above template; · Multi-model GL-CCCM from left template. The encoder performs SATD search for the 6 GL-CCCM modes along with the existing CCCM modes to find the best candidates for full RD tests. 2.1.13.8 CCCM with Multiple Downsampling Filters Multiple downsampling filters are applied to a group of reconstructed luma samples in a CCCM. The linear combination of these downsampled reconstructed samples is multiplied by derived filter coefficients to form the final chroma predictor. The horizontal or vertical location of the center luma sample are also considered in the tested model. The cross-component models shown below are tested as additional CCCM modes with a mode index signalled in the bitstream: (1) Model 1: predChroma = c0 *H (C) + c1 * G1 (C) + c2 * G2 (C) + c3 * G3 (C) + c4 * P (H (C) ) + c5 *P (G1 (C) ) + c6 *P (G2 (C) ) + c7 *X + c8 *Y + c9 *B; (2) Model 2: predChroma = c0 *H (C) + c1 * H (W) + c2 * H (E) + c3 * G1 (C) + c4 * G1 (W) + c5 *G1 (E) + c6 * P (H (C) ) + c7 * P (H (W) ) + c8 * P (H (E) ) + c9 *X + c10 *B; (3) Model 3: predChroma = c0 *H (C) + c1 * H (NE) + c2 * H (SW) + c3 * G3 (C) + c4 * G3 (NE) + c5 *G3 (SW) + c6 * P (H (C) ) + c7 * P (H (NE) ) + c8 * P (H (SW) ) + c9 *Y + c10 *B; where H (·) , G1 (·) , G2 (·) , G3 (·) are various downsampling filters, C denotes the current chroma sample position, and N, S, W, E, NE, SW are the positions around C, ci are filter coefficients, P and B are nonlinear term and bias term, and X and Y are the horizontal and vertical locations of the center luma sample with respect to the top-left coordinates of the block. FIG. 22 illustrates various downsampling filters used in cross-component models. 2.1.14 Local-Boosting Cross-Component Prediction (LB-CCP) Prediction samples of MM-CCLM / MM-CCCM can be filtered with neighbouring samples. A 3×3 low-pass filter is applied to filter prediction samples generated by MM-CCLM / MM-CCCM. For a sample at a top / left boundary, the filtering window may involve neighbouring reconstructed samples. For inner samples, the filtering window only involves prediction samples, which may be padded. A flag is signaled to indicate whether filtering is applied or not for a block coded with MM-CCLM / MM-CCCM. FIG. 23 illustrates a filter on samples of MM-CCLM / MM-CCCM. 2.1.15 Cross-Component Prediction (CCP) merge (a. k. a., non-local CCP) mode For chroma coding, a flag is signalled to indicate whether CCP mode (including the CCLM, CCCM, GLM and their variants) or non-CCP mode (conventional chroma intra prediction mode, fusion of chroma intra prediction mode) is used. If the CCP mode is selected, one more flag is signalled to indicate how to derive the CCP type and parameters, i.e., either from a CCP merge list or signalled / derived on-the-fly. A CCP merge candidate list is constructed from the spatial adjacent, temporal, spatial non-adjacent, history-based m or shifted temporal candidates. After including these candidates, default models are further included to fill the remaining empty positions in the merge list. In order to remove redundant CCP models in the list, pruning operation is applied. After constructing the list, the CCP models in the list are reordered depending on the SAD costs, which are obtained using the neighbouring template of the current block. More details are described below. Spatial adjacent and non-adjacent candidates The positions and inclusion order of the spatial adjacent and non-adjacent candidates are the same as those defined in ECM for regular inter merge prediction candidates. Temporal and shifted temporal candidates Temporal candidates are selected from the collocated picture. The position and inclusion order of the temporal candidates are the same as those defined in ECM for regular inter merge prediction candidates. The shifted temporal candidates are also selected from the collocated picture. The position of temporal candidates is shifted by a selected motion vector which is derived from motion vectors of neighboring blocks. History-based candidates A history-based table is maintained to include the recently used CCP models, and the table is reset at the beginning of each CTU row. If the current list is not full after including spatial adjacent and non-adjacent candidates, the CCP models in the history-based table are added into the list. Default candidates CCLM candidates with default scaling parameters are considered, only when the list is not full after including the spatial adjacent, spatial non-adjacent, or history-based candidates. If the current list has no candidates with the single model CCLM mode, the default scaling parameters are {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, 4 / 8, -4 / 8, 5 / 8, -5 / 8, 6 / 8} . Otherwise, the default scaling parameters are {0, the scaling parameter of the first CCLM candidate + {1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, 4 / 8, -4 / 8, 5 / 8, -5 / 8, 6 / 8} . It is noted that the LB-CCP flag is inherited from a CCP candidate in the CCP merge candidate list. A flag is signaled to indicate whether the CCP merge mode is applied or not. If CCP merge mode is applied, an index is signaled to indicate which candidate model is used by the current block. In addition, CCP merge mode is not allowed for the current chroma coding block when the current CU is coded by intra sub-partitions (ISP) with single tree, or the current chroma coding block size is less than or equal to 16. For a CCP merge coded block, one CCP-merge fusion flag is further signalled to indicate whether a fusion mode is applied. In the fusion mode, the final prediction is generated by a weighted sum of the CCP-merge prediction and either the MM-CCCM prediction or the DIMD prediction. A CCP-merge fusion type flag is further signalled if the CCP-merge fusion flag is true, to indicate whether the MM-CCCM prediction or the DIMD prediction is selected and fused with the CCP-merge prediction. 2.1.16 Decoder derived CCP mode In this method, a candidate list of cross-component prediction (CCP) modes is constructed, and to select the best candidate from the list a template cost is calculated to compare the reconstructed samples and the prediction values generated by the evaluated CCP mode. The template is shown in FIG. 24. The CCP mode list is constructed from the already existed in ECM modes by single model CCLM, single model CCCM, multi-model CCCM, single model GLCCCM, single model CCCM applied with LBCCP, and multi-model CCCM applied with LBCCP. In the second aspect of the method, various decoder-derived CCP fusion candidates are added. A fusion candidate is the combination of two CCP modes selected from the existing CCP mode lists reordered by template costs. Mode flag and a fusion flag are signalled to indicate the mode usage. 2.1.17 Spatial Geometric partitioning mode (SGPM) SGPM is an intra mode that resembles the inter coding tool of GPM, where the two prediction parts are generated from intra predicted process. In this mode, a candidate list is built with each entry containing one partition split and two intra prediction modes as shown in FIG. 25.26 partition modes and 9 of intra prediction modes are used to form the combinations. the length of the candidate list is set equal to 16. The selected candidate index is signalled. The list is reordered using template (FIG. 26) where SAD between the prediction and reconstruction of the template is used for ordering. The template size is fixed to 1. For each partition mode, an IPM list is derived for each part using the same intra-inter GPM list derivation. The IPM list size is set to 3. In the list, TIMD derived mode is replaced by 2 derived modes with horizontal and vertical orientations. The list is further augmented with block-vector based prediction candidates obtained from the adjacent and non-adjacent merge candidates coded in IntraTMP or IBC mode. The template cost is employed to select the up to 6 block vectors. The final list contains up to 9 predictors: 3 regular intra modes and up to 6 block vectors based predictors. The SGPM mode is applied with a restricted blocks size: 4<=width<=64, 4<=height<=64, width<height*8, height<width*8, width*height>=32. A PPS flag is coded to indicate whether no blending of two intra predictions is allowed. When this PPS flag is set to false, the following adaptive blending is also used for spatial GPM, where blending depth τ shown in FIG. 27 is derived as follows: · If min (width, height) ==4, 1 / 2 τ is selected; · else if min (width, height) ==8, τ is selected; · else if min (width, height) ==16, 2 τ is selected; · else if min (width, height) ==32, 4 τ is selected; · else, 8 τ is selected. Otherwise (the PPS flag is set to true) , 1 / 4 τ is always used for spatial GPM coded blocks to make sure no blending is used when SGPM block has partition angle completely horizontal or vertical, and much narrower blending width is used when SGPM block has other partition angles. It is noted that the flag is set to true in current Common Test Conditions (CTC) for the screen content videos. 2.1.18 Directional planar mode Two additional planar modes where only the horizontal interpolation or only the vertical interpolation are used to obtain the predicted samples. For planar horizontal mode, only the horizontal linear interpolation is performed based on the left reference sample and the top-right reference sample to predict the current sample as: pred (x, y) = ( (W-1-x) *rec (-1, y) + (x+1) *rec (W, -1) + (W>>1) ) >>log2 (W) . For planar vertical mode, only the vertical linear interpolation is performed based on the above reference sample and the bottom-left reference sample to predict the current sample as: pred (x, y) = ( (H-1-y) *rec (x, -1) + (y+1) *rec (-1, H) + (H>>1) ) >>log2 (H) . The transform kernel selection for planar horizontal and planar vertical mode is shown in FIG. 28. If an intra prediction mode of a current block is the planar vertical mode, the horizontal intra prediction mode is used to derive a transform kernel in MTS set and LFNST set. Also, if an intra prediction mode of a current block is the planar horizontal mode, the vertical intra prediction mode is used to derive a transform kernel in MTS set and LFNST set. 2.1.19 Direct block vector (DBV) for chroma block The direct block vector is used for chroma blocks. A flag is signaled to indicate whether a chroma block is coded using IBC mode. If one of the luma blocks in five locations shown in FIG. 29 is coded with IBC or intraTMP mode, its block vector is scaled and is used as block vector for the chroma block. Template matching is used to perform block vector scaling. 2.1.20 Extrapolation filter-based intra prediction (EIP) mode In the EIP mode, the samples in a CU are predicted from the top-left position to the bottom-right position by applying an extrapolation filter to neighboring reconstructed samples or predicted samples. The EIP mode uses a 15-tap filter for prediction as below: , where pred (x, y) is the predicted value at position (x, y) in the CU, ci is the filter coefficient, and the is the reconstructed samples or predicted samples. Predicted sample values are clipped to the range of the reference samples instead of the full sample value range. Reference sample area used for determining the range is the same that is used when generating the filter coefficients. The EIP filter can be derived from the neighboring reconstructed samples or be inherited from the previous EIP coded blocks. There are three EIP filter shapes and three types of reconstructed area supported in ECM as shown in FIG. 30 and FIG. 31, respectively. For a CU coded in the EIP mode, an EIP merge flag is signaled to indicate whether the EIP filter is inherited from previous blocks coded in EIP mode. When the EIP merge flag is true, an EIP merge list is constructed from the spatial adjacent, spatial non-adjacent, temporal and history candidates. The position and inclusion order of these candidates are the same as those used in CCP merge candidate list. An EIP merge index is further signaled to indicate which EIP merge candidate is selected. The filter shape and the filter coefficients of the selected candidate are then inherited to code the CU. When the EIP merge flag is false, the EIP filter is derived from the neighboring reconstructed samples and the relevant syntax element is signaled to indicate which one of the three types of reconstructed area and which one of the three filter shapes are used for the CU. The selected filter moves in the selected reconstructed area either horizontally or vertically with a one-pixel step to construct the auto-correlation matrix and the cross-correlation vector. The calculation of coefficients from the auto-correlation matrix and the cross-correlation vector is similar to that of CCCM with added L-2 regularization. The coefficients are computed as Where λ is the regularization parameters. L2-regularization is achieved by a diagonal matrix λI being added to the ‘ATA’ matrix. A subset of the coefficients can be relaxed (unregularized) . In this case, when the ith coefficient is relaxed, the (ith row, ith column) entry of the diagonal matrix λI is set to zero. In regularized EIP scheme: · λ=M·pEIP, where pEIP=15 is the number of filter taps in EIP. ○ M=192 when the number of input samples nSamples≤2024, and thus λ=2880. ○ M=128 otherwise, and thus λ=1920. · The bias term of EIP is relaxed (unregularized) : the bottom-right entry of the diagonal matrix is set to zero. After generating the prediction samples of the CU using the EIP filter, an intra prediction mode is derived by applying the DIMD process to the prediction samples. Specifically, a horizontal gradient and a vertical gradient are calculated for each predicted sample to build a histogram of gradient. Then the intra prediction mode corresponding to the largest histogram count is used to determine the LFNST, NSPT or MTS transform set. 2.1.21 Matrix based intra prediction replacing existing conventional intra modes A matrix of weights, which are defined for a block shape and intra mode, is introduced, those weights are multiplied by the neighbour reference template to derive the prediction samples replacing conventional intra prediction. The weights are applied to the reference samples of the L shaped causal neighborhood template as shown in the FIG. 32. The reference samples in the causal neighborhood are denoted as r, and F (x, y) is the matrix of weights. Then the prediction P (x, y) can be derived as P (x, y) = ∑k F (x, y, k) *r (k) , where k denotes the index of the reference sample in the template. In the test, this prediction is used for block size with both width and height up to 32 (except for 4x32, 32x4, 8x32 and 32x8) . The template size is 2 for blocks with both width and height up to 16 and it is only used for mode 0, 1, and (2+2*k) . For other blocks, template size is set to 1; is used for mode 0, 1, and (2+4*k) ; prediction is only performed for 16x16 positions, and the rest of the samples are generated by bilinear interpolation. For all block sizes, block shape and mode-based symmetry is used. Reference length is set to W and H for modes greater than 18 and less than 50 and set to 2*W and 2*H otherwise. 2.1.22 Modifications to matrix-based intra prediction Matrix sizes of the MIP modes are increased for the blocks with sizes up to 32x32, excluding of 4x32, 32x4, 8x32 and 32x8. The proposed matrices use the L-shaped causal template as input to generate the WxH prediction block. The prediction of a sample P (x, y) can be derived as: P (x, y) = ∑k F (x, y, k) *r (k) , Where r (k) is the kth item in the L-shaped template, and F (x, y) is the matrix weights corresponding to the position (x, y) . The size of the prediction block generated by matrix multiplication equals to the current block size. 2.1.23 Adaptive reordering of non-CCP chroma modes The non-CCP chroma intra modes are adaptively reordered with template matching. FIG. 33 illustrates the collocated luma block and adjacent chroma block. A candidate list is firstly constructed with the following modes: - DBV, DM, DIMD, Planar, horizontal, vertical, DC; - luma modes from the position of TL, TR, BL and BR in the collocated luma block; - chroma modes from the adjacent chroma blocks in the position of L', T', BL', TR'a nd TL'; - luma BVs from the positions of C, TL, TR, BL and BR in the collocated luma block. Then, these chroma modes in the candidate list are used to predict the collocated luma block and the template of the chroma block. This template comprises one top row and one left column of the chroma block. The modes in the list are reordered based on a combined SATD cost, calculated as follows: cost=8*costY+ (costCbTop+costCrTop) << (logH+2) + (costCbLeft+costCrLeft) << (logW+2) . After reordering, the first N modes in the list are kept and are allowed for the chroma block. The value of N is set to equal to 7 if at least one of the luma modes from five collocated luma positions is coded with IBC or intraTMP mode. Otherwise, the value of N is set to equal to 6. 2.2 InterCCCM InterCCCM applies the CCCM method for predicting chroma samples from reconstructed luma samples when the CU uses inter prediction or intra block copy (IBC) . FIG. 34 illustrates the decoder side of the method. The cross-component filters are derived using the prediction blocks of luma and chroma. The derived filters are applied to the reconstructed luma block and blended with the prediction blocks of chroma to produce the final chroma prediction blocks. In the blending process the filtered reconstructed luma blocks use blending weight of 0.75 and chroma prediction blocks use blending weight of 0.25. The 8-tap filter consist of 6 spatial luma samples, a nonlinear term, and a bias term. The spatial luma samples (L0, …, L5) are obtained from the luma grid selecting the 6 luma samples closest to the chroma position C without down sampling as shown in FIG. 35. The predicted chroma value is obtained as, predChromaVal = c0 L0+ c1L1 + c2L2 + c3L3 + c4L4 + c5L5 + c6 nonlinear ( (L0+L3+1) >> 1) + c7 B, where nonlinear is CCCM’s nonlinear operator and B is bias. The filter coefficients are derived using ECM’s division-free Gaussian elimination method and the necessary offsets are applied to samples prior to filter derivation. The offsets for division-free Gaussian elimination method are obtained using a four-point average of the luma and chroma prediction blocks, where the four points correspond to the top-left, top-right, bottom-left and bottom-right corners of the blocks. For filter coefficient derivation at most 256 chroma samples are used. Usage of the mode is signalled with a CABAC coded TU level flag. One new CABAC context was included to support this. The InterCCCM flag is only signalled if the TU’s luma Cbf is non-zero and the CU’s predMode is either MODE_INTER or MODE_IBC. The encoder performs an RD decision in the transform selection loop for the chroma components when luma Cbf is non-zero and the CU’s predMode is either MODE_INTER or MODE_IBC. 2.3 InterCCP merge mode The intraCCP merge mode is extended to inter coding blocks, where the final chroma inter prediction combines motion-compensation predicted signals and cross-component predicted signals derived using an inherited CCP model from a CCP merge list. A decoder derived intraCCCM candidate is inserted at the front of the interCCP candidate list. In addition to the decoder derived intraCCCM candidate, a set of CCP candidates (i.e. spatial adjacent, temporal, spatial non-adjacent, history-based, shifted temporal, and default candidates) are inherited from previous coded blocks. 2.4 Block vector guided CCCM The block vector guided CCCM (BVG-CCCM) method uses block vectors of the co-located luma blocks, coded in IBC or intraTMP modes, to determine the reference area for calculating the CCCM parameters. Then the reference area in luma and corresponding area in chroma channel is used to calculate the CCCM parameters. The prediction uses the calculated model parameters and co-located luma samples to do the CCCM prediction. FIG. 36 illustrates the reference area in BVG-CCCM method. The mode is enabled only in intra slices. Moreover, an SPS-level flag is introduced for enabling or disabling the mode. The BVG-CCCM mode uses an 11-tap filter for cross-component prediction as below: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P (C) + c6P (N) + c7P (S) + c8P (W) + c9P (E) + c10B. The input to the spatial 5-tap component of the filter consists of a center (C) luma sample which is collocated with the chroma sample to be predicted and its above / north (N) , below / south (S) , left / west (W) and right / east (E) neighbors as illustrated in FIG. 37. The nonlinear term P is represented as power of two of the corresponding luma sample and B is the bias term. Similar to Direct Block Vector (DBV) mode in ECM-9.0, five locations, as shown in FIG. 38, in collocated luma block area are scanned and the associated block vectors are then used for determining the reference area for parameter calculation in BVG-CCCM method. The mode can use block vector (s) of both IBC and intraTMP coded blocks from co-located luma area. 2.4.1 Bitstream Signalling Usage of the mode is signalled with a CABAC coded PU level flag. The BVG-CCCM flag is signalled if co-located block is coded in IBC or intraTMP modes and the cross-component index is LM_CHROMA_IDX or MMLM_CHROMA_IDX. 2.4.2 Encoder Operation The encoder performs two additional RD for the BVG-CCCM for single-model and multi-model CCCM variants. 2.5 Adaptive loop filter 2.5.1 ALF simplification removal ALF gradient subsampling and ALF virtual boundary processing are removed. Block size for classification is reduced from 4x4 to 2x2. Filter size for both luma and chroma, for which ALF coefficients are signalled, is increased to 9x9. 2.5.2 ALF with fixed filters To filter a luma sample, three different classifiers (C0, C1 and C2) and three different sets of filters (F0, F1 and F2) are used. Sets F0 and F1 contain fixed filters, with coefficients trained for classifiers C0 and C1. Coefficients of filters in F2 are signalled. Which filter from a set Fi is used for a given sample is decided by a class Ci assigned to this sample using classifier Ci. The number of bits used to represent the fractional part of a luma coefficient is adaptive from 5 to 8, inclusively. For each luma filter set, which contains up to 25 filters, a 2-bit syntax element is signalled in APS to indicate the number of bits used for the coefficients in this set. The value range of a coefficient is not changed. 2.5.3 Filtering At first, two 13x13 diamond shape fixed filters F0 and F1 are applied to derive two intermediate samples R0 (x, y) and R1 (x, y) . After that, F2 is applied to R0 (x, y) , R1 (x, y) , and neighboring samples to derive a filtered sample as where fi, j is the clipped difference between a neighboring sample and current sample R (x, y) and gi is the clipped difference between Ri-20 (x, y) and current sample. The filter coefficients ci, i=0, …21, are signalled. 2.5.4 Classification Based on directionality Di and activity aclass Ci is assigned to each 2x2 block: where MD, i represents the total number of directionalities Di. As in VVC, values of the horizontal, vertical, and two diagonal gradients are calculated for each sample using 1-D Laplacian. The sum of the sample gradients within a 4×4 window that covers the target 2×2 block is used for classifier C0 and the sum of sample gradients within a 12×12 window is used for classifiers C1 and C2. The sums of horizontal, vertical and two diagonal gradients are denoted, respectively, as and The directionality Di is determined by comparing with a set of thresholds. The directionality D2 is derived as in VVC using thresholds 2 and 4.5. For D0 and D1, horizontal / vertical edge strength and diagonal edge strength are calculated first. Thresholds Th= [1.25, 1.5, 2, 3, 4.5, 8] are used. Edge strength is 0 if otherwise, is the maximum integer such that Edge strength is 0 if otherwise, is the maximum integer such that When i.e., horizontal / vertical edges are dominant, the Di is derived by using Table 2 (a) ; otherwise, diagonal edges are dominant, the Di is derived by using Table 2 (b) . Table 2. Mapping of and to Di To obtain the sum of vertical and horizontal gradients Ai is mapped to the range of 0 to n, where n is equal to 4 for and 15 for and In an ALF_APS, up to 4 luma filter sets are signalled, each set may have up to 25 filters. 2.5.5 Alternative 2x2 ALF classifier Classification in ALF is extended with an additional alternative classifier. For a signalled luma filter set, a flag is signalled to indicate whether the alternative classifier is applied. Geometrical transformation is not applied to the alternative band classifier. When the band-based classifier is applied, the sum of sample values of a 2x2 luma block is calculated at first. Then the class index is calculated as below, class_index = (sum *25) >> (sample bit depth + 2) . 2.5.6 Residual based classifier A third classifier based on luma residual sample values. For each 2x2 luma block, the sum of absolute values of the residual samples in a neighbouring 8x8 window is calculated, and the class index is derived as: classIdx = sum >> (sample bit depth –4) . The value of classIdx is in the range of 0 to 24. The classifier usage is signalled for each luma filter set in APS. 2.5.7 CCALF with long tap filter Different from VVC wherein only luma samples are involved in CCALF, in ECM, the CCALF process uses a linear filter to filter luma sample values, luma residual samples and generate a residual correction for the chroma samples. In addition, the CCALF filter shape is constructed by 23 luma spatial taps and 5 luma residual taps, which is illustrated in FIG. 39. For a given slice, the encoder can collect the statistics of the slice, analyze them and can signal up to 16 filters through APS. The number of bits used to represent the fractional part of a CCALF coefficient can vary from 7 to 10 adaptively. In an ALF adaptation parameter set (APS) , a 2-bit syntax element is signalled for each chroma component to indicate the number of bits used for the CCALF coefficients for this component. In addition, the power of 2 constraint is removed. Chroma SAOoutput samples applied to 4 taps in a 3x3 asymmetric cross shape are added as additional inputs to CCALF, as illustrated in FIG. 40. 2.5.8 Adaptive filter shape switch and using samples before deblocking filter for adaptive loop filter Two candidate filter shapes: a diamond shape as shown in FIG. 41 and a new cross shape as shown in FIG. 42, can be adaptively selected by the luma filters in ALF. The number of coefficients of a luma filter is 22 for both the filter shapes. Please note that these 22 taps are constituted with 20 spatial taps and 2 fixed filters based taps in both shapes. In each adaptation parameter set (APS) , a shape index for the derived luma filters is signaled to decoder. Each APS contains the luma filters that are associated with the filter shape index. For each CTB, an APS index is signaled to indicate which luma filter shape is used to filter the current CTB. When filtering a luma sample, the coefficients and clip indices are also rearranged according to the corresponding filter shape. The diamond shape luma ALF is replaced by the longer filter shown in FIG. 42. The samples before deblocking filters are used as additional inputs for ALF. A final ALF sample is derived by weighting the regular ALF and the filter applied to the samples before the deblocking filter. Specifically, a filtered sample is derived as , where fi, j is the clipped difference between a neighboring sample and current sample R (x, y) , gi is the clipped difference between an intermediate sample and current sample R (x, y) and hi, j is the clipped difference between a neighboring sample before DBF and current sample R (x, y) . The filter coefficients ci, i=0, …24 are signalled. In this test, 3x3 diamond shape is applied to samples before deblocking filter. In an APS, a flag is signalled to indicate whether samples before DBF are used for ALF which is always set as true at encoder. 2.5.9 Extended Fixed-Filter-Output based Taps for ALF In ALF online-trained filters consist of 4 kinds of filter taps: spatial taps, reconstruction-before-DBF based taps, residual based taps and fixed-filter-output based taps as shown in FIG. 43A and FIG. 43B. 2.5.10 ALF with residual samples The residual samples are used as additional inputs to the ALF. A filtered sample is derived as , where ri is the clipped neighboring residual sample value and rFilteredi is the clipped residual sample filtered by the fixed-filter. For residual samples, the fixed filter reuses the offline fixed filter trained for 2.5.11 ALF residuals scaling A scaling factor is signalled in slice header, the scaling factors is applied to the difference between the ALF input and ALF output, and the scaled residual is added to the ALF input (it produces a scaled ALF filtering) . A similar scaling process is applied to NN filtering in NNVC. Different luma scaling factors may be associated with different group of class indexes, and the ALF output is derived as follows: rec’ (s) = rec (s) + (corr (s) *tab [sfi [class (s) ] ] + 4) >> 3, where ALF residual correction ‘corr (s) ’ is scaled using the scaling factor associated to the class index of the sample, ‘tab’ is the predefined LUT for mapping scaling index ‘sfi’ to scaling factor. 2.5.12 Additional fixed filter for ALF Additional fixed filter with a shape of diamond 7x7 is introduced, the filter parameters are stored at both encoder and decoder. There is no classification for the newly added fixed filter. An online filter of the proposed method is shown in FIG. 44A and FIG. 44B, where spatial taps (i.e., tap #0 ~#19) , reconstruction-before-DBF-based taps (i.e., tap #26, #27, #36) , residual-based taps (i.e., #37 ~ #38) and fixed-filter-output-based taps (i.e., tap #20 ~ #25, #34, #35) are kept the same as the ECM-8.0, and several extended taps (i.e., tap #28 ~ #33, #39) are introduced into luma online-trained filters. The reconstruction before DBF is fed into the additional fixed filter to produce the filter outputs, then these filter outputs are used as input for newly extended taps. This filter is always enabled without any filter shape switching. 2.5.13 Improved fixed filters for ALF Two Laplacian-based classifiers (one for each fixed filter) are applied to a 2x2 block. In each classifier, activity and directionality values are derived based on vertical, horizontal, and diagonal gradients using a window surrounding each 2x2 block. For each 2x2 block, the mean value of a surrounding window is calculated. Then, for each sample of this window, the difference between the sample value and the mean value is calculated. A scaling factor is determined based on the activity value derived from a Laplacian classifier. The square root of the sum of the squared differences is further quantized to C′by a scaling factor. The value of C′ is an integer between 0 and 7, inclusively. With i=0, 1, let Ci denote the classifier from the classifier of i-th fixed filter in ECM-9.0. Then the proposed class index Ci′is derived as C′i= C′*896+Ci. The total number of the fixed filters is not changed. Then a class index is determined based on the activity and directionality values. Two diamond shaped fixed filters are selected from the two filter sets by using the derived two class indices. Both fixed filters are applied to samples before DBF and ALF input, where additional diamond 9x9 filter is used for the samples before DBF. The shape of the first fixed filter applied to the ALF input samples is reduced from 13x13 to 9x9, and the shape of the second fixed filter, which is 13x13, applied to ALF input is unchanged as shown in the table below. Fixed filter f1 is applied to outputs of f0 (instead of ALF input) and samples before DBF. Finally, a signalled filter is applied to the ALF input samples, samples before the deblocking filter (DBF) , outputs of the two fixed filters, output of a gaussian filter and the residual data. 2.5.14 Chroma ALF fixed filter A classifier based on Laplacian values and variance is applied to a 2x2 chroma block. Compared to the luma classifier of a fixed filter, when calculating the activity value, the sum of the chroma vertical and horizonal Laplacian values is multiplied by 2 before scaling. Similarly, the chroma variance is multiplied by 2 before scaling. The derived class index is then used to select a fixed filter from a chroma filter set. A chroma fixed filter is applied to chroma ALF input samples in a 13x13 diamond shape and DBF input samples in a 7x7 diamond shape. The first luma classifier is applied to each 2x2 chroma block. The derived class index is then used to select a fixed filter from the luma fixed filter set related to this classifier. A fixed filter is applied to chroma ALF input sample in a 9x9 diamond shape and DBF input samples in a 9x9 diamond shape. In a signalled chroma filter, 5x5 crossing extra taps are introduced, which are applied to the fixed filter output. 2.6 ALF-CCCM FIG. 45 illustrates the decoder-side diagram of ALF-CCCM. The CCCM filters are derived using SAO / CC-SAO outputs. The CCCM filtering uses the ALF luma output samples as its input and resulting samples are blended with the SAO / CC-SAO outputs before ALF chroma processing. In the ALF-CCCM filtering scheme, each CTU is divided into non-overlapping blocks and for each block the cross-component filter coefficients are derived using the SAO / CC-SAO outputs. The output samples of luma ALF are used as input to the ALF-CCCM filtering. To obtain the final ALF-CCCM output samples, the cross-component prediction samples are blended with the SAO / CC-SAO chroma output samples. In the blending the weights are equal to 0.5 for both the SAO / CC-SAO chroma output samples and the cross-component prediction samples. The encoder decides the best block size for each CTU using a rate-distortion optimization loop. There are eight possible blocks sizes 2x2, 3x3, 4x4, 8x8, 16x16, 32x32, 64x64, 128x128. For filter derivation the blocks are extended by one sample on each side. For example, with 2x2 blocks the filter is derived using 4x4 blocks. The block extension is clipped against CTU boundaries. For each CTU, the encoder’s RDO decides the best cross-component model from eight possible models. The models are listed in Table 1, where the cardinal directions indicate co-located luma sample position in the chroma grid (north up, south down) . The nonlinear and bias terms are the same as in the 7-tap CCCM model. The CCCM solver is used for deriving the filter coefficients. The 6-tap downsampling filter (used in CCCM and CCLM) is used for mapping the co-located luma into the chroma grid. For each CTU, the choice of the block size and the choice of the cross-component model are signaled using CABAC coded flags. The proposed method and its signaling are skipped for CTUs where luma ALF is not applied. For intra-coding, a CTU may inherit both the block size and the model type from the above or left CTU. This choice is signaled using a single CABAC coded flag for each CTU when ALF-CCCM is present in the left and / or above CTU. If two CTU candidates available, another CABAC coded flag is signaled to indicate the choice. For inter-coding, a picture may inherit the block sizes and model types for all CTUs from a reference picture. The reference picture is derived from the L0 and L1 lists. Only reference pictures with ALF-CCCM present in at least one CTU are considered. The reference picture with the smallest POC distance to the current POC is selected. If activated, the picture level inheritance will skip the CTU-level signaling completely for the current picture. This choice is signaled using a single CABAC coded flag if a reference picture with at least one CTU utilizing ALF-CCCM is present. The existing operations of SAO, CC-SAO, ALF and CC-ALF of ECM-14.0 are kept unchanged with exception that ALF chroma operates on ALF-CCCM output. 2.7 Bilateral filter The filter is carried out in the sample adaptive offset (SAO) loop-filter stage, as shown in FIG. 46. The bilateral filter (BIF) , SAO and CC-SAO are using samples from deblocking as input. Each filter creates an offset per sample, and these are added to the input sample and then clipped, before proceeding to ALF. In detail, the output sample SOUT is obtained as SOUT=clip (SIN+δSAO+δCCSAO+δBIF) , where SIN is the input sample from deblocking, δBIF is the offset from the bilateral filter, δSAO is the offset from SAO and δCCSAO is offset from CC-SAO. The implementation provides the possibility for the encoder to enable or disable filtering at the CTU and slice level. The encoder takes a decision by evaluating the RDO cost. For CTUs that are filtered, the filtering process proceeds as follows. At the picture border, where samples are unavailable, the bilateral filter uses extension (sample repetition) to fill in unavailable samples. For virtual boundaries, the behavior is the same as for SAO, i.e., no filtering occurs. When crossing horizontal CTU borders, the bilateral filter can access the same samples as SAO is accessing. As an example, if the center sample S0, 0 (see FIG. 47) is located on the top line of a CTU, S-1, 1, S0, 1 and S1, 1 are read from the CTU above, just like SAO does, but S0, 2 is padded, so no extra line buffer is needed. The filter shape and the samples surrounding the center sample S0, 0 are denoted according to FIG. 47. This diamond shape in FIG. 47 is different from which used a square filter support. The BIF offset equals a sum of 12 offsets δBIF= (CTU· [∑sign (Si, j-S0, 0) ·FBIF, i, j, QP (|Si, j-S0, 0|) ] +128) >>8, (1) where FBIF, i, j, QP is based on a 26x16 LUT of 8-bit integers denoted by LUTbase, QP (for 26 QPs from 17 to 42 and 16 levels of sample differences) . Three scale factors (C1, 0, C1, 1 and C2, 0) are used to pre-compute three LUTs for three different neighbor distances (1, and 2) , i.e. : LUTi, j, QP (k) = (Ci, j·LUTbase, QP (k) +4) >>3. Further, the averaging linear interpolation is used to double the number of level for |Si, j-S0, 0| from 16 to 32, i.e. : For chroma, the number of cut off bits is decreased from 3 to 2, i.e. The TU-based scale factor CTU in formula (1) is defined as follows: where is based on the TU’s shape sizes and is based on the mean absolute difference (MAD) of the TU.Both and are calculated using LUTs. More precisely, let LUTw, h be a 2D 8×8 lookup table with non-negative 8-bit integer values, and let LUTMAD be a 1D 16-entry lookup table with non-negative 8-bit integer values. Then these scale factors are defined as follows: The MAD of a (h×w) -size TU with the channel samples denoted by si, j is defined as follows: In total, four 64-byte tables LUTw, h and four 16-byte tables LUTMAD are introduced (for luma / chroma component, for intra / inter prediction) . Note that CTU is a constant for all samples of the same channel inside one TU. According to the parameters values provided in the source code of ECM 13.0, the maximal value of CTU is 20 (chroma, intra prediction) and the maximal value of ∑FBIF, i, j, QP (|Si, j-S0, 0|) is 820, however the value obtained in the modified version of formula (1) before right bit-shifting belongs to the interval [-13812, 14068] in the worst case. Therefore, all arithmetic operations can be successfully accomplished in 15-bit singed integers. The only multiplication in formula (1) can be implemented as a multiplication of 5-bit unsigned integer and 11-bit signed integer with the result saved into 15-bit signed integer. Finally, bilateral_filter_strength is signalled in the pps and can be 0 or 1. For full strength filtering, we use exactly formula (1) . For the half-strength filtering (bilateral_filter_strength = 0) , formula (1) is slightly modified to decrease the value in two times approximately: δBIF= (CTU· [∑sign (Si, j-S0, 0) ·FBIF, i, j, QP (|Si, j-S0, 0|) ] +256) >>9. 2.8 Bilateral inloop filter on chroma Same as BIF-luma, proposed BIF-chroma is also performed in parallel with the SAO and CCSAO process as shown in FIG. 48. BIF-chroma, CCSAO and SAO use the same chroma samples produced by the deblocking filter as input and generate three offsets per chroma sample in parallel. Then these three offsets are added to the input chroma sample to obtain a sum, which is then clipped to form the final output chroma sample value. The proposed BIF-chroma provides an on / off control mechanism on CTU level and slice level. The filtering process of BIF-chroma is similar to that of BIF-luma. For a chroma sample, a 5×5 diamond shape filter is used for generating the filtering offset. The difference between the central sample and each surrounding sample is calculated first. The coefficient for each reference sample is extracted from a pre-defined look-up-table based on the calculated difference directly. The coefficients used for chroma components are retrained, different from those from BIF-luma. In the BIF-luma design, the block-level filtering strength parameter c is determined based on luma TU size and CU mode. While in the BIF-chroma design, the parameter for chroma components is determined based the chroma TU size and mode when dual-tree partitioning is enabled for the current slice and based on the corresponding luma TU size and mode when dual-tree partitioning is disabled. 2.9 Cross-Component Sample Adaptive Offset (CCSAO) Cross-component Sample Adaptive Offset (CCSAO) is used to refine reconstructed chroma samples. Similarly to SAO, the CCSAO classifies the reconstructed samples into different categories, derives one offset for each category and adds the offset to the reconstructed samples in that category. However, different from SAO which only uses one single luma / chroma component of current sample as input, the CCSAO utilizes all three components to classify the current sample into different categories. To facilitate the parallel processing, the output samples from the de-blocking filter are used as the input of the CCSAO. FIG. 49 shows the diagram of the decoding workflow when the CCSAO is applied. In CCSAO, only BO is used to enhance the quality of the reconstructed samples. For a given luma / chroma sample, three candidate samples are selected to classify the given sample into different categories: one collocated Y sample, one collocated U sample, and one collocated V sample. The sample values of these three selected samples are then classified into three different bands {bandY, bandU, bandV} , and a joint index i represents the category of the given sample. One offset is signaled and added to the reconstructed samples that fall into that category, which can be formulated as bandY= (Ycol·NY) >>BD bandU= (Ucol·NU) >>BD bandV= (Vcol·NV) >>BD i=bandY· (NU·NV) +bandU·NV+bandV C′rec=Clip1 (Crec+σCCSAO [i] ) , where {Ycol, Ucol, Vcol} are the three selected collocated samples used to classify current sample; {NY, NU, NV} are the numbers of equally divided bands applied to {Ycol, Ucol, Vcol} full range respectively; BD is the internal coding bit-depth; Crec and Cr′ec are the reconstructed samples before and after the CCSAO is applied; σCCSAO [i] is the value of the CCSAO offset applied to i-th BO category. The collocated luma sample can be chosen from 9 candidate positions, while the collocated chroma sample positions are fixed, as depicted in FIG. 50. Similar to SAO, different classifiers can be applied to different local region to further enhance the whole picture quality. The parameters for each classifier (i.e., the position of Ycol, NY, NU, NV, and offsets) are signaled at picture level, and the classifier to be used is explicitly signaled and switched at CTB level. For each classifier, the maximum of {NY, NU, NV} is set to {16, 4, 4} , and offsets are constrained to be within the range [-15, 15] . At most 4 classifiers are used per frame. SAO, Bilateral filter (BIF) and CCSAO offset are computed in parallel, added to the reconstructed chroma samples and jointly clipped, as shown in FIG. 51. Similar to the edge classifier of SAO in VVC the edge-based classifier of CCSAO also uses the four 1-D directional patterns for sample classification: horizontal, vertical, 135° diagonal and 45° diagonal, as shown in FIG. 52. FIG. 52 illustrates four 1-D directional patterns for CCSAO EO sample classification: horizontal (EO class = 0) , vertical (EO class =1) , 135° diagonal and 45° diagonal. For every 1-D pattern, each sample is classified based on the sample difference between the luma sample value labeled as “c” and its two neighbor luma samples labeled as “a” and “b” along the selected 1-D pattern.Similar to SAO, the encoder may decide the best 1-D directional pattern using the rate-distortion optimization (RDO) and signal this additional information in each classifier / set. Both the sample differences “a-c” and “b-c” are compared against a pre-defined threshold value (Th) to derive the final “class_idx” information. The encoder selects the best “Th” value from an array of pre-defined threshold values based on RDO and the index into the “Th” array is signalled. Furthermore, an additional difference between CCSAO edge-based classifier and the SAO edge classifier in VVC is that, in the former, Chroma samples use the co-located Luma samples for deriving the edge information (samples “a” , “c” and “b” are the co-located luma samples) whereas, in the later Chroma samples use its own neighboring samples for deriving the edge information. Two edge-based classifiers are supported in ECM, and the edge-based classifier is decided by RDO and is signaled in slice header. The first edge-based classifier process is formulated as follows: Ea= (a-c<0) ? (a-c< (-Th) ? 0: 1) : (a-c< (Th) ? 2: 3) – (1) Eb= (b-c<0) ? (b-c< (-Th) ? 0: 1) : (b-c< (Th) ? 2: 3) – (2) class_idx = iB *16 + Ea *4 + Eb – (3) C′rec=Clip1 (Crec+σCCSAO [class_idx] ) – (4) variable “iB” in equation (3) is derived as follows. iB= (cur ·Ncur) >>BD (or) iB= (col1 ·Ncol1) >>BD (or) iB= (col2 ·Ncol2) >>BD, – (5) wherein, sample “cur” is the current sample being processed, col1 and col2 are the co-located samples. When Luma samples are processed, col1 and col2 are the co-located Cb and Cr samples respectively. When Chroma (Cb) samples are processed, col1 and col2 are the co-located Y andCrsamples respectively. Similarly When Chroma (Cr) samples are processed, col1 and col2 are the co-located Y and Cb samples respectively. Based on RDO, encoder signals one among the samples “cur” , “col1", "col2"used in deriving the band information. The second edge-based classifier is a subset of the first edge-based classifier with less edge range divisions, and is formulated as follows: Ea= (a-c< (-Th) ? 0: 1) – (6) Eb= (b-c< (-Th) ? 0: 1) – (7) class_idx = iB *4 + Ea *2 + Eb – (8) C′rec=Clip1 (Crec+σCCSAO [class_idx] ) – (9) To reduce the signaling overhead, the CCASO offsets and classifier parameters can be inherited from previous coded pictures. A FIFO buffer is used to store the CCASO parameters. An index is signaled in slice header to indicate which candidate in the FIFO buffer is selected for the current slice. 2.9.1 Adaptive clipping A bit-depth based clipping defined as [0, 2^bitdepth-1] is applied in filtering, sampling, interpolation, weighted prediction, weighted combination, reconstruction, and other stages, to ensure that the generated prediction and reconstruction samples remain in the defined dynamic range. The range [min, max] of the clipping is signalled as delta values in picture header. For I pictures, the min / max is signalled as delta values compared to [16 * (1 << (bitdepth -8) ) , 235 * (1 << (bitdepth -8) ) ] , for other pictures, the delta values compared to the min / max values of the collocated picture are signalled. When clipping is performed in LMCS domain, the min / max values are derived by applying forward LMCS look-up table to the signalled min / max values. At encoder, the min / max values are obtained by scanning the sample values in the original picture. If motion compensated temporal filter (MCTF) is applied to a picture, the min / max values are obtained by scanning the sample values in the MCTF pre-filtered picture. The min / max values are used as the lower and upper bounds for clipping at the reconstruction stage (when adding the residual signal to the prediction) and before saving the reconstructed picture into the decoded picture buffer. The min / max values are derived from MCTF prefiltered picture. 3 Problems Below issues exist in the current video coding techniques and can be improved. 1) ALF-CCCM is introduced right before the ALF chroma stage, and the output of ALF luma is used as the input for the ALF-CCCM, which leads to latency issue and breaks the parallel processing of ALF luma and ALF chroma. 4 Detailed solutions The detailed embodiments below should be considered as examples to explain general concepts. These embodiments should not be interpreted in a narrow way. Furthermore, these embodiments can be combined in any manner. The terms “video unit” or “coding unit” or “block” may represent a picture, a slice, a tile, a coding tree block (CTB) , a coding tree unit (CTU) , a coding block (CB) , a CU, a PU, a TU, a PB, or a TB. The term “CCP” may refer to any cross-component prediction method such as any kind of luma-to-chroma prediction, chroma-to-luma prediction. It could be based on linear model, non-linear model, convolutional model, and so on. For example, it could be LM / CCLM / CCCM / MMLM / GLM / GL-CCCM / CCPmerge. It could be a type of CCP based fusion mode. The term “linear / non-linear model / filter” may refer to a filter / model wherein its coefficients are derived based on a linear regression model or a non-linear model. It is noted that the terminologies mentioned below are not limited to the specific ones defined in existing standards. Any variance of the coding tool is also applicable. 1) At least one cross-component model based (CCM-based) filter may be introduced in the loop-filter stage. a) The cross-component model may refer to at least one of the following methods: i) GLM. ii) CCLM. iii) CCCM (e.g., GL-CCCM, MDF-CCCM, NS-CCCM, etc. ) . iv) MMLM. v) MM-CCCM. vi) CCP. vii) CCP merge. viii) Any variant of the above. ix) Any combination of the above. x) In one example, the used model / combination may be signaled / derived on the fly / pre-defined. b) The filter coefficients of the cross-component model may be derived with at least one element from A and at least one element from B: i) For example, A may be (1) Luma samples before LMCS. (2) Luma samples after LMCS. (3) Luma samples before deblocking filter. (4) Luma samples after deblocking filter. (5) Luma samples before bilateral filter. (6) Luma samples after luma bilateral filter. (7) Luma samples before SAO or chromaSAO or CCSAO filter. (8) Luma samples after SAO or chromaSAO or CCSAO filter. (9) Luma samples before ALF or chromaALF or CCALF filter. (10) Luma samples after ALF or chromaALF or CCALF filter. (11) Luma samples before neural network based loop filter. (12) Luma samples after neural network based loop filter. (13) For example, the above mentioned samples may be adjacent or non-adjacent or temporal samples neighboring to the current video unit. (14) For example, the above mentioned samples may be reconstruction samples of the current video unit. ii) For example, B may be (1) Chroma samples before LMCS. (2) Chroma samples after LMCS. (3) Chroma samples before deblocking filter. (4) Chroma samples after deblocking filter. (5) Chroma samples before bilateral filter. (6) Chroma samples after chroma bilateral filter. (7) Chroma samples before SAO or chromaSAO or CCSAO filter. (8) Chroma samples after SAO or chromaSAO or CCSAO filter. (9) Chroma samples before ALF or chromaALF or CCALF filter. (10) Chroma samples after ALF or chromaALF or CCALF filter. (11) Chroma samples before neural network based loop filter. (12) Chroma samples after neural network based loop filter. (13) For example, the above mentioned samples may be adjacent or non-adjacent or temporal samples neighboring to the current video unit. (14) For example, the above mentioned samples may be reconstruction samples of the current video unit. c) The input of the cross-component model based filter application may be one or more of the followings: (1) Luma samples before LMCS. (2) Luma samples after LMCS. (3) Luma samples before deblocking filter. (4) Luma samples after deblocking filter. (5) Luma samples before bilateral filter. (6) Luma samples after luma bilateral filter. (7) Luma samples before SAO or chromaSAO or CCSAO filter. (8) Luma samples after SAO or chromaSAO or CCSAO filter. (9) Luma samples before ALF or chromaALF or CCALF filter. (10) Luma samples after ALF or chromaALF or CCALF filter. (11) Luma samples before neural network based loop filter. (12) Luma samples after neural network based loop filter. (13) Chroma samples before LMCS. (14) Chroma samples after LMCS. (15) Chroma samples before deblocking filter. (16) Chroma samples after deblocking filter. (17) Chroma samples before bilateral filter. (18) Chroma samples after chroma bilateral filter. (19) Chroma samples before SAO or chromaSAO or CCSAO filter. (20) Chroma samples after SAO or chromaSAO or CCSAO filter. (21) Chroma samples before ALF or chromaALF or CCALF filter. (22) Chroma samples after ALF or chromaALF or CCALF filter. (23) Chroma samples before neural network based loop filter. (24) Chroma samples after neural network based loop filter. (25) For example, the above mentioned samples may be reconstruction samples of the current video unit. d) The output of the cross-component model based filter may be treated as the input of one or more of the following process: i) LMCS. ii) Deblocking filter. iii) Bilateral filter (e.g., luma bilateral, or chroma bilateral filter) . iv) SAO or chromaSAO or CCSAO filter. v) ALF or chromaALF or CCALF filter. vi) neural network based loop filter. vii) alternatively, the output of the cross-component model based filter may be directly stored into the decoded picture buffer (DPB) . 2) Luma sample values before downsampling may be used as the input of the cross-component model based loop filter. i) For example, one of the CCP model may be designed as a N-tap filter, consisting of X spatial terms of luma sample values before downsampling, Y nonlinear terms and Z bias term, wherein N = X+Y+Z, and X, Y, Z are constants. 3) Luma sample values down-sampled by different downsampling filters may be used as the input of the cross-component model based loop filter. a) The down-sampling filter may be derived at decoder. b) The down-sampling filter may be signaled. 4) The basic unit of the CCM based loop filter may be adaptively determined (e.g., without signalling) . i) For example, a CTU may be divided into K subblocks and each subblock may derive its own CCM. However, the value of K may be adaptively determined based on a decoder derived method. (1) For example, K may be equal to 1 or 2 or 4 or 8 or 16 or 32 or 64. (2) For example, K may be depend on content type. (a) For example, less value of K may be determined for simple content, while greater value of K may be determined for complex content. (b) In one example, the determination may consider gradient / band / residual information of Luma samples in current CTU. (c) In one example, the determination may consider gradient / band / residual information of Chroma samples in current CTU. 5) For example, several subblocks may be grouped together and share a same type of CCM-basedfilter. a) For example, an indication of CCM-based filter type may be defined to identify which filter is used for a certain subblock. b) For example, at most K types of CCM-based filter may be used for a certain CTU. 6) For example, for a certain type CCM-based filter, the filter coefficients may be on-line computed or off-line trained. a) For example, the filter model / shape (e.g., number of filter taps, linear terms, non-linear terms, bias terms) of the certain type CCM-based filter may be fixed and used for all appropriate subblocks. b) For example, the filter coefficients of the certain type of CCM-based filter may be on-line calculated or offline trained. c) For example, the filter coefficients of the certain type of CCM-based filter may be inherited from previous coded picture / slice / tile / APS. 7) For example, for a certain type CCM-based filter, multiple models may be used. a) For example, the training samples may be classified into more than one group and each group of samples share a set of filter coefficients. For model application, the samples are separated into multiple groups according to the same rule, and each group apply its own filter coefficients to derive the final output. 8) For example, the cross-component model used in a CCM-based filter applied on one current sample may be inherit from cross-component model (s) used in a CCM-based filter applied on one previous sample. a) For example, it may be inherited from an adjacent or non-adjacent neighbouring block. b) For example, it may be inherited from a history table. 9) For example, the output of a CCM-based loop-filter may or may not blended with another chroma sample value. a) For example, whether and / or how to blend or not may be explicitly signalled in the bitstream (e.g., an indicator) . b) For example, whether and / or how to blend or not may be implicitly determined based on decoding information (e.g., block width / height, etc. ) without explicit signalling. c) For example, the blending weights may be pre-defined, such as {0.25, 0.75} , {0.75, 0.25} , {0.5, 0.5} , {0, 1} , {1, 0} , and so on, wherein the first factor may be used for the output of a CCP based loop-filter and the second factor may be used for the other chroma value to be blended.General aspects: 1) The disclosed method may be used in single tree. 2) The disclosed method may be used in dual tree. 3) The disclosed method may be used for chroma coding. 4) The disclosed method may be used for luma coding. 5) The disclosed method may be used for intra block coding. 6) The disclosed method may be used for inter block coding. 7) The disclosed method may be used for IBC block coding. 8) The disclosed method may be used in a inter (such as B or P) slice. 9) The disclosed method may be used in an intra (such as I) slice. 10) Whether to and / or how to apply the disclosed methods above may be signalled at sequence level / group of pictures level / picture level / slice level / tile group level, such as in sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / tile group header. a. For example, whether the disclosed method is applied to a sequence (or, group of pictures, etc. ) may be dependent on the SPS (or PPS, etc. ) flag. b. For example, the disclosed methods may be applied to SCC sequences only. i. For example, it may be controlled by the SPS / PPS flag. ii. For example, it may be determined based on an implicit rule which does not require a syntax element signalling (e.g., an implicit SCC content detection, etc. ) . 11) Whether to and / or how to apply the disclosed methods above may be signalled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / tile / sub-picture / other kinds of region contain more than one sample or pixel. 12) Whether to and / or how to apply the disclosed methods above may be dependent on coded information, such as block size, colour format, single / dual tree partitioning, colour component, slice / picture type.
[0110] More details of the embodiments of the present disclosure will be described below which are related to video coding. The embodiments of the present disclosure should be considered as examples to explain the general concepts and should not be interpreted in a narrow way. Furthermore, these embodiments can be applied individually or combined in any manner.
[0111] As used herein, the term “block” may represent a coding tree block (CTB) , a coding tree unit (CTU) , a coding block (CB) , a coding unit (CU) , a prediction unit (PU) , a transform unit (TU) , a prediction block (PB) , a transform block (TB) , a subblock, a tile, a slice, a subpicture, a video processing unit comprising multiple samples / pixels, and / or the like. A block may be rectangular or non-rectangular.
[0112] Fig. 53 illustrates a flowchart of a method 5300 for video processing in accordance with embodiments of the present disclosure. The method 5300 is implemented during a conversion between a current video block of a video and a bitstream of the video. As shown in FIG. 53, the method 5300 starts at block 5310, where reconstructed samples of the current video block are obtained. In some embodiments, the samples may be reconstructed based on prediction samples and residuals of the current video block.
[0113] At block 5320, at least one cross-component model based (CCM-based) filter is applied on the reconstructed samples. In some embodiments, the at least one CCM-based filter may be applied in a loop filter process of the conversion. For example, the at least one CCM-based filter may be applied in combination with at least one of the following: a deblocking filter, a sample adaptive offset (SAO) , or an adaptive loop filter (ALF) . Details of the CCM-based filter will be described in detail below.
[0114] At block 5330, the conversion is performed based on the applying. In some embodiments, the conversion may include encoding the current video block into the bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream. It should be understood that the above illustrations and / or examples are described merely for purpose of description. The scope of the present disclosure is not limited in this respect.
[0115] In view of the above, at least one CCM-based filter is applied on reconstructed samples of the current video block. Compared with the conventional solution, the proposed method can advantageously introduce a CCM-based filter (s) into the filtering stage, and improve the filtering quality. Thereby, the coding quality can be improved.
[0116] In some embodiments, the cross-component model may comprise at least one of the following: a gradient linear model (GLM) , a cross-component linear model (CCLM) , a convolutional cross-component model (CCCM) , a gradient and location based convolutional cross-component model (GL-CCCM) , a non-downsampled convolutional cross-component model (NS-CCCM) , a multiple downsample filter based convolutional cross-component model (MDF-CCCM) , a multi-model convolutional cross-component model (MM-CCCM) , a multi-model linear model (MMLM) , a cross-component prediction scheme, a CCP merge scheme. It should be understood that the possible implementation of the cross-component model described here is merely illustrative and therefore should not be construed as limiting the present disclosure in any way.
[0117] In some embodiments, the cross-component model may be indicated in the bitstream. Alternatively, the cross-component model may be determined on-the-fly. Alternatively, the cross-component model may be predetermined.
[0118] In some embodiments, coefficients of the cross-component model may be determined based on at least one first element and at least one second element, where the at least one first element comprises at least one of the following: luma samples before luma mapping with chroma scaling (LMCS) , luma samples after LMCS, luma samples before a deblocking filter, luma samples after a deblocking filter, luma samples before a bilateral filter, luma samples after a luma bilateral filter, luma samples before sample adaptive offset (SAO) , luma samples before chroma SAO, luma samples before a cross-component SAO (CCSAO) filter, luma samples after SAO, luma samples after chroma SAO, luma samples after a CCSAO filter, luma samples before adaptive loop filter (ALF) , luma sample s before chroma ALF, luma samples before a cross-component ALF (CCALF) filter, luma samples after ALF, luma samples after chroma ALF, luma samples after a CCALF filter, luma samples before neural network based loop filter, or luma samples after neural network based loop filter, and where the at least one second element comprises at least one of the following: chroma samples before LMCS, chroma samples after LMCS, chroma samples before a deblocking filter, chroma samples after a deblocking filter, chroma samples before a bilateral filter, chroma samples after a chroma bilateral filter, chroma samples before SAO, chroma samples before chroma SAO, chroma samples before a CCSAO filter, chroma samples after SAO, chroma samples after chroma SAO, chroma samples after a CCSAO filter, chroma samples before ALF, chroma samples before chroma ALF, chroma samples before a CCALF filter, chroma samples after ALF, chroma samples after chroma ALF, chroma samples after a CCALF filter, chroma samples before neural network based loop filter, or chroma samples after neural network based loop filter.
[0119] In some embodiments, the luma samples may comprise at least one of the following: spatial adjacent luma samples of the current video block, spatial non-adjacent luma samples of the current video block, temporal luma samples neighboring to the current video block, or reconstructed luma samples of the current video block. Additionally and alternatively, the chroma samples may comprise at least one of the following: spatial adjacent chroma samples of the current video block, spatial non-adjacent chroma samples of the current video block, temporal chroma samples neighboring to the current video block, or reconstructed chroma samples of the current video block.
[0120] In some embodiments, an input of the at least one CCM-based filter may comprise at least one of the following: luma samples before luma mapping with chroma scaling (LMCS) , luma samples after LMCS, luma samples before a deblocking filter, luma samples after a deblocking filter, luma samples before a bilateral filter, luma samples after a luma bilateral filter, luma samples before sample adaptive offset (SAO) , luma samples before chroma SAO, luma samples before a cross-component SAO (CCSAO) filter, luma samples after SAO, luma samples after chroma SAO, luma samples after a CCSAO filter, luma samples before adaptive loop filter (ALF) , luma samples before chroma ALF, luma samples before a cross-component ALF (CCALF) filter, luma samples after ALF, luma samples after chroma ALF, luma samples after a CCALF filter, luma samples before neural network based loop filter, luma samples after neural network based loop filter, chroma samples before LMCS, chroma samples after LMCS, chroma samples before a deblocking filter, chroma samples after a deblocking filter, chroma samples before a bilateral filter, chroma samples after a chroma bilateral filter, chroma samples before SAO, chroma samples before chroma SAO, chroma samples before a CCSAO filter, chroma samples after SAO, chroma samples after chroma SAO, chroma samples after a CCSAO filter, chroma samples before ALF, chroma samples before chroma ALF, chroma samples before a CCALF filter, chroma samples after ALF, chroma samples after chroma ALF, chroma samples after a CCALF filter, chroma samples before neural network based loop filter, or chroma samples after neural network based loop filter.
[0121] In some embodiments, the luma samples may comprise reconstructed luma samples of the current video block, and the chroma samples may comprise reconstructed chroma samples of the current video block.
[0122] In some embodiments, an output of the at least one CCM-based filter may be inputted to at least one of the following: LMCS, deblocking filter, bilateral filter, SAO, chroma SAO, CCSAO filter, ALF, chroma ALF, CCALF filter, or neural network based loop filter. Alternatively, the output of the at least one CCM-based filter may be directly stored into a decoded picture buffer (DPB) .
[0123] In some embodiments, an input of the at least one CCM-based filter may comprise luma sample values before downsampling.
[0124] In some embodiments, one of the at least one CCM-based filter may be an N-tap filter comprising X spatial terms for the luma sample values before downsampling, Y nonlinear terms and Z bias term. For example, X, Y, Z may be constants and N may be equal to a sum of X, Y, and Z.
[0125] In some embodiments, luma sample values downsampled by different downsampling filters may be allowed to be used as the input of the at least one CCM-based filter.
[0126] In some embodiments, information regarding luma sample values downsampled by which of different downsampling filters may be used as the input of the at least one CCM-based filter is determined based on predefined rule or indicated in the bitstream.
[0127] In some embodiments, at least one of the different downsampling filters may be determined at a decoder or indicated in the bitstream.
[0128] In some embodiments, a basic processing unit of the at least one CCM-based filter may be determined without being indicated in the bitstream.
[0129] In some embodiments, a coding tree unit (CTU) may be partitioned into K subblocks, each of the K subblocks may be used as the basic processing unit of the at least one CCM-based filter, and K may be a number determined based on a decoder side deriving scheme.
[0130] In some embodiments, a cross-component model may be derived for each of the K subblocks, respectively. In some embodiments, a value of K may be one of the following: 1, 2, 4, 8, 16, 32, 64, or the like. It should be understood that the specific values recited herein are intended to be examples rather than limiting the scope of the present disclosure. In some other embodiments, a value of K may be dependent on content type associated with the CTU. In some further embodiments, a value of K may be positively correlated with a content complexity associated with the CTU.
[0131] In some embodiments, the content type of the content complexity may be determined based on at least one of the following: gradient information of luma samples within the CTU, band information of luma samples within the CTU, residual information of luma samples within the CTU, gradient information of chroma samples within the CTU, band information of chroma samples within the CTU, or residual information of chroma samples within the CTU. For example, the band information may be derived either from the output of preceding processing steps (such as CCSAO, regular ALF, and SAO) or computed for the current video block via histogram analysis. By way of example rather than limitation, the theoretical luma range is partitioned into a plurality of fixed band regions, after which statistical analysis is performed to determine which band each sample of the current block falls into and count the total number of samples contained in each band, so as to derive the band information.
[0132] In some embodiments, a CTU may be partitioned into a plurality of subblocks, and more than one of the plurality of subblocks may be grouped together and share a same type of CCM-based filter.
[0133] In some embodiments, the bitstream may comprise an indication indicating a CCM-based filter type for a subblock.
[0134] In some embodiments, the maximum allowed number of CCM-based filter types for a CTU may be predetermined.
[0135] In some embodiments, for a predetermined type of CCM-based filter, at least one filter coefficient may be determined on-line. Alternatively, for a predetermined type of CCM-based filter, the at least one filter coefficient may be trained off-line. Alternatively, for a predetermined type of CCM-based filter, the at least one filter coefficient may be inherited from one of the following: a previously coded picture, a previously coded slice, a previously coded tile, or a previously coded adaptation parameter set (APS) .
[0136] In some embodiments, at least one of the following parameters of the predetermined type of CCM-based filter may be fixed and used for a set of subblocks: a filter mode, a filter shape, the number of filter taps, the number of linear terms, the number of non-linear terms, or the number of bias terms.
[0137] In some embodiments, a plurality set of filter coefficients may be allowed for a predetermined type of CCM-based filter.
[0138] In some embodiments, in a training stage of the predetermined type of CCM-based filter, training samples may be classified into more than one group of samples according to a predetermined classification rule, and each group of samples of the more than one group of samples may be used to train a set of filter coefficients.
[0139] In some embodiments, in an application stage of the predetermined type of CCM-based filter, samples to be filtered may be classified into more than one group of samples according to the predetermined classification rule, and each group of samples of the more than one group of samples may be filtered with a corresponding set of filter coefficients.
[0140] In some embodiments, a cross-component model for a CCM-based filter applied on a first sample of the current video block may be inherited from at least one cross-component model for a CCM-based filter applied on a previously filtered sample.
[0141] In some embodiments, the previously filtered sample may be within an adjacent block of the current video block or a non-adjacent block of the current video block.
[0142] In some embodiments, a cross-component model for a CCM-based filter applied on a first sample of the current video block may be inherited from a history table.
[0143] In some embodiments, an output of a CCM-based filter may not be blended with a chroma sample value. In some other embodiments, an output of a CCM-based filter may be blended with a chroma sample value.
[0144] In some embodiments, information regarding at least one of the following may be indicated in the bitstream: whether to blend an output of a CCM-based filter with a chroma sample value, or how to blend the output of the CCM-based filter with the chroma sample value.
[0145] In some other embodiments, information regarding at least one of the following may be determined without being indicated in the bitstream: whether to blend an output of a CCM-based filter with a chroma sample value, or how to blend the output of the CCM-based filter with the chroma sample value.
[0146] In some embodiments, the information may be determined based on coding information of the current video block.
[0147] In some embodiments, the coding information may comprise at least one of the following: a block width or a block height.
[0148] In some embodiments, blending weights for blending the output of the CCM-based filter with the chroma sample value may be predetermined. By way of example, the predetermined blending weights may comprise one of the following: {0.25, 0.75} , {0.75, 0.25} , {0.5, 0.5} , {0, 1} , {1, 0} , or the like. It should be understood that the specific values recited herein are intended to be examples rather than limiting the scope of the present disclosure.
[0149] In some embodiments, the chroma sample value may be a chroma sample value before the CCM-based filter.
[0150] In some embodiments, the method may be applied for at least one of the following: a single tree partition, a dual tree partition, a chroma coding, a luma coding, an inter block coding, an intra block coding, an IBC coding, an intra slice, or an inter slice.
[0151] In some embodiments, whether to and / or how to apply the method may be indicated at one of the following: a sequence level, a group of pictures level, a picture level, a slice level, or a tile group level.
[0152] In some embodiments, whether to and / or how to apply the method may be indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS) , a video parameter set (VPS) , a decoding parameter set (DPS) , a decoding capability information (DCI) , a picture parameter set (PPS) , an adaptation parameter set (APS) , a slice header, or a tile group header.
[0153] In some embodiments, whether the method is applied to a sequence may be dependent on an SPS flag. Alternatively, whether the method is applied to a group of pictures may be dependent on a PPS flag.
[0154] In some embodiments, the method may be applied to a screen content coding (SCC) sequence.
[0155] In some embodiments, whether the method is applied to a screen content coding (SCC) sequence may be controlled by an SPS flag or a PPS flag. In some other embodiments, whether the method is applied to a screen content coding (SCC) sequence may be determined based on an implicit rule which does not require syntax element signalling.
[0156] In some embodiments, whether to and / or how to apply the method may be indicated at a region containing more than one sample or pixel. By way of example, the region may comprise at least one of the following: a prediction block (PB) , a transform block (TB) , a coding block (CB) , a prediction unit (PU) , a transform unit (TU) , a coding unit (CU) , a virtual pipeline data unit (VPDU) , a coding tree unit (CTU) , a CTU row, a slice, a tile, or a sub-picture.
[0157] In some embodiments, whether to and / or how to apply the method may be dependent on coded information. By way of example, the coded information may comprise at least one of the following: a block size, a color format, a single tree partitioning, a dual tree partitioning, a color component, a slice type, or a picture type.
[0158] In view of the above, the solutions in accordance with some embodiments of the present disclosure can advantageously improve coding efficiency and coding quality.
[0159] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video which is generated by a method performed by an apparatus for video processing. The method comprises: obtaining reconstructed samples of a current video block of the video; applying at least one cross-component model based (CCM-based) filter on the reconstructed samples; and generating the bitstream based on the applying.
[0160] According to still further embodiments of the present disclosure, a method for storing bitstream of a video is provided. The method comprises: obtaining reconstructed samples of a current video block of the video; applying at least one cross-component model based (CCM-based) filter on the reconstructed samples; generating the bitstream based on the applying; and storing the bitstream in a non-transitory computer-readable recording medium.
[0161] Implementations of the present disclosure can be described in view of the following clauses, the features of all of the following clauses can be combined in any reasonable manner.
[0162] Clause 1. A method for video processing, comprising: obtaining, for a conversion between a current video block of a video and a bitstream of the video, reconstructed samples of the current video block; applying at least one cross-component model based (CCM-based) filter on the reconstructed samples; and performing the conversion based on the applying.
[0163] Clause 2. The method of clause 1, wherein the at least one CCM-based filter is applied in a loop filter process of the conversion.
[0164] Clause 3. The method of any of clauses 1-2, wherein the cross-component model comprises at least one of the following: a gradient linear model (GLM) , a cross-component linear model (CCLM) , a convolutional cross-component model (CCCM) , a gradient and location based convolutional cross-component model (GL-CCCM) , a non-downsampled convolutional cross-component model (NS-CCCM) , a multiple downsample filter based convolutional cross-component model (MDF-CCCM) , a multi-model convolutional cross-component model (MM-CCCM) , a multi-model linear model (MMLM) , a cross-component prediction scheme, a CCP merge scheme.
[0165] Clause 4. The method of clause 3, wherein the cross-component model is indicated in the bitstream, or the cross-component model is determined on-the-fly, or the cross-component model is predetermined.
[0166] Clause 5. The method of any of clauses 1-4, wherein coefficients of the cross-component model is determined based on at least one first element and at least one second element, wherein the at least one first element comprises at least one of the following: luma samples before luma mapping with chroma scaling (LMCS) , luma samples after LMCS, luma samples before a deblocking filter, luma samples after a deblocking filter, luma samples before a bilateral filter, luma samples after a luma bilateral filter, luma samples before sample adaptive offset (SAO) , luma samples before chroma SAO, luma samples before a cross-component SAO (CCSAO) filter, luma samples after SAO, luma samples after chroma SAO, luma samples after a CCSAO filter, luma samples before adaptive loop filter (ALF) , luma samples before chroma ALF, luma samples before a cross-component ALF (CCALF) filter, luma samples after ALF, luma samples after chroma ALF, luma samples after a CCALF filter, luma samples before neural network based loop filter, or luma samples after neural network based loop filter, and wherein the at least one second element comprises at least one of the following: chroma samples before LMCS, chroma samples after LMCS, chroma samples before a deblocking filter, chroma samples after a deblocking filter, chroma samples before a bilateral filter, chroma samples after a chroma bilateral filter, chroma samples before SAO, chroma samples before chroma SAO, chroma samples before a CCSAO filter, chroma samples after SAO, chroma samples after chroma SAO, chroma samples after a CCSAO filter, chroma samples before ALF, chroma samples before chroma ALF, chroma samples before a CCALF filter, chroma samples after ALF, chroma samples after chroma ALF, chroma samples after a CCALF filter, chroma samples before neural network based loop filter, or chroma samples after neural network based loop filter.
[0167] Clause 6. The method of clause 5, wherein the luma samples comprise at least one of the following: spatial adjacent luma samples of the current video block, spatial non-adjacent luma samples of the current video block, temporal luma samples neighboring to the current video block, or reconstructed luma samples of the current video block, and / or wherein the chroma samples comprise at least one of the following: spatial adjacent chroma samples of the current video block, spatial non-adjacent chroma samples of the current video block, temporal chroma samples neighboring to the current video block, or reconstructed chroma samples of the current video block.
[0168] Clause 7. The method of any of clauses 1-6, wherein an input of the at least one CCM-based filter comprises at least one of the following: luma samples before luma mapping with chroma scaling (LMCS) , luma samples after LMCS, luma samples before a deblocking filter, luma samples after a deblocking filter, luma samples before a bilateral filter, luma samples after a luma bilateral filter, luma samples before sample adaptive offset (SAO) , luma samples before chroma SAO, luma samples before a cross-component SAO (CCSAO) filter, luma samples after SAO, luma samples after chroma SAO, luma samples after a CCSAO filter, luma samples before adaptive loop filter (ALF) , luma samples before chroma ALF, luma samples before a cross-component ALF (CCALF) filter, luma samples after ALF, luma samples after chroma ALF, luma samples after a CCALF filter, luma samples before neural network based loop filter, luma samples after neural network based loop filter, chroma samples before LMCS, chroma samples after LMCS, chroma samples before a deblocking filter, chroma samples after a deblocking filter, chroma samples before a bilateral filter, chroma samples after a chroma bilateral filter, chroma samples before SAO, chroma samples before chroma SAO, chroma samples before a CCSAO filter, chroma samples after SAO, chroma samples after chroma SAO, chroma samples after a CCSAO filter, chroma samples before ALF, chroma samples before chroma ALF, chroma samples before a CCALF filter, chroma samples after ALF, chroma samples after chroma ALF, chroma samples after a CCALF filter, chroma samples before neural network based loop filter, or chroma samples after neural network based loop filter.
[0169] Clause 8. The method of clause 7, wherein the luma samples comprise reconstructed luma samples of the current video block, and the chroma samples comprise reconstructed chroma samples of the current video block.
[0170] Clause 9. The method of any of clauses 1-8, wherein an output of the at least one CCM-based filter is inputted to at least one of the following: LMCS, deblocking filter, bilateral filter, SAO, chroma SAO, CCSAO filter, ALF, chroma ALF, CCALF filter, or neural network based loop filter, or wherein the output of the at least one CCM-based filter is directly stored into a decoded picture buffer (DPB) .
[0171] Clause 10. The method of any of clauses 1-9, wherein an input of the at least one CCM-based filter comprises luma sample values before downsampling.
[0172] Clause 11. The method of clause 10, wherein one of the at least one CCM-based filter is an N-tap filter comprising X spatial terms for the luma sample values before downsampling, Y nonlinear terms and Z bias term, wherein X, Y, Z are constants and N is equal to a sum of X, Y, and Z.
[0173] Clause 12. The method of any of clauses 1-11, wherein luma sample values downsampled by different downsampling filters are allowed to be used as the input of the at least one CCM-based filter.
[0174] Clause 13. The method of clause 12, wherein information regarding luma sample values downsampled by which of different downsampling filters are used as the input of the at least one CCM-based filter is determined based on predefined rule or indicated in the bitstream.
[0175] Clause 14. The method of any of clauses 12-13, wherein at least one of the different downsampling filters is determined at a decoder or indicated in the bitstream.
[0176] Clause 15. The method of any of clauses 1-14, wherein a basic processing unit of the at least one CCM-based filter is determined without being indicated in the bitstream.
[0177] Clause 16. The method of clause 15, wherein a coding tree unit (CTU) is partitioned into K subblocks, each of the K subblocks is used as the basic processing unit of the at least one CCM-based filter, and K is a number determined based on a decoder side deriving scheme.
[0178] Clause 17. The method of clause 16, wherein a cross-component model is derived for each of the K subblocks, respectively.
[0179] Clause 18. The method of any of clauses 16-17, wherein a value of K is one of the following: 1, 2, 4, 8, 16, 32, or 64.
[0180] Clause 19. The method of any of clauses 16-18, wherein a value of K is dependent on content type associated with the CTU.
[0181] Clause 20. The method of any of clauses 16-19, wherein a value of K is positively correlated with a content complexity associated with the CTU.
[0182] Clause 21. The method of any of clauses 19-20, wherein the content type of the content complexity is determined based on at least one of the following: gradient information of luma samples within the CTU, band information of luma samples within the CTU, residual information of luma samples within the CTU, gradient information of chroma samples within the CTU, band information of chroma samples within the CTU, or residual information of chroma samples within the CTU.
[0183] Clause 22. The method of any of clauses 1-21, wherein a CTU is partitioned into a plurality of subblocks, and more than one of the plurality of subblocks is grouped together and share a same type of CCM-based filter.
[0184] Clause 23. The method of clause 22, wherein the bitstream comprises an indication indicating a CCM-based filter type for a subblock.
[0185] Clause 24. The method of any of clauses 22-23, wherein the maximum allowed number of CCM-based filter types for a CTU is predetermined.
[0186] Clause 25. The method of any of clauses 1-24, wherein for a predetermined type of CCM-based filter, at least one filter coefficient is determined on-line, or for a predetermined type of CCM-based filter, the at least one filter coefficient is trained off-line, or for a predetermined type of CCM-based filter, the at least one filter coefficient is inherited from one of the following: a previously coded picture, a previously coded slice, a previously coded tile, or a previously coded adaptation parameter set (APS) .
[0187] Clause 26. The method of clause 25, wherein at least one of the following parameters of the predetermined type of CCM-based filter is fixed and used for a set of subblocks: a filter mode, a filter shape, the number of filter taps, the number of linear terms, the number of non-linear terms, or the number of bias terms.
[0188] Clause 27. The method of any of clauses 1-26, wherein a plurality set of filter coefficients are allowed for a predetermined type of CCM-based filter.
[0189] Clause 28. The method of clause 27, wherein in a training stage of the predetermined type of CCM-based filter, training samples are classified into more than one group of samples according to a predetermined classification rule, and each group of samples of the more than one group of samples is used to train a set of filter coefficients.
[0190] Clause 29. The method of clause 28, wherein in an application stage of the predetermined type of CCM-based filter, samples to be filtered are classified into more than one group of samples according to the predetermined classification rule, and each group of samples of the more than one group of samples is filtered with a corresponding set of filter coefficients.
[0191] Clause 30. The method of any of clauses 1-29, wherein a cross-component model for a CCM-based filter applied on a first sample of the current video block is inherited from at least one cross-component model for a CCM-based filter applied on a previously filtered sample.
[0192] Clause 31. The method of clause 30 , wherein the previously filtered sample is within an adjacent block of the current video block or a non-adjacent block of the current video block.
[0193] Clause 32. The method of any of clauses 1-29, wherein a cross-component model for a CCM-based filter applied on a first sample of the current video block is inherited from a history table.
[0194] Clause 33. The method of any of clauses 1-32, wherein an output of a CCM-based filter is not blended with a chroma sample value.
[0195] Clause 34. The method of any of clauses 1-32, wherein an output of a CCM-based filter is blended with a chroma sample value.
[0196] Clause 35. The method of any of clauses 1-34, wherein information regarding at least one of the following is indicated in the bitstream: whether to blend an output of a CCM-based filter with a chroma sample value, or how to blend the output of the CCM-based filter with the chroma sample value.
[0197] Clause 36. The method of any of clauses 1-34, wherein information regarding at least one of the following is determined without being indicated in the bitstream: whether to blend an output of a CCM-based filter with a chroma sample value, or how to blend the output of the CCM-based filter with the chroma sample value.
[0198] Clause 37. The method of clause 36, wherein the information is determined based on coding information of the current video block.
[0199] Clause 38. The method of clause 37, wherein the coding information comprises at least one of the following: a block width or a block height.
[0200] Clause 39. The method of any of clauses 34-38, wherein blending weights for blending the output of the CCM-based filter with the chroma sample value is predetermined.
[0201] Clause 40. The method of clause 39, wherein the predetermined blending weights comprise one of the following: {0.25, 0.75} , {0.75, 0.25} , {0.5, 0.5} , {0, 1} , {1, 0} .
[0202] Clause 41. The method of any of clauses 33-40, wherein the chroma sample value is a chroma sample value before the CCM-based filter.
[0203] Clause 42. The method of any of clauses 1-41, wherein the method is applied for at least one of the following: a single tree partition, a dual tree partition, a chroma coding, a luma coding, an inter block coding, an intra block coding, an IBC coding, an intra slice, or an inter slice.
[0204] Clause 43. The method of any of clauses 1-42, wherein whether to and / or how to apply the method is indicated at one of the following: a sequence level, a group of pictures level, a picture level, a slice level, or a tile group level.
[0205] Clause 44. The method of any of clauses 1-43, wherein whether to and / or how to apply the method is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS) , a video parameter set (VPS) , a decoding parameter set (DPS) , a decoding capability information (DCI) , a picture parameter set (PPS) , an adaptation parameter set (APS) , a slice header, or a tile group header.
[0206] Clause 45. The method of any of clauses 1-44, wherein whether the method is applied to a sequence is dependent on an SPS flag, or whether the method is applied to a group of pictures is dependent on a PPS flag.
[0207] Clause 46. The method of clause 45, wherein the method is applied to a screen content coding (SCC) sequence.
[0208] Clause 47. The method of clause 45, wherein whether the method is applied to a screen content coding (SCC) sequence is controlled by an SPS flag or a PPS flag.
[0209] Clause 48. The method of clause 45, wherein whether the method is applied to a screen content coding (SCC) sequence is determined based on an implicit rule which does not require syntax element signalling.
[0210] Clause 49. The method of any of clauses 1-45, wherein whether to and / or how to apply the method is indicated at a region containing more than one sample or pixel.
[0211] Clause 50. The method of clause 49, wherein the region comprises at least one of the following: a prediction block (PB) , a transform block (TB) , a coding block (CB) , a prediction unit (PU) , a transform unit (TU) , a coding unit (CU) , a virtual pipeline data unit (VPDU) , a coding tree unit (CTU) , a CTU row, a slice, a tile, or a sub-picture.
[0212] Clause 51. The method of any of clauses 1-50, wherein whether to and / or how to apply the method is dependent on coded information.
[0213] Clause 52. The method of clause 51, wherein the coded information comprises at least one of the following: a block size, a color format, a single tree partitioning, a dual tree partitioning, a color component, a slice type, or a picture type.
[0214] Clause 53. The method of any of clauses 1-52, wherein the conversion includes encoding the current video block into the bitstream.
[0215] Clause 54. The method of any of clauses 1-52, wherein the conversion includes decoding the current video block from the bitstream.
[0216] Clause 55. An apparatus for video processing comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to perform a method in accordance with any of clauses 1-54.
[0217] Clause 56. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method in accordance with any of clauses 1-54.
[0218] Clause 57. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by an apparatus for video processing, wherein the method comprises: obtaining reconstructed samples of a current video block of the video; applying at least one cross-component model based (CCM-based) filter on the reconstructed samples; and generating the bitstream based on the applying.
[0219] Clause 58. A method for storing a bitstream of a video, comprising: obtaining reconstructed samples of a current video block of the video; applying at least one cross-component model based (CCM-based) filter on the reconstructed samples; generating the bitstream based on the applying; and storing the bitstream in a non-transitory computer-readable recording medium. Example Device
[0220] Fig. 54 illustrates a block diagram of a computing device 5400 in which various embodiments of the present disclosure can be implemented. The computing device 5400 may be implemented as or included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300) .
[0221] It would be appreciated that the computing device 5400 shown in Fig. 54 is merely for purpose of illustration, without suggesting any limitation to the functions and scopes of the embodiments of the present disclosure in any manner.
[0222] As shown in Fig. 54, the computing device 5400 includes a general-purpose computing device 5400. The computing device 5400 may at least comprise one or more processors or processing units 5410, a memory 5420, a storage unit 5430, one or more communication units 5440, one or more input devices 5450, and one or more output devices 5460.
[0223] In some embodiments, the computing device 5400 may be implemented as any user terminal or server terminal having the computing capability. The server terminal may be a server, a large-scale computing device or the like that is provided by a service provider. The user terminal may for example be any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, station, unit, device, multimedia computer, multimedia tablet, Internet node, communicator, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, personal communication system (PCS) device, personal navigation device, personal digital assistant (PDA) , audio / video player, digital camera / video camera, positioning device, television receiver, radio broadcast receiver, E-book device, gaming device, or any combination thereof, including the accessories and peripherals of these devices, or any combination thereof. It would be contemplated that the computing device 5400 can support any type of interface to a user (such as “wearable” circuitry and the like) .
[0224] The processing unit 5410 may be a physical or virtual processor and can implement various processes based on programs stored in the memory 5420. In a multi-processor system, multiple processing units execute computer executable instructions in parallel so as to improve the parallel processing capability of the computing device 5400. The processing unit 5410 may also be referred to as a central processing unit (CPU) , a microprocessor, a controller or a microcontroller.
[0225] The computing device 5400 typically includes various computer storage medium. Such medium can be any medium accessible by the computing device 5400, including, but not limited to, volatile and non-volatile medium, or detachable and non-detachable medium. The memory 5420 can be a volatile memory (for example, a register, cache, Random Access Memory (RAM) ) , a non-volatile memory (such as a Read-Only Memory (ROM) , Electrically Erasable Programmable Read-Only Memory (EEPROM) , or a flash memory) , or any combination thereof. The storage unit 5430 may be any detachable or non-detachable medium and may include a machine-readable medium such as a memory, flash memory drive, magnetic disk or another other media, which can be used for storing information and / or data and can be accessed in the computing device 5400.
[0226] The computing device 5400 may further include additional detachable / non-detachable, volatile / non-volatile memory medium. Although not shown in Fig. 54, it is possible to provide a magnetic disk drive for reading from and / or writing into a detachable and non-volatile magnetic disk and an optical disk drive for reading from and / or writing into a detachable non-volatile optical disk. In such cases, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0227] The communication unit 5440 communicates with a further computing device via the communication medium. In addition, the functions of the components in the computing device 5400 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 5400 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs) or further general network nodes.
[0228] The input device 5450 may be one or more of a variety of input devices, such as a mouse, keyboard, tracking ball, voice-input device, and the like. The output device 5460 may be one or more of a variety of output devices, such as a display, loudspeaker, printer, and the like. By means of the communication unit 5440, the computing device 5400 can further communicate with one or more external devices (not shown) such as the storage devices and display device, with one or more devices enabling the user to interact with the computing device 5400, or any devices (such as a network card, a modem and the like) enabling the computing device 5400 to communicate with one or more other computing devices, if required. Such communication can be performed via input / output (I / O) interfaces (not shown) .
[0229] In some embodiments, instead of being integrated in a single device, some or all components of the computing device 5400 may also be arranged in cloud computing architecture. In the cloud computing architecture, the components may be provided remotely and work together to implement the functionalities described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage service, which will not require end users to be aware of the physical locations or configurations of the systems or hardware providing these services. In various embodiments, the cloud computing provides the services via a wide area network (such as Internet) using suitable protocols. For example, a cloud computing provider provides applications over the wide area network, which can be accessed through a web browser or any other computing components. The software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote position. The computing resources in the cloud computing environment may be merged or distributed at locations in a remote data center. Cloud computing infrastructures may provide the services through a shared data center, though they behave as a single access point for the users. Therefore, the cloud computing architectures may be used to provide the components and functionalities described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.
[0230] The computing device 5400 may be used to implement video encoding / decoding in embodiments of the present disclosure. The memory 5420 may include one or more video coding modules 5425 having one or more program instructions. These modules are accessible and executable by the processing unit 5410 to perform the functionalities of the various embodiments described herein.
[0231] In the example embodiments of performing video encoding, the input device 5450 may receive video data as an input 5470 to be encoded. The video data may be processed, for example, by the video coding module 5425, to generate an encoded bitstream. The encoded bitstream may be provided via the output device 5460 as an output 5480.
[0232] In the example embodiments of performing video decoding, the input device 5450 may receive an encoded bitstream as the input 5470. The encoded bitstream may be processed, for example, by the video coding module 5425, to generate decoded video data. The decoded video data may be provided via the output device 5460 as the output 5480.
[0233] While this disclosure has been particularly shown and described with references to example embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be covered by the scope of this present application. As such, the foregoing description of embodiments of the present application is not intended to be limiting.
Claims
1.A method for video processing, comprising:obtaining, for a conversion between a current video block of a video and a bitstream of the video, reconstructed samples of the current video block;applying at least one cross-component model based (CCM-based) filter on the reconstructed samples; andperforming the conversion based on the applying.2.The method of claim 1, wherein the at least one CCM-based filter is applied in a loop filter process of the conversion.3.The method of any of claims 1-2, wherein the cross-component model comprises at least one of the following: a gradient linear model (GLM) , a cross-component linear model (CCLM) , a convolutional cross-component model (CCCM) , a gradient and location based convolutional cross-component model (GL-CCCM) , a non-downsampled convolutional cross-component model (NS-CCCM) , a multiple downsample filter based convolutional cross-component model (MDF-CCCM) , a multi-model convolutional cross-component model (MM-CCCM) , a multi-model linear model (MMLM) , a cross-component prediction scheme, a CCP merge scheme.4.The method of claim 3, wherein the cross-component model is indicated in the bitstream, orthe cross-component model is determined on-the-fly, orthe cross-component model is predetermined.5.The method of any of claims 1-4, wherein coefficients of the cross-component model is determined based on at least one first element and at least one second element,wherein the at least one first element comprises at least one of the following:luma samples before luma mapping with chroma scaling (LMCS) ,luma samples after LMCS,luma samples before a deblocking filter,luma samples after a deblocking filter,luma samples before a bilateral filter,luma samples after a luma bilateral filter,luma samples before sample adaptive offset (SAO) ,luma samples before chroma SAO,luma samples before a cross-component SAO (CCSAO) filter,luma samples after SAO,luma samples after chroma SAO,luma samples after a CCSAO filter,luma samples before adaptive loop filter (ALF) ,luma samples before chroma ALF,luma samples before a cross-component ALF (CCALF) filter,luma samples after ALF,luma samples after chroma ALF,luma samples after a CCALF filter,luma samples before neural network based loop filter, orluma samples after neural network based loop filter, andwherein the at least one second element comprises at least one of the following:chroma samples before LMCS,chroma samples after LMCS,chroma samples before a deblocking filter,chroma samples after a deblocking filter,chroma samples before a bilateral filter,chroma samples after a chroma bilateral filter,chroma samples before SAO,chroma samples before chroma SAO,chroma samples before a CCSAO filter,chroma samples after SAO,chroma samples after chroma SAO,chroma samples after a CCSAO filter,chroma samples before ALF,chroma samples before chroma ALF,chroma samples before a CCALF filter,chroma samples after ALF,chroma samples after chroma ALF,chroma samples after a CCALF filter,chroma samples before neural network based loop filter, orchroma samples after neural network based loop filter.6.The method of claim 5, wherein the luma samples comprise at least one of the following: spatial adjacent luma samples of the current video block, spatial non-adjacent luma samples of the current video block, temporal luma samples neighboring to the current video block, or reconstructed luma samples of the current video block, and / orwherein the chroma samples comprise at least one of the following: spatial adjacent chroma samples of the current video block, spatial non-adjacent chroma samples of the current video block, temporal chroma samples neighboring to the current video block, or reconstructed chroma samples of the current video block.7.The method of any of claims 1-6, wherein an input of the at least one CCM-based filter comprises at least one of the following:luma samples before luma mapping with chroma scaling (LMCS) ,luma samples after LMCS,luma samples before a deblocking filter,luma samples after a deblocking filter,luma samples before a bilateral filter,luma samples after a luma bilateral filter,luma samples before sample adaptive offset (SAO) ,luma samples before chroma SAO,luma samples before a cross-component SAO (CCSAO) filter,luma samples after SAO,luma samples after chroma SAO,luma samples after a CCSAO filter,luma samples before adaptive loop filter (ALF) ,luma samples before chroma ALF,luma samples before a cross-component ALF (CCALF) filter,luma samples after ALF,luma samples after chroma ALF,luma samples after a CCALF filter,luma samples before neural network based loop filter,luma samples after neural network based loop filter,chroma samples before LMCS,chroma samples after LMCS,chroma samples before a deblocking filter,chroma samples after a deblocking filter,chroma samples before a bilateral filter,chroma samples after a chroma bilateral filter,chroma samples before SAO,chroma samples before chroma SAO,chroma samples before a CCSAO filter,chroma samples after SAO,chroma samples after chroma SAO,chroma samples after a CCSAO filter,chroma samples before ALF,chroma samples before chroma ALF,chroma samples before a CCALF filter,chroma samples after ALF,chroma samples after chroma ALF,chroma samples after a CCALF filter,chroma samples before neural network based loop filter, orchroma samples after neural network based loop filter.8.The method of claim 7, wherein the luma samples comprise reconstructed luma samples of the current video block, and the chroma samples comprise reconstructed chroma samples of the current video block.9.The method of any of claims 1-8, wherein an output of the at least one CCM-based filter is inputted to at least one of the following: LMCS, deblocking filter, bilateral filter, SAO, chroma SAO, CCSAO filter, ALF, chroma ALF, CCALF filter, or neural network based loop filter, orwherein the output of the at least one CCM-based filter is directly stored into a decoded picture buffer (DPB) .10.The method of any of claims 1-9, wherein an input of the at least one CCM-based filter comprises luma sample values before downsampling.11.The method of claim 10, wherein one of the at least one CCM-based filter is an N-tap filter comprising X spatial terms for the luma sample values before downsampling, Y nonlinear terms and Z bias term, wherein X, Y, Z are constants and N is equal to a sum of X, Y, and Z.12.The method of any of claims 1-11, wherein luma sample values downsampled by different downsampling filters are allowed to be used as the input of the at least one CCM-based filter.13.The method of claim 12, wherein information regarding luma sample values downsampled by which of different downsampling filters are used as the input of the at least one CCM-based filter is determined based on predefined rule or indicated in the bitstream.14.The method of any of claims 12-13, wherein at least one of the different downsampling filters is determined at a decoder or indicated in the bitstream.15.The method of any of claims 1-14, wherein a basic processing unit of the at least one CCM-based filter is determined without being indicated in the bitstream.16.The method of claim 15, wherein a coding tree unit (CTU) is partitioned into K subblocks, each of the K subblocks is used as the basic processing unit of the at least one CCM-based filter, and K is a number determined based on a decoder side deriving scheme.17.The method of claim 16, wherein a cross-component model is derived for each of the K subblocks, respectively.18.The method of any of claims 16-17, wherein a value of K is one of the following: 1, 2, 4, 8, 16, 32, or 64.19.The method of any of claims 16-18, wherein a value of K is dependent on content type associated with the CTU.20.The method of any of claims 16-19, wherein a value of K is positively correlated with a content complexity associated with the CTU.21.The method of any of claims 19-20, wherein the content type of the content complexity is determined based on at least one of the following:gradient information of luma samples within the CTU,band information of luma samples within the CTU,residual information of luma samples within the CTU,gradient information of chroma samples within the CTU,band information of chroma samples within the CTU, orresidual information of chroma samples within the CTU.22.The method of any of claims 1-21, wherein a CTU is partitioned into a plurality of subblocks, and more than one of the plurality of subblocks is grouped together and share a same type of CCM-based filter.23.The method of claim 22, wherein the bitstream comprises an indication indicating a CCM-based filter type for a subblock.24.The method of any of claims 22-23, wherein the maximum allowed number of CCM-based filter types for a CTU is predetermined.25.The method of any of claims 1-24, wherein for a predetermined type of CCM-based filter, at least one filter coefficient is determined on-line, orfor a predetermined type of CCM-based filter, the at least one filter coefficient is trained off-line, orfor a predetermined type of CCM-based filter, the at least one filter coefficient is inherited from one of the following: a previously coded picture, a previously coded slice, a previously coded tile, or a previously coded adaptation parameter set (APS) .26.The method of claim 25, wherein at least one of the following parameters of the predetermined type of CCM-based filter is fixed and used for a set of subblocks: a filter mode, a filter shape, the number of filter taps, the number of linear terms, the number of non-linear terms, or the number of bias terms.27.The method of any of claims 1-26, wherein a plurality set of filter coefficients are allowed for a predetermined type of CCM-based filter.28.The method of claim 27, wherein in a training stage of the predetermined type of CCM-based filter, training samples are classified into more than one group of samples according to a predetermined classification rule, and each group of samples of the more than one group of samples is used to train a set of filter coefficients.29.The method of claim 28, wherein in an application stage of the predetermined type of CCM-based filter, samples to be filtered are classified into more than one group of samples according to the predetermined classification rule, and each group of samples of the more than one group of samples is filtered with a corresponding set of filter coefficients.30.The method of any of claims 1-29, wherein a cross-component model for a CCM-based filter applied on a first sample of the current video block is inherited from at least one cross-component model for a CCM-based filter applied on a previously filtered sample.31.The method of claim 30 , wherein the previously filtered sample is within an adjacent block of the current video block or a non-adjacent block of the current video block.32.The method of any of claims 1-29, wherein a cross-component model for a CCM-based filter applied on a first sample of the current video block is inherited from a history table.33.The method of any of claims 1-32, wherein an output of a CCM-based filter is not blended with a chroma sample value.34.The method of any of claims 1-32, wherein an output of a CCM-based filter is blended with a chroma sample value.35.The method of any of claims 1-34, wherein information regarding at least one of the following is indicated in the bitstream:whether to blend an output of a CCM-based filter with a chroma sample value, orhow to blend the output of the CCM-based filter with the chroma sample value.36.The method of any of claims 1-34, wherein information regarding at least one of the following is determined without being indicated in the bitstream:whether to blend an output of a CCM-based filter with a chroma sample value, orhow to blend the output of the CCM-based filter with the chroma sample value.37.The method of claim 36, wherein the information is determined based on coding information of the current video block.38.The method of claim 37, wherein the coding information comprises at least one of the following: a block width or a block height.39.The method of any of claims 34-38, wherein blending weights for blending the output of the CCM-based filter with the chroma sample value is predetermined.40.The method of claim 39, wherein the predetermined blending weights comprise one of the following: {0.25, 0.75} , {0.75, 0.25} , {0.5, 0.5} , {0, 1} , {1, 0} .41.The method of any of claims 33-40, wherein the chroma sample value is a chroma sample value before the CCM-based filter.42.The method of any of claims 1-41, wherein the method is applied for at least one of the following: a single tree partition, a dual tree partition, a chroma coding, a luma coding, an inter block coding, an intra block coding, an IBC coding, an intra slice, or an inter slice.43.The method of any of claims 1-42, wherein whether to and / or how to apply the method is indicated at one of the following:a sequence level,a group of pictures level,a picture level,a slice level, ora tile group level.44.The method of any of claims 1-43, wherein whether to and / or how to apply the method is indicated in one of the following:a sequence header,a picture header,a sequence parameter set (SPS) ,a video parameter set (VPS) ,a decoding parameter set (DPS) ,a decoding capability information (DCI) ,a picture parameter set (PPS) ,an adaptation parameter set (APS) ,a slice header, ora tile group header.45.The method of any of claims 1-44, wherein whether the method is applied to a sequence is dependent on an SPS flag, orwhether the method is applied to a group of pictures is dependent on a PPS flag.46.The method of claim 45, wherein the method is applied to a screen content coding (SCC) sequence.47.The method of claim 45, wherein whether the method is applied to a screen content coding (SCC) sequence is controlled by an SPS flag or a PPS flag.48.The method of claim 45, wherein whether the method is applied to a screen content coding (SCC) sequence is determined based on an implicit rule which does not require syntax element signalling.49.The method of any of claims 1-45, wherein whether to and / or how to apply the method is indicated at a region containing more than one sample or pixel.50.The method of claim 49, wherein the region comprises at least one of the following:a prediction block (PB) ,a transform block (TB) ,a coding block (CB) ,a prediction unit (PU) ,a transform unit (TU) ,a coding unit (CU) ,a virtual pipeline data unit (VPDU) ,a coding tree unit (CTU) ,a CTU row,a slice,a tile, ora sub-picture.51.The method of any of claims 1-50, wherein whether to and / or how to apply the method is dependent on coded information.52.The method of claim 51, wherein the coded information comprises at least one of the following:a block size,a color format,a single tree partitioning,a dual tree partitioning,a color component,a slice type, ora picture type.53.The method of any of claims 1-52, wherein the conversion includes encoding the current video block into the bitstream.54.The method of any of claims 1-52, wherein the conversion includes decoding the current video block from the bitstream.55.An apparatus for video processing comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to perform a method in accordance with any of claims 1-54.56.A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method in accordance with any of claims 1-54.57.A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by an apparatus for video processing, wherein the method comprises:obtaining reconstructed samples of a current video block of the video;applying at least one cross-component model based (CCM-based) filter on the reconstructed samples; andgenerating the bitstream based on the applying.58.A method for storing a bitstream of a video, comprising:obtaining reconstructed samples of a current video block of the video;applying at least one cross-component model based (CCM-based) filter on the reconstructed samples;generating the bitstream based on the applying; andstoring the bitstream in a non-transitory computer-readable recording medium.