Method and device for video processing and medium

The CCP model uses rows or columns of non-adjacent samples to convert and predict video blocks, which solves the problem of improving encoding and decoding efficiency in the existing technology and achieves more efficient video encoding and decoding.

CN120826902APending Publication Date: 2025-10-21DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480016439.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-04
Filing Date
2024-03-01
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing video coding and decoding technologies have room for improvement in coding and decoding efficiency, especially in cross-component prediction. It is difficult to effectively use non-adjacent samples for prediction to improve coding and decoding efficiency.

Method used

A Cross-Component Prediction (CCP) model is proposed to improve the coding and decoding efficiency by converting and predicting video blocks based on rows or columns of non-adjacent samples.

Benefits of technology

The CCP model improves the effectiveness and efficiency of video encoding and decoding, and enhances the performance of video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120826902A_ABST
    Figure CN120826902A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. In the method, for a conversion between a current video block of the video and a bitstream of the video, a cross-component prediction (CCP) model for the current video block is determined based on at least one row of samples non-adjacent to the current video block, the at least one row comprising at least one of a row of non-adjacent samples, or a column of non-adjacent samples. The transformation is performed based on the CCP model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to a cross-component prediction (CCP) model. Background Art

[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is provided. The method includes: determining a cross-component prediction (CCP) model for a current video block of a video based on at least one row of samples non-adjacent to the current video block, the at least one row comprising at least one of the following: rows of non-adjacent samples or columns of non-adjacent samples; and performing the conversion based on the CCP model. The method according to the first aspect of the present disclosure enables determining the CCP model based on rows or columns of samples non-adjacent to the current video block. Thus, codec effectiveness and codec efficiency can be improved.

[0005] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.

[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, the instructions causing a processor to execute the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining a cross-component prediction (CCP) model for a current video block of the video based on at least one row of samples that are non-adjacent to the current video block of the video, the at least one row comprising at least one of the following: rows of non-adjacent samples or columns of non-adjacent samples; and generating a bitstream based on the CCP model.

[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining a cross-component prediction (CCP) model for a current video block of the video based on at least one row of samples that are non-adjacent to the current video block, the at least one row comprising at least one of the following: rows of non-adjacent samples or columns of non-adjacent samples; generating a bitstream based on the CCP model; and storing the bitstream in a non-transitory computer-readable recording medium.

[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0011] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;

[0012] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;

[0013] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;

[0014] Figure 4 The nominal vertical and horizontal positions of the 4:2:2 luminance and chrominance samples in the picture are shown;

[0015] Figure 5 An example of an encoder block diagram is shown;

[0016] Figure 6 67 intra prediction modes are shown;

[0017] Figure 7 The reference samples used for wide-angle intra prediction are shown;

[0018] Figure 8 The problem of discontinuity is shown in the case of orientations exceeding 45°;

[0019] Figure 9 The positions of the sample points used to derive α and β are shown;

[0020] Figure 10 An example of classifying neighboring points into two groups is shown;

[0021] Figure 11A is a diagram showing the definition of sample points used by PDPC applied to a diagonal upper right mode;

[0022] Figure 11B is a diagram showing the definition of sample points used by PDPC applied to a diagonal lower left mode;

[0023] Figure 11C is a diagram showing the definition of sample points used by PDPC applied to an adjacent diagonal upper right pattern;

[0024] Figure 11D is a diagram showing the definition of sample points used by PDPC applied to an adjacent diagonal lower left pattern;

[0025] Figure 12 is a schematic diagram illustrating a gradient method for non-vertical / non-horizontal patterns;

[0026] Figure 13 is a diagram showing the value of nScale relative to nTbH and the number of modes; for all cases where nScale<0, the gradient method is used;

[0027] Figure 14 is a schematic diagram showing a flow chart of the current PDPC and the proposed PDPC;

[0028] Figure 15 is a schematic diagram showing neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list;

[0029] Figure 16 is a schematic diagram showing an example of the proposed intra reference mapping;

[0030] Figure 17 is a diagram showing an example of four reference rows adjacent to a prediction block;

[0031] Figure 18A is a diagram showing examples of sub-partitioning for 4×8 and 8×4 CUs;

[0032] Figure 18B is a diagram showing an example of sub-partitioning for a CU other than 4×8, 8×4, and 4×4;

[0033] Figure 19 is a schematic diagram illustrating a matrix-weighted intra prediction process;

[0034] Figure 20 is a schematic diagram showing target samples, template samples, and reference samples of the template used in DIMD;

[0035] Figure 21 is a schematic diagram illustrating the proposed intra-block decoding process;

[0036] Figure 22 is a schematic diagram showing the calculation of HoG from a template with a width of 3 pixels;

[0037] Figure 23 is a schematic diagram illustrating prediction fusion by weighted averaging of two HoG modes and a plane;

[0038] Figure 24 is a schematic diagram showing the spatial portion of a convolutional filter;

[0039] Figure 25 is a schematic diagram showing a reference area (with its filling) for deriving filter coefficients;

[0040] Figure 26 is a schematic diagram showing four Sobel-based gradient modes for GLM;

[0041] Figure 27 is a schematic diagram showing spatial sample points for GL-CCCM;

[0042] Figure 28 is a schematic diagram showing luma samples that have not been downsampled;

[0043] Figure 29 The spatial GPM candidates are shown;

[0044] Figure 30 A GPM template is shown;

[0045] Figure 31 GPM mixing is shown;

[0046] Figure 32 shows the binarization of the cross-component prediction mode in ECM, where Figure 32 "CCLM" can be replaced by "CCCM";

[0047] Figure 33 shows an example of luminance samples to be prepared;

[0048] Figure 34 Examples of potential candidate regions (shared blocks) are shown;

[0049] Figures 35A to 35C Possible templates are shown respectively;

[0050] Figure 36 Various downsampling filters used in the proposed cross-component model are shown;

[0051] Figure 37 The positions of the chroma samples are shown;

[0052] Figure 38 The filters on the samples of MM-CCLM / MM-CCCM are shown;

[0053] Figure 39 A template for a block is shown;

[0054] Figures 40A to 40C Possible templates are shown respectively;

[0055] Figure 41 Adjacent neighboring blocks are shown;

[0056] Figure 42 The specific row of sample points used to derive the CCLM model is shown;

[0057] Figure 43 shows the sample points of a particular row used to derive the CCLM model including the upper left sample point;

[0058] Figure 44 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and

[0059] Figure 45 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.

[0060] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION

[0061] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.

[0062] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0063] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.

[0064] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0065] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment

[0066] Figure 1 is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0067] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0068] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.

[0069] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.

[0070] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.

[0071] Figure 2 is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.

[0072] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0073] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.

[0074] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0075] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.

[0076] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0077] The mode selection unit 203 can, for example, select one of a plurality of coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).

[0078] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.

[0079] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.

[0080] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.

[0081] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0082] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.

[0083] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0084] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0085] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0086] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0087] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0088] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0089] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.

[0090] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0091] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0092] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0093] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0094] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.

[0095] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0096] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.

[0097] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which motion information includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.

[0098] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.

[0099] Motion compensation unit 302 may calculate interpolated values ​​for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.

[0100] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.

[0101] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0102] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.

[0103] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview The present disclosure relates to video coding techniques. Specifically, it relates to cross-component prediction. It can be applied to existing video coding standards such as HEVC or Versatile Video Codec (VVC). It can also be applied to future video coding standards or video codecs. 2. Introduction Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, ISO / IEC developed MPEG-1 and MPEG-4 Vision, and the two organizations jointly developed the H.262 / MPEG-2 Video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard with the goal of reducing bit rate by 50% compared to HEVC. 2.1. Color Space and Chroma Downsampling A color space, also called a color model (or color system), is an abstract mathematical model that simply describes the range of colors as a tuple of numbers, usually 3 or 4 values ​​or color components (e.g., RGB). Basically, a color space is a refinement of a coordinate system and subspace. For video compression, the most commonly used color spaces are YCbCr and RGB. YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luma component, and CB and CR are the blue-difference and red-difference chroma components. Y' (with a prime) is distinguished from Y, which is luma, meaning that light intensity is encoded nonlinearly based on the gamma-corrected RGB primaries. Chroma downsampling is the practice of encoding an image at a lower resolution for chroma information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to differences in color than to luminance. 2.1.1.4:4:4 Each of the three Y'CbCr components has the same sampling rate, so there is no chroma downsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2.4:2:2 The two chroma components are sampled at half the sampling rate of luma: the horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third with little or no visual difference. Examples of nominal vertical and horizontal positions for a 4:2:2 color format are given in the VVC working draft. Figure 4 Depicted in. Figure 4 The nominal vertical and horizontal positions of the 4:2:2 luma and chroma samples in the picture are shown. 2.1.3.4:2:0 In 4:2:0, horizontal sampling is doubled compared to 4:1:1, but vertical resolution is halved because the Cb and Cr channels are sampled only on alternate lines. Therefore, the data rate remains the same. Cb and Cr are downsampled by a factor of 2 both horizontally and vertically. There are three variants of the 4:2:0 scheme, with different horizontal and vertical positioning. In MPEG-2, Cb and Cr are co-located in the horizontal direction. Cb and Cr are located between pixels in the vertical direction (at the interstitial position). In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located in interstitial positions, in the middle of alternate luma samples. In 4:2:0 DV, Cb and Cr are co-located horizontally. Vertically, Cb and Cr are co-located on alternate lines. Table 1 SubWidthC and SubHeightC values ​​derived from chroma_format_idc and separate_colour_plane_flag chroma_format_idc separate_colour_plane_flag Chroma format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1 2.2. Encoding and decoding process of typical video codecs Figure 5 An example of a VVC encoder block diagram is shown, which contains three loop filtering blocks: deblocking filter (DF), sample adaptive offset (SAO), and ALF. Unlike DF, which uses a predefined filter, SAO and ALF use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offset and applying a finite impulse response (FIR) filter, respectively, and using the encoded side information to signal the offset and filter coefficients. ALF is located at the last processing stage of each picture and can be seen as a tool that attempts to capture and repair artifacts caused by previous stages. Intra-mode codec with 67 intra-prediction modes In order to capture arbitrary edge directions presented in natural videos, such as Figure 6As shown, the number of directional intra modes is extended from 33 to 65 in HEVC. Figure 6 67 intra prediction modes are shown, and planar and DC modes remain unchanged. These more dense directional intra prediction modes apply to all block sizes and both luma and chroma intra prediction. In HEVC, each intra-coded block has a square shape, and the length of each side is a power of 2. Therefore, no division operation is required to generate intra prediction values ​​using DC mode. In VVC, blocks can have a rectangular shape, and in general, a division operation is required for each block. To avoid division operations for DC prediction, only the longer side is used to calculate the average value of non-square blocks. 2.3.1. Wide-angle intra prediction Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index also depends on the block shape. Conventional angular intra prediction directions are defined to be from 45 degrees to -135 degrees in a clockwise direction. In VVC, for non-square blocks, multiple conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes. The replaced modes are signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e. 67, and the intra mode encoding and decoding method remains unchanged. Figure 7 To support these prediction directions, a top reference with a length of 2W+1 and a left reference with a length of 2H+1 are defined, as shown in Figure 7 shown. The number of modes replaced in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 2. Table 2 Intra-frame prediction modes replaced by wide-angle mode Figure 8 The problem of discontinuity is shown in the case of directions exceeding 45°. Figure 8 As shown in Figure 2, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the increased gap Δp. αnegative impact. If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that meet this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted using these modes, the samples in the reference buffer are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of the non-fractional mode in the regular prediction mode with the wide-angle mode. In VVC, in addition to 4:2:0, 4:2:2 and 4:4:4 chroma formats are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, expanding the number of entries from 35 to 67 to align with the expansion of intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the range of luma intra prediction modes from 2 to 5 is mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values ​​of the entries of the mapping table to more accurately convert the prediction angles for chroma blocks. 2.4. Intra-frame prediction mode encoding and decoding for chroma components For the chroma component of an intra PU, the encoder selects the best chroma prediction mode from five modes, including planar, DC, horizontal, vertical, and direct copy of the intra prediction mode for the luma component. The mapping between the intra prediction direction of chroma and the intra prediction mode number is shown in Table 3. When the intra prediction mode number for the chroma component is 4, the intra prediction direction for the luma component is used for intra prediction sample generation for the chroma component. When the intra prediction mode number for the chroma component is not 4 and is the same as the intra prediction mode number for the luma component, the intra prediction direction of 66 is used for intra prediction sample generation for the chroma component. 2.5. Inter-frame prediction For each inter-predicted CU, the motion parameters consist of a motion vector, a reference picture index and a reference picture list usage index, as well as additional information required by the new codec features of VVC that will be used for inter-prediction sample generation. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector deltas or reference picture indices. A Merge mode is specified whereby the motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates, as well as the additional scheduling introduced in VVC. Merge mode can be applied to any inter-predicted CU, not just skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index and reference picture list usage flag for each reference picture list, as well as other required information, are explicitly signaled for each CU. 2.6. Intra-block copy (IBC) Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that IBC significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed inside the current picture. The luminance block vector of the CU encoded and decoded by IBC is in integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. The CU encoded and decoded by IBC is regarded as a third prediction mode different from the intra prediction mode or the inter prediction mode. The IBC mode is applicable to CUs whose width and height are both less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD check on blocks with a width or height of no more than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed. In a hash-based search, the hash key matching (32-bit CRC) between the current block and the reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on a 4x4 sub-block. For larger current block sizes, a hash key is determined to match the hash key of a reference block when all hash keys of all 4x4 sub-blocks match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the smallest cost is selected. In the block matching search, the search range is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using a flag, and it can be signaled as IBC AMVP mode or IBC Skip / Merge mode as shown below: -IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and pairwise candidates. – IBC AMVP mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the upper neighbor (if encoded with IBC). When either neighbor is unavailable, the default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index. 2.7. Cross-component linear model prediction To reduce cross-component redundancy, a cross-component linear model (CCLM) prediction mode is used in VVC. For this CCLM prediction mode, chroma samples are predicted based on the reconstructed luma samples of the same CU using the following linear model: pred C (i,j)=α·rec L ′(i,j)+β (2-1) where pred C (i, j) represents the predicted chroma sample in CU, and rec L (i, j) represents the downsampled reconstructed luma samples of the same CU. The CCLM parameters (α and β) are derived using up to four adjacent chroma samples and their corresponding downsampled luma samples. Assuming the current chroma block dimensions are WxH, W' and H' are set to – When LM mode is applied, W'=W, H'=H; – When LM_T mode is applied, W'=W+H; – When LM_L mode is applied, H'=H+W. The upper adjacent position is represented as S[0,-1]…S[W'-1,-1], and the left adjacent position is represented as S[-1,0]…S[-1,H'-1]. Then the four sample points are selected as – When LM mode is applied and both the upper and left neighboring samples are available, S[W' / 4,-1], S[3*W' / 4,-1], S[-1,H' / 4], S[-1,3*H' / 4]; – When LM_T mode is applied or only upper neighboring samples are available, S[W' / 8,-1], S[3*W' / 8,-1], S[5*W' / 8,-1], S[7*W' / 8,-1]; – When LM_L mode is applied or only left neighbor samples are available, S[-1,H' / 8], S[-1,3*H' / 8], S[-1,5*H' / 8], S[-1,7*H' / 8]. The four adjacent brightness samples at the selected position are downsampled and compared four times to find the two larger values: x 0 A and x 1 A , and two smaller values: x 0 B and x 1 B Their corresponding chroma sample values ​​are represented by y 0 A 、y 1 A 、y 0 B and y 1 B Then x A 、x B 、y A and y B is derived as: X a =(x 0 A +x 1 A +1)>>1;X b =(x 0 B +x 1 B +1)>>1;Y a =(y 0 A +y 1 A +1)>>1;Y b =(y 0B +y 1 B +1)>>1 (2-2). Finally, the linear model parameters α and β are obtained according to the following formula. β=Y b -α·X b (2-4). Figure 9 An example of the positions of the left sample points and the upper sample points involved in the CCLM mode and the samples of the current block is shown. Figure 9 The locations of the sample points used to derive α and β are shown. The division operation for calculating the parameter α and is implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented by exponential notation. For example, diff is approximated using a 4-bit significant part and an exponent. Thus, the table for 1 / diff is reduced to 16 elements for the 16 values ​​of the significant number, as shown below: DivTable[]={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0} (2-5). This will have the advantage of reducing the complexity of the calculations and the memory size required to store the required tables. In addition to the upper template and the left template being used together to calculate the linear model coefficients, they can also be alternatively used in the other two LM modes (called LM_T and LM_L modes). In LM_T mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is expanded to (W+H) samples. In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is expanded to (H+W) samples. In LM mode, the left template and the upper template are used to calculate the linear model coefficients. To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters are applied to the luma samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The choice of downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "Type 0" and "Type 2" content, respectively. Note that when the upper reference line is at a CTU boundary, only one luma line (common line buffer in intra prediction) is used to make the downsampled luma samples. This parameter calculation is performed as part of the decoding process, not just as an encoder search operation. As a result, no syntax is used to convey the α and β values ​​to the decoder. For chroma intra mode coding and decoding, a total of 8 intra modes are allowed for chroma intra mode coding and decoding. These modes include five conventional intra modes and three cross-component linear model modes (LM, LM_T and LM_L). The chroma mode signaling and derivation process are shown in Table 3. Chroma mode coding and decoding depends directly on the intra prediction mode of the corresponding luminance block. Since the separate block partitioning structure of the luminance component and the chrominance component in the I strip is enabled, one chroma block can correspond to multiple luminance blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chroma block is directly inherited. Table 3 Derivation of chroma prediction mode from luma mode when CCLM is enabled As shown in Table 4, a single binarization table is used regardless of the value of sps_cclm_enabled_flag. Table 4 Unified binarization table for chroma prediction mode The value of intra_chroma_pred_mode Binary string 4 00 0 0100 1 0101 2 0110 3 0111 5 10 6 110 7 111 In Table 4, the first binary bit indicates whether it is normal (0) or LM mode (1). If it is LM mode, the next binary bit indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next 1 binary bit indicates whether it is LM_L (0) or LM_T (1). For this case, when sps_cclm_enabled_flag is 0, the first binary bit of the binarization table corresponding to intra_chroma_pred_mode can be discarded before entropy coding. Or, in other words, the first binary bit is inferred to be 0 and therefore not coded. This single binarization table is used for the cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two binary bits in Table 4 are context coded using their own context model, and the remaining binary bits are bypass coded. Additionally, to reduce luma-chroma latency in dual trees, when a 64x64 luma codec tree node is split using Not Split (and ISP is not used for 64x64 CUs) or QT, the chroma CUs in the 32x32 / 32x16 chroma codec tree nodes are allowed to use CCLM in the following manner: – If a 32x32 chroma node is not split or is split by a QT partition, all chroma CUs in the 32x32 node can use CCLM. If a 32x32 chroma node is split horizontally with BT, and the 32x16 child node is not split or uses vertical BT, then all chroma CUs in the 32x16 chroma node can use CCLM. Under all other luma and chroma codec tree partitioning conditions, CCLM is not allowed for chroma CUs. 2.8. Multi-model Linear Model (MMLM) With MMLM, there can be more than one linear model between the luma samples and chroma samples in a CU. In this method, the neighboring luma samples and chroma samples of the current block are classified into several groups, each of which is used as a training set to derive a linear model (i.e., a specific α and β are derived for a specific group). In addition, the samples of the current luma block are also classified based on the same rules as the classification of the neighboring luma samples. Neighboring samples can be classified into M groups, where M is 2 or 3. In addition to the original LM mode, the MMLM method with M = 2 and M = 3 is designed as two additional chroma prediction modes, called MMLM2 and MMLM3. The encoder selects the best mode during the RDO process and transmits it through the signal. When M is equal to 2, Figure 10 An example of classifying neighboring samples into two groups is shown. The threshold is calculated as the average value of the neighboring reconstructed luminance samples. Rec'L[x,y]<=threshold Rec' L Neighboring samples with [x,y]≤threshold are classified into group 1; and Rec′ L [x,y]>threshold Rec'L[x,y]>threshold adjacent samples are classified into group 2. Similar to CCLM, there are 3 modes in MMLM, namely MMLM, MMLM_T and MMLM_L. The two models are derived as The threshold is the average of the reconstructed luminance of neighboring samples. A linear model for each class is derived using the Least Mean Squares (LMS) method (if enabled) or the Min / Max method of VVC. 2.9. Position-dependent intra prediction combination In VVC, the results of intra prediction for DC, planar, and multiple angle modes are further modified by the Position Dependent Intra Prediction Combination (PDPC) method. PDPC is an intra prediction method that calls for a combination of boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, intra angles less than or equal to horizontal, and intra angles greater than or equal to vertical and less than or equal to 80. PDPC is not applied if the current block is in BDPCM mode or the MRL index is greater than 0. The prediction sample pred(x', y') is predicted using the intra prediction mode (DC, planar, angular) and a linear combination of the reference samples according to the following equation 2-8: pred(x',y')=Clip(0,(1< <BitDepth)–1,(wL×R -1,y '+wT×R x ' ,-1 +(64-wL-wT)×pred (x', y')+32)>>6) (2-9) where R x,-1 、R -1,y Respectively represent the reference sample points located at the top and left boundaries of the current sample point (x, y). If PDPC is applied to DC, planar, horizontal and vertical intra modes, additional boundary filters are not required as required in the case of HEVC DC mode boundary filters or horizontal / vertical mode edge filters. The PDPC process is the same for DC mode and planar mode. For angular mode, if the current angular mode is HOR_IDX or VER_IDX, the left reference sample or top reference sample is not used respectively. The PDPC weights and scaling factors depend on the prediction mode and the size of the block. PDPC is applied to blocks with width and height both greater than or equal to 4. Figure 11A is a diagram showing the definition of samples used by PDPC applied to the diagonal upper right mode. Figure 11B is a diagram showing the definition of sample points used by PDPC applied to the diagonal bottom-left mode. Figure 11C is a diagram showing the definition of samples used by PDPC applied to the adjacent diagonal upper-right pattern. Figure 11D is a diagram showing the definition of samples used by PDPC applied to the adjacent diagonal bottom-left pattern. 11A to 11D The reference samples (R x,-1 and R -1,y ) definition. The prediction sample point pred(x', y') is located at (x', y') in the prediction block. For example, the reference sample point R x,-1 The coordinate x of is given by the following formula: x=x'+y'+1, and the reference point R -1,y The coordinate y of is similarly given by the following formula: y = x' + y' + 1 (for diagonal mode). For other angle modes, the reference point R x,-1 and R -1,y Can be located at fractional sample positions. In this case, the sample value at the nearest integer sample position is used. Gradient PDPC like Figure 12 As shown in Figure 1, the gradient-based method is extended for non-vertical / non-horizontal modes. Here, the gradient is calculated as r(-1,y)–r(-1+d,-1), where d is the horizontal displacement depending on the angular direction. A few points should be noted here: The gradient term r(-1,y)–r(-1+d,-1) needs to be calculated once for each row since it does not depend on the x position. The calculation of d is already part of the original intra prediction process and can be reused, so there is no need to calculate d separately. Therefore, d has 1 / 32 pixel accuracy. When d is in fractional position, two-tap (linear) filtering is used, i.e., if dPos is the displacement with 1 / 32 pixel precision, dInt is the (floored) integer part (dPos>>5), and dFract is the fractional part with 1 / 32 pixel precision (dPos>31), then r(-1+d) is calculated as: r(-1+d)=(32–dFrac)*r(-1+dInt)+dFrac*r(-1+dInt+1). As explained in a, this 2-tap filtering is performed once for each row (if necessary). Finally, the prediction signal is calculated. p(x,y)=Clip(((64–wL(x))*p(x,y)+wL(x)*(r(-1,y)-r(-1+d,-1))+32)>>6) Where wL(x)=32>>((x<<1)>>nScale2), and nScale2=(log2(nTbH)+log2(nTbW)–2)>>2, which is the same as the vertical / horizontal mode. In short, the same process is applied compared to the vertical / horizontal mode (in fact, d=0 indicates the vertical / horizontal mode). Secondly, when (nScale < 0) or when PDPC cannot be applied due to the unavailability of secondary reference samples, the gradient-based approach is activated for non-vertical / non-horizontal modes. The value of nScale relative to the TB size and angle mode is as follows Figure 13 As shown, to better visualize the use of gradient method. Additionally, in Figure 14 In FIG, flow charts for the current PDPC and the proposed PDPC are shown. Figure 14 is a schematic diagram showing the flow chart of the current PDPC (left) and the proposed PDPC (right). 2.11. Secondary MPM The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list includes 16 entries. First, a general MPM list with 22 entries is constructed, and then the first 6 entries in the general MPM list are included in the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the plane mode. The remaining entries are composed of the following: Figure 15 Shown are the intra modes for the left (L), above (A), below left (BL), above right (AR), and above left (AL) neighboring blocks, the directional modes with offsets added from the first two available directional modes for the neighboring blocks, and the default mode. If the CU block is vertically oriented, the order of neighboring blocks is A, L, BL, AR, AL; otherwise, the order of neighboring blocks is L, A, BL, AR, AL. The PMPM flag is parsed first, if equal to 1, the PMPM index is parsed to determine which entry of the PMPM list is selected, otherwise the SPMPM flag is parsed to determine whether to parse the SMPM index or the remaining modes. 2.12.6 Tap Intra-frame Interpolation Filter In order to improve the prediction accuracy, a 6-tap interpolation filter is proposed to replace the 4-tap cubic interpolation filter. The filter coefficients are derived based on the same polynomial regression model, but the polynomial order is 6. The filter coefficients are shown below, {0,0,256,0,0,0}, / / 0 / 32 position {0,-4,253,9,-2,0}, / / 1 / 32 position {1,-7,249,17,-4,0}, / / 2 / 32 position {1,-10,245,25,-6,1}, / / 3 / 32 position {1,-13,241,34,-8,1}, / / 4 / 32 position {2,-16,235,44,-10,1}, / / 5 / 32 position {2,-18,229,53,-12,2}, / / 6 / 32 position {2,-20,223,63,-14,2}, / / 7 / 32 position {2,-22,217,72,-15,2}, / / 8 / 32 position {3,-23,209,82,-17,2}, / / 9 / 32 position {3,-24,202,92,-19,2}, / / 10 / 32 position {3,-25,194,101,-20,3}, / / 11 / 32 position {3,-25,185,111,-21,3}, / / 12 / 32 position {3,-26,178,121,-23,3}, / / 13 / 32 position {3,-25,168,131,-24,3}, / / 14 / 32 position {3,-25,159,141,-25,3}, / / 15 / 32 position {3,-25,150,150,-25,3}, / / half pixel position The reference samples used for interpolation come from reconstructed samples or are padded as in HEVC, so that a conditional check of reference sample availability is not necessary. It is proposed to use a 4-tap cubic interpolation filter instead of using the nearest integer operation to derive the extended intra-frame reference samples. Figure 16 As shown in the example in , a four-tap interpolation filter is used to derive the value of the reference sample point P, while in JEM-3.0 or HM, P is directly set to X1. 2.13. Multiple Reference Line (MRL) Intra Prediction Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. Figure 17 In

[15] , an example of 4 reference lines is depicted, where the samples of segments A and F are not obtained from reconstructed neighboring samples, but are filled with the closest samples from segments B and E, respectively. HEVC intra picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 2) are used. The index of the selected reference row (mrl_idx) is signaled and used to generate intra prediction values. For reference row indices greater than 0, only additional reference row modes are included in the MPM list, and only the MPM index is signaled, while the remaining modes are not signaled. The reference row index is signaled before the intra prediction mode, and if a non-zero reference row index is signaled, planar mode is excluded from the intra prediction mode. MRL is disabled for the first row of blocks inside a CTU to prevent the use of extended reference samples outside the current CTU row. In addition, PDPC is disabled when additional rows are used. For MRL mode, the derivation of DC values ​​in DC intra prediction mode for non-zero reference row indices is aligned with the derivation for reference row index 0. MRL requires storage of 3 adjacent luma reference rows along with the CTU to generate the prediction. The Cross Component Linear Model (CCLM) tool also requires 3 adjacent luma reference rows for its downsampling filter. The definition of MRL using the same 3 rows is aligned with CCLM to reduce the storage requirements of the decoder. 2.14. Intra-frame sub-segmentation (ISP) Intra sub-partitioning (ISP) divides the luma intra prediction block into 2 or 4 sub-partitions vertically or horizontally depending on the block size. For example, the minimum block size for ISP is 4x8 (or 8x4). If the block size is larger than 4x8 (or 8x4), the corresponding block is divided into 4 sub-partitions. It has been noted that M×128 (with M≤64) and 128×N (with N≤64) ISP blocks may generate potential problems with 64×64 VDPU. For example, an M×128 CU in the single-tree case has an M×128 luma TB and two corresponding Chroma TB. If the CU uses ISP, the luma TB will be divided into four M×32TBs (only horizontal division is possible), each of which is smaller than a 64×64 block. However, in the current ISP design, the chroma blocks are not divided. Therefore, both chroma components will have a size larger than a 32×32 block. Similarly, a 128×NCU using ISP can cause a similar situation. Therefore, these two situations are problems for a 64×64 decoder pipeline. For this reason, the CU size that can use ISP is limited to a maximum of 64×64. Figure 18A and Figure 18B Examples of two possibilities are shown. All sub-segments satisfy the condition of having at least 16 samples. In ISP, dependencies of 1xN / 2xN sub-block predictions on reconstructed values ​​of previously decoded 1xN / 2xN sub-blocks of the codec block are disallowed, resulting in a minimum prediction width of 4 samples for each sub-block. For example, an 8xN (N>4) codec block encoded using ISP with vertical partitioning is partitioned into two prediction regions of size 4xN each and four transforms of size 2xN. Furthermore, a 4xN codec block encoded using ISP with vertical partitioning is predicted using a full 4xN block; four 1xN transforms are used. While both 1xN and 2xN transform sizes are allowed, transforms for these blocks within a 4xN region can be performed in parallel. For example, when a 4xN prediction region contains four 1xN transforms, no transform is performed horizontally; the transform in the vertical direction can be performed as a single 4xN transform in the vertical direction. Similarly, when a 4xN prediction region contains two 2xN transform blocks, the transform operations for the two 2xN blocks in each direction (horizontally and vertically) can be performed in parallel. Therefore, processing these smaller blocks does not increase latency compared to processing intra blocks for 4x4 regular codecs. Figure 18A is a diagram showing examples of sub-partitioning for 4x8 and 8x4 CUs. Figure 18B is a diagram illustrating an example of sub-partitioning for a CU other than 4x8, 8x4, and 4x4. Table 5 Entropy coding and decoding coefficient group size Block size Coefficient group size 1×N,N≥16 1×16 N×1,N≥16 16×1 2×N,N≥8 2×8 N×2,N≥8 8×2 All other possible M×N situations 4×4 For each sub-partition, the reconstructed samples are obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated by processes such as entropy decoding, inverse quantization and inverse transformation. Therefore, the reconstructed sample values ​​of each sub-partition can be used to generate the prediction of the next sub-partition, and each sub-partition is processed repeatedly. In addition, the first sub-partition to be processed is the sub-partition containing the upper left samples of the CU, and then continues downwards (horizontal partitioning) or to the right (vertical partitioning). As a result, the reference samples used to generate the sub-partition prediction signal are only located to the left and above the row. All sub-partitions share the same intra mode. The following is an overview of the interaction of ISP with other codec tools. – Multiple Reference Line (MRL): If a block has an MRL index other than 0, the ISP codec mode will be inferred to be 0, so the ISP mode information will not be sent to the decoder. – Entropy coding coefficient group size: The size of the entropy coding sub-blocks has been modified so that they have 16 samples in all possible cases, as shown in Table 5. Note that the new size only affects blocks generated by ISP where one dimension is less than 4 samples. In all other cases, the coefficient group remains 4×4 in dimension. –CBF codec: It is assumed that at least one subpartition has a non-zero CBF. Thus, if n is the number of subpartitions and the first n-1 subpartitions have yielded zero CBF, the CBF of the nth subpartition is assumed to be 1. – Transform size restriction: All ISP transforms with length greater than 16 points use DCT-II. –MTS flag: If the CU uses ISP codec mode, the MTS CU flag will be set to 0 and will not be sent to the decoder. Therefore, the encoder will not perform RD tests for the different available transforms for each resulting sub-split. Instead, the transform selection for ISP mode will be fixed and selected based on the utilized intra mode, processing order, and block size. Therefore, no signaling is required. For example, let t H Hehet V and are the horizontal and vertical transforms selected for the w×h and sub-partitions, respectively, where w and are the widths, and h and are the heights. The transforms are then selected according to the following rules: If w=1 or h=1, there is no horizontal transform or vertical transform, respectively. – If w ≥ 4 and w ≤ 16, then t H =DST-VII, otherwise, t H =DCT-II. – If h ≥ 4 and h ≤ 16, then t V =DST-VII, otherwise, t V =DCT-II. In ISP mode, all 67 intra prediction modes are allowed. PDPC is also applied if the corresponding width and height are at least 4 samples long. In addition, the conditions for reference sample filtering (reference smoothing) and intra interpolation filter selection no longer exist, and in ISP mode, the cubic (DCT-IF) filter is always applied for fractional position interpolation. 2.15. Matrix Weighted Intra Prediction (MIP) The matrix weighted intra prediction (MIP) method is a newly added intra prediction technology in VVC. In order to predict the samples of a rectangular block of width W and height H, the matrix weighted intra prediction (MIP) takes a row of H reconstructed neighboring boundary samples on the left side of the block and a row of W reconstructed neighboring boundary samples above the block as input. If the reconstructed samples are not available, the reconstructed samples are generated in the same way as in conventional intra prediction. The generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication and linear interpolation, as shown in Figure 19 shown. 2.15.1. Averaging neighboring points Among the boundary samples, four or eight samples are selected by averaging based on the block size and shape. Specifically, the boundary bdry is input by averaging the adjacent boundary samples according to a predefined rule depending on the block size. top Hehe bdry left and is reduced to a smaller bound and Then, the two shrinking boundaries and Spliced ​​to the reduced boundary vector bdry red , so for a block of shape 4×4, bdry red Size is 4, and for all other block shapes, bdry red The size is 8. If mode refers to MIP mode, the splicing is defined as follows: Matrix Multiplication Using the averaged samples as input, a matrix-vector multiplication is performed, followed by the addition of an offset. The result is a scaled-down prediction signal on the downsampled set of samples in the original block. Based on the scaled-down input vector bdry red and, the reduced prediction signal pred red , and is generated, pred red , and the width is W red The sum and height are H red Here, W red and H red is defined as: Reduced prediction signal pred red It is calculated by taking the matrix-vector product and adding the offset: pred red =A·bdry red +b (2-13) Here, if W=H=4, then A is a red ·H red rows and 4 columns, and in all other cases a matrix with 8 columns. b is a matrix of size W red ·H red The matrix A and the offset vector b are taken from one of the sets S0, S1, and S2. The index idx = idx(W,H) is defined as follows: Here, each coefficient of the matrix A is represented with 8 bits of precision. Set S0 consists of 16 matrices and 16 offset vectors Each matrix has 16 rows and 4 columns, and each offset vector has a size of 16. This set of matrices and offset vectors is used for blocks of size 4×4. Set S1 consists of 8 matrices and 8 offset vectors Each matrix has 16 rows and 8 columns, and each offset vector has a size of 16. Set S2 consists of 6 matrices and 6 offset vectors Each matrix has 64 rows and 8 columns, and each offset vector has a size of 64. 2.15.3. Interpolation The prediction signals at the remaining positions are generated from the prediction signals on the downsampled set by linear interpolation, which is a single-step linear interpolation in each direction. Regardless of the block shape or block size, the interpolation is performed first in the horizontal direction and then in the vertical direction. 2.15.4.MIP Mode Signaling and Coordination with Other Codec Tools For each codec unit (CU) in intra mode, a flag indicating whether the MIP mode is to be applied is sent. If the MIP mode is to be applied, the MIP mode (predModeIntra) is signaled. For the MIP mode, a transposed flag (isTransposed) that determines whether the mode is transposed, and a MIP mode Id (modeId) that determines which matrix to use for a given MIP mode are derived as shown below isTransposed=predModeIntra&1 modeId=predModeIntra>>1 (2-15) The MIP codec mode is coordinated with other codecs by taking into account the following aspects: – For MIPs on large blocks, LFNST is enabled. Here, the planar LFNST transform is used. - Reference sample derivation for MIP is performed exactly the same as for regular intra prediction modes. – For the upsampling step used in MIP prediction, the original reference samples are used instead of the downsampled reference samples. – Clipping is performed before upsampling, rather than after upsampling. – Regardless of the maximum transform size, MIPs are allowed to be at most 64x64. For sizeId=0, the number of MIP patterns is 32, for sizeId=1, the number of MIP patterns is 16, and for sizeId=2, the number of MIP patterns is 12. 2.16. Decoder-side intra-mode derivation In JEM-2.0, intra modes are expanded from 35 in HEVC to 67 modes, and they are derived at the encoder and explicitly signaled to the decoder. In JEM-2.0, a significant amount of overhead is spent on intra mode encoding and decoding. For example, in a full intra codec configuration, the intra mode signaling overhead can be as high as 5-10% of the total bitrate. This contribution proposes a decoder-side intra mode derivation method to reduce the intra mode encoding and decoding overhead while maintaining prediction accuracy. To reduce the overhead of intra mode signaling, this contribution proposes a decoder-side intra mode derivation (DIMD) approach. In the proposed approach, instead of explicitly signaling the intra mode, the information is derived from neighboring reconstructed samples of the current block at both the encoder and the decoder. The intra mode derived by DIMD is used in two ways: 1) For a 2Nx2N CU, when the corresponding CU-level DIMD flag is turned on, DIMD mode is used as the intra mode for intra prediction; 2) For NxN CU, DIMD mode is used to replace a candidate in the existing MPM list to improve the efficiency of intra mode encoding and decoding. 2.16.1. Template-based intra-mode derivation Figure 20 FIG is a schematic diagram showing target points, template points, and reference points of the template used in DIMD. Figure 20 As shown, the target represents the current block (block size is N) for which the intra prediction mode is to be estimated. Figure 20 The patterned area indication in ( ) specifies a set of reconstructed samples that are used to derive the intra mode. The template size is expressed as the number of samples within the template that extend above and to the left of the target block, i.e., L. In the current implementation, for 4x4 and 8x8 blocks, a template size of 2 (i.e., L=2) is used, and for 16x16 and larger blocks, a template size of 4 (i.e., L=4) is used. The reference of the template (given by Figure 20 The dashed area in the image (indicated by the dashed area in the image) refers to a set of neighboring samples from above and to the left of the template defined by JEM-2.0. Unlike template samples, which are always from the reconstructed area, the template's reference samples may not have been reconstructed when encoding / decoding the target block. In this case, the existing reference sample replacement algorithm of JEM-2.0 is utilized to replace unavailable reference samples with available reference samples. For each intra prediction mode, DIMD calculates the SAD between the reconstructed template samples and its predicted samples obtained from the template's reference samples. The intra prediction mode that produces the smallest SAD is selected as the final intra prediction mode for the target block. 2.16.2. DIMD for Intra 2N×2N CU For intra 2Nx2N CUs, DIMD is used as an additional intra mode that is adaptively selected by comparing the DIMD intra mode with the best normal intra mode (i.e., the intra mode explicitly signaled). For each intra 2Nx2N CU, a flag is signaled to indicate the use of DIMD. If the flag is 1, the intra mode derived by DIMD is used to predict the CU; otherwise, DIMD is not applied and the intra mode explicitly signaled in the bitstream is used to predict the CU. When DIMD is enabled, the chroma components always reuse the same intra mode derived for the luma component, i.e., the DM mode. Additionally, for each DIMD-encoded CU, blocks in the CU can adaptively choose to derive their intra mode at the PU level or the TU level. Specifically, when the DIMD flag is 1, another CU-level DIMD control flag is signaled to indicate the level at which DIMD is performed. If the flag is 0, it means that DIMD is performed at the PU level and all TUs in the PU use the same derived intra mode for their intra prediction; otherwise (i.e., the DIMD control flag is 1), it means that DIMD is performed at the TU level and each TU in the PU derives its own intra mode. Furthermore, when DIMD is enabled, the number of angular directions increases to 129, and DC mode and planar mode remain unchanged. To accommodate the increased granularity of angular intra modes, the precision of intra interpolation filtering for DIMD-encoded CUs is increased from 1 / 32 pixel to 1 / 64 pixel. Additionally, in order to use the derived intra modes of a DIMD-encoded CU as MPM candidates for neighboring intra blocks, these 129 directions of the DIMD-encoded CU are converted to "normal" intra modes (i.e., 65 angular intra directions) before being used as MPMs. 2.16.3. DIMD for Intra N×N CU In the proposed method, the intra mode of an intra NxN CU is always signaled. However, to improve the efficiency of intra mode encoding and decoding, the intra mode derived from DIMD is used as an MPM candidate for predicting the intra mode of the four PUs in the CU. In order not to increase the overhead of MPM index signaling, the DIMD candidate is always placed at the first position in the MPM list and the last existing MPM candidate is removed. In addition, deduplication is performed so that if a DIMD candidate is redundant, it will not be added to the MPM list. 2.16.4.DIMD Intra-frame Pattern Search Algorithm To reduce encoding / decoding complexity, a straightforward fast intra mode search algorithm is used for DIMD. First, an initial estimation process is performed to provide a good starting point for the intra mode search. Specifically, an initial candidate list is created by selecting N fixed modes from the allowed intra modes. Then, the SAD is calculated for all candidate intra modes, and the one that minimizes the SAD is selected as the starting intra mode. To achieve a good complexity / performance trade-off, the initial candidate list consists of 11 intra modes, including DC, planar, and every 4th mode of the 33 angular intra directions as defined in HEVC, i.e., intra modes 0, 1, 2, 6, 10…30, 34. If the starting Intra mode is DC or Planar, it is used as the DIMD mode. Otherwise, based on the starting Intra mode, a refinement process is then applied, where the optimal Intra mode is identified through an iterative search. It works by comparing the SAD values ​​of three Intra modes separated by a given search interval at each iteration and maintaining the Intra mode that minimizes the SAD. The search interval is then reduced to half, and the selected Intra mode from the last iteration will be used as the center Intra mode for the current iteration. For the current DIMD implementation with 129 angular Intra directions, a maximum of 4 iterations are used in the refinement process to find the optimal DIMD Intra mode. 2.17. Decoder-side Intra-mode Derivation by Computing Gradients of Neighboring Samples The three angle modes are selected from the Histogram of Gradients (HoG) calculated from the neighboring pixels of the current block. Once the three modes are selected, their predictions are calculated normally, and then a weighted average of the predictions is used as the final prediction for the block. To determine the weights, the corresponding magnitude in the HoG is used for each of the three modes. DIMD mode is used as an alternative prediction mode and is always checked in FullRD mode. The current version of DIMD has modified some aspects of signaling, HoG computation, and prediction fusion. The purpose of these modifications is to improve codec performance and address the complexity issues raised during the last meeting (i.e., throughput of 4x4 blocks). The following sections describe each of these modifications. 2.17.1. Signaling Figure 21 is a schematic diagram illustrating the proposed intra block decoding process. Figure 21 The order of parsing flags / indexes integrated with the proposed DIMD in VTM5 is shown. As can be seen, the DIMD flag of the block is first parsed using a single CABAC context, which is initialized to a default value of 154. If flag == 0, parsing continues normally. Otherwise (if flag == 1), only the ISP index is parsed and the following flags / indexes are assumed to be zero: BDPCM flag, MIP flag, MRL index. In this case, the entire IPM parsing is also skipped. During the parsing phase, when a regular non-DIMD block asks its DIMD neighbor's IPM, the pattern PLANAR_IDX is used as the DIMD block's virtual IPM. 2.17.2. Texture Analysis Figure 22 : is a schematic diagram showing the calculation of HoG from a template with a width of 3 pixels. Texture analysis of DIMD includes the calculation of the Histogram of Gradients (HoG) ( Figure 22 The HoG calculation is performed by applying horizontal and vertical Sobel filters to the pixels in a template of width 3 around the block. However, if the upper template pixels fall into a different CTU, they will not be used in the texture analysis. Once calculated, the IPMs corresponding to the two highest histogram bins are selected for the block. In previous versions, all pixels in the middle row of the template participated in the HoG calculation. However, the current version improves the throughput of this process by applying the Sobel filter more sparsely on 4x4 blocks. For this purpose, only one pixel from the left and one pixel from above are used. This is in Figure 22 is shown in . Besides reducing the number of operations for gradient computation, this property also simplifies the selection of the best 2 patterns from the HoG, since the resulting HoG cannot have more than two non-zero magnitudes. 2.17.3. Prediction Fusion The current method uses a fusion of three prediction values ​​for each block. However, the choice of prediction mode is different and utilizes the proposed combined hypothetical intra prediction method, where planar mode is considered to be used in combination with other modes when computing intra prediction candidates. In the current version, the two IPMs corresponding to the two highest HoG slices are combined with planar mode. Prediction fusion is applied as a weighted average of the above three predictions. For this purpose, the weight of the plane is fixed to 21 / 64 (~1 / 3). Then, the remaining weight of 43 / 64 (~2 / 3) is shared between the two HoGIPMs, proportional to the amplitude of their HoG stripes. Figure 23 Visualize this process. Figure 23 is a schematic diagram illustrating prediction fusion by weighted averaging of two HoG modes and a plane. 2.18. Template-based Intra Mode Derivation (TIMD) This contribution proposes a template-based intra mode derivation (TIMD) method using MPM, where the TIMD mode is derived from the MPM using neighboring templates. The TIMD mode is used as an additional intra prediction method for a CU. 2.18.1. TIMD Mode Derivation For each intra prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The intra prediction mode with the smallest SATD is selected as the TIMD mode and used for intra prediction of the current CU. Position-dependent intra prediction combining (PDPC) is included in the derivation of the TIMD mode. 2.18.2.TIMD Signaling A flag is signaled in the sequence parameter set (SPS) to enable / disable the proposed method. When the flag is true, a CU-level flag is signaled to indicate whether the proposed TIMD method is used. The TIMD flag is signaled immediately after the MIP flag. If the TIMD flag is true, the remaining syntax elements related to the luma intra prediction mode (including MRL, ISP, and the normal parsing phase for luma intra prediction mode) are skipped. 2.18.3. Interaction with new codec tools The DIMD method with prediction fusion using planes is integrated in EE2. When the EE2 DIMD flag is equal to true, the proposed TIMD flag is not signaled and is set equal to false. Similar to PDPC, gradient PDPC is also included in the derivation of TIMD mode. When the secondary MPM is enabled, both the primary MPM and the secondary MPM are used to derive the TIMD mode. The 6-tap interpolation filter is not used for the derivation of TIMD mode. 2.18.4. Modification of MPM list construction in TIMD mode derivation During the construction of the MPM list, the intra prediction mode of the neighboring blocks is derived as a plane when they are inter-coded. To improve the accuracy of the MPM list, when the neighboring blocks are inter-coded, the propagated intra prediction mode is derived using the motion vector and reference picture and used in the construction of the MPM list. This modification is only applied to the derivation of TIMD mode. 2.18.5. TIMD with Fusion Instead of selecting the only mode with the smallest SATD cost, this contribution proposes to select the first two modes with the smallest SATD cost for the intra modes derived using the TIMD method, then fuse them using weights, and such weighted intra prediction is used to encode and decode the current CU. The costs of the two selected modes are compared with the threshold, and in the test a cost factor of 2 is applied as follows: costMode2<2×costMode1. If this condition is true, then fusion is applied, otherwise only mode1 is used. The weights of the modes are calculated from their SATD costs as follows: weight1=costMode2 / (costMode1+costMode2), weight2=1–weight1. 2.19. Convolutional Cross-Component Model (CCCM) for Intra Prediction It is proposed to apply a convolutional cross-component model (CCCM) to predict chroma samples from reconstructed luma samples in a similar spirit as done by the current CCLM mode. Like CCLM, when chroma downsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid. In addition, similar to CCLM, there is an option to use a single model or a multi-model variant of CCCM. The multi-model variant uses two models, one model is derived for samples above the average luminance reference value, and the other model is derived for the remaining samples (following the spirit of CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples. 2.19.1. Convolutional Filters The proposed convolutional 7-tap filter consists of a 5-tap plus sign-shaped spatial component, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter consists of the center (C) luminance sample co-located with the chrominance sample to be predicted and its upper / north (N), lower / south (S), left / west (W), and right / east (E) neighbors, as shown in Figure 24 As shown, Figure 24 The spatial portion of the convolution filter is shown. The nonlinear term P is expressed as the square of the center luminance sample C and scaled to the content's sample value range: P=(C*C+midVal)>>bitDepth. That is, for 10-bit content, it is calculated as: P=(C*C+512)>>10. The bias term B represents a scalar offset between the input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content). The output of the filter is calculated as the filter coefficient c i Convolution with the input value and clipped to the range of valid chroma samples: predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B. 2.19.2. Calculation of filter coefficients Filter coefficient c i is calculated by minimizing the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region. Figure 25 is a schematic diagram showing the reference area (with its filling) used to derive filter coefficients. Figure 25 A reference region consisting of six rows of chroma samples above and to the left of the PU is shown. The reference region extends one PU width to the right at the PU boundary and one PU height below the PU boundary. The region is adjusted to include only available samples. The extension of the region shown in blue is required to support the "side samples" of the shape spatial filter and is padded when in unavailable areas. MSE minimization is performed by computing the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is ​​decomposed by LDL, and the final filter coefficients are calculated using inverse substitution. The process roughly follows the calculation of the ALF filter coefficients in ECM, however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. The proposed method uses only integer arithmetic. 2.19.3. Bitstream Signaling The use of the mode is signaled using a PU-level flag via the CABAC codec. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered a submode of CCLM. That is, the CCCM flag is signaled only when the intra prediction mode is LM_CHROMA_IDX (to enable single-mode CCCM) or MMLM_CHROMA_IDX (to enable multi-mode CCCM). 2.20. Gradient Linear Model (GLM) Compared to CCLM, GLM uses the gradient of luma samples to infer the linear model instead of the downsampled luma values. Specifically, when GLM is applied, the input of the CCLM process (i.e., the downsampled luma samples L) is replaced by the luma sample gradient G. The other parts of CCLM (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged. C=α·G+β For signaling, when CCLM mode is enabled for the current CU, two flags are transmitted separately for the Cb component and the Cr component through signals to indicate whether GLM is enabled for each component; if GLM is enabled for a component, a syntax element is further transmitted through signals to select one of the four gradient filters for gradient calculation. like Figure 26 As shown, four gradient filters are enabled for the GLM. GLM with brightness In ECM-6.0, GLM uses the gradient of luma samples to predict chroma samples as follows: pred C (i, j) = α·G(i, j) + β where pred C (i, j) represents the predicted value of the chroma sample, G(i, j) represents the gradient of the corresponding reconstructed luma sample, and the linear model parameters α and β are derived from the adjacent reconstructed samples based on the same linear minimum mean square error (LMMSE) method as CCLM. A new GLM model is proposed, where the chroma samples are based on the gradient G(i,j) of the luma samples and the reconstructed values ​​rec′ of the downsampled luma samples with different parameters L (i,j) and predicted: pred C (i,j)=α0·G(i,j)+α1·rec′ L (i,j)+α2·midValue The model parameters α0, α1, and α2 are derived from six rows and six columns of adjacent sample points based on the same LDL decomposition method as the CCCM model in ECM-6.0. 2.21. Gradient and Position-Based Convolutional Cross-Component Model (GL-CCCM) for Intra Prediction The proposed GL-CCCM method uses gradient and position information to replace the four spatial neighboring samples in the CCCM filter. The GL-CCCM filter for prediction is: predChromaVal=c0C+c1G y +c2G x +c3Y+c4X+c5P+c6B. Among them G y and G x are the vertical and horizontal gradients respectively, and are calculated as: G y =(2N+NW+NE)–(2S+SW+SE), G x =(2W+NW+SW)–(2E+NE+SE). Furthermore, the Y parameter and the X parameter are the vertical position and the horizontal position of the center luma sample, and they are calculated relative to the upper left coordinate of the block. The remaining parameters are the same as those of the CCCM tool. The reference area used for parameter calculation is the same as that of the CCCM method. Figure 27 is a schematic diagram showing spatial sampling points for GL-CCCM. Bitstream signaling The use of this mode is signaled using a PU-level flag via the CABAC codec. A new CABAC context is included to support this. When it comes to signaling, GL-CCCM is considered a submode of CCCM. That is, the GL-CCCM flag is signaled only when the original CCCM flag is true. Encoder Operation The encoder performs two new RD checks in the chroma prediction mode loop, one RD check for checking the single-model GL-CCCM mode and one RD check for checking the multi-model GL-CCCM mode. 2.22. CCCM using unsubsampled luma samples 2.22.1. Block Level In this contribution, a CCCM using non-downsampled luma samples is proposed, where the chroma samples are predicted directly from the original reconstructed luma samples, ie, without downsampling. Figure 28 Schematic diagram showing the luminance samples that have not been downsampled. Figure 28 As shown in Figure 1, the proposed CCCM filter consists of a 6-tap spatial term, two nonlinear terms, and a bias term. The 6-tap spatial term corresponds to the 6 neighboring luma samples (i.e., L0, L1, ..., L5) of the chroma sample to be predicted (i.e., C). where α i Yes and L i The associated coefficients are β and β is the offset. As in the existing CCCM design, the chroma samples no more than 6 rows / columns above and to the left of the current CU are applied to derive the filter coefficients. The filter coefficients are derived based on the same LDL decomposition method used in CCCM. In this contribution, the proposed method is signaled as an additional CCCM model in addition to the existing CCCM model. For signaling, when CCCM is selected, a single flag is signaled and used for both chroma components to indicate whether the default CCCM model or the proposed CCCM model is applied. 2.22.2. High-level control For content with sharp details (such as SCC content), downsampling of the luma component may not be optimal for CCCM model derivation. In this contribution, it is proposed to disable luma downsampling and derive and apply the model directly on the un-downsampled luma samples. If downsampling is not applied, the CCCM model shape is a diamond 5x5. The SPS flag is signaled to indicate whether luma downsampling is applied to the CCCM. 2.23. Airspace GPM (SGPM) In the spatial domain GPM, a candidate list including partitioning and two intra prediction modes is constructed. An MPM with no more than 11 intra prediction modes is used to form a combination, and the length of the candidate list is set to be equal to 16. The selected candidate index is transmitted through the signal. Figure 29 The spatial domain GPM candidates are shown. Figure 29 The templates shown reorder the list. The GPM blending process is not used in the templates, and the SAD between the prediction and reconstruction of the templates is used for sorting. The SGPM mode is applied to blocks whose width and height satisfy the same restrictions as in inter-frame GPM. Figure 30 A GPM template is shown. Consider the following projects: Airspace GPM segmentation mode: 26 predefined modes. An adaptive inference algorithm based on the ratio of horizontal gradient to vertical gradient. Intra-frame prediction mode selection: List of IPMs with and without TIMD: For each segmentation mode, an IPM list is derived for each part using the intra-inter GPM list. The IPM list size is 3. In the list, the TIMD-derived mode is replaced by 2 derived modes with horizontal and vertical directions (using the top or left template), or the TIMD-derived mode is excluded. MPM List: A unified MPM list (maximum 11 elements) is used for all segmentation modes. Template size (left and top): 1 or 4. Extended block size: Spatial GPM is extended to be further applied to 4x8, 8x4, 4x16 and 16x4 blocks, which can be described as 4<=width<=64, 4<=height<=64, width<height*8, height<width*8, width*height>=32. Adaptive Mixing: Adaptive mixing is tested for spatial GPM, where the mixing depth τ is derived as follows: ■If min(width, height) == 4, 1 / 2τ is selected. ■ Otherwise, if min(width, height) == 8, then τ is selected. ■ Otherwise, if min(width, height) == 16, then 2τ is selected. ■ Otherwise, if min(width, height) == 32, then 4τ is selected. ■Otherwise, 8τ is selected. Figure 31 GPM mixing is shown. 2.24. Signaling of cross-component prediction modes in ECM Figure 32 Binarization of cross-component prediction modes in ECM is shown. Figure 32 "CCLM" in the format can be replaced by "CCCM". In ECM-7, cross-component modes include CCLM, CCLM-L, CCLM-T, MM-CCLM, MM-CCLM-L, MM-CCLM-T and CCCM, CCCM-L, CCCM-T, MM-CCCM, MM-CCCM-L, MM-CCCM-T. A flag is transmitted through the signal to determine whether it is a CCCM mode or a CCLM mode. The truncation unary code is used to indicate such Figure 32 CCLM mode or CCCM mode shown. CCLM or CCCM: 0. MM-CCLM or MM-CCCM: 10. CCLM-L or CCCM-L: 110. CCLM-T or CCCM-T: 1110. MM-CCLM-L or MM-CCCM-L: 11110. MM-CCLM-T or MM-CCCM-T: 11110. 2.25. Slope Adjustment for CCLM CCLM uses a 2-parameter model to map luma values ​​to chroma values. The slope parameter "a" and the bias parameter "b" define the mapping as follows: chromaVal=a*lumaVal+b. An adjustment of the slope parameter “u” via signal transmission is proposed to update the model to the following form: chromaVal=a'*lumaVal+b' in a'=a+u, b'=bu*y r . With this choice, the mapping function is centered around the value y with the brightness r The points are tilted or rotated. It is proposed to use the average of the reference brightness samples used in model creation as y r , to provide meaningful modifications to the model. The following image illustrates this process. 2.26. Fusion of Chroma Intra Prediction Modes In test 1.2b, it is proposed that the DM mode and the four default modes can be fused with the MMLM_LT mode as follows: pred=(w0*pred0+w1*pred1+(1<<(shift-1)))>>shift Where pred0 is the prediction value obtained by applying the non-LM mode, pred1 is the prediction value obtained by applying the MMLM_LT mode, and pred and are the final prediction values ​​of the current chroma block. The two weights w0 and w1 are determined by the intra prediction mode of the adjacent chroma blocks, and shift and are set to be equal to 2. Specifically, when the upper adjacent block and the left adjacent block are both encoded and decoded using the LM mode, {w0, w1} = {1, 3}; when the upper adjacent block and the left adjacent block are both encoded and decoded using the non-LM mode, {w0, w1} = {3, 1}; otherwise, {w0, w1} = {2, 2}. For syntax design, if non-LM mode is selected, a flag is signaled to indicate whether merging is applied, and the proposed merging is only applied to I slices. 2.27. History-based Cross-Component Prediction (H-CCP) 1. It is proposed that the model(s) of cross-component prediction (CCP) (such as CCLM or CCCM) in a block can be stored into a history table (HT). a.HT is a list with ordered entries. i. Each entry has an index. For example, the first entry has an index of 0, and subsequent entries have indices of 1, 2, 3, ... b. The model parameters of CCLM and its variants can include a, b, and shift to control the calculation accuracy. c. The model parameters of CCLM and its variants may include a linear part (such as c0-c4) and a nonlinear part (such as c5). d. Models may include models for different color components such as Cb and Cr. i. For example, the models for Cb and Cr can be coupled in the entry. e. In one example, different CCPs such as CCLM and CCCM can share the same HT. i. In one example, a segment in an entry of an HT may reflect the type of CCP model(s) stored in the entry. f. In one example, different CCPs such as CCLM and CCCM may have different HTs. i. In one example, a CCLM_HT can store models of CCLM and its variants (such as CCLM-L or CCLM-T). ii. In one example, one CCCM_HT can store models of CCCM and its variants (such as CCCM-T or CCCM-T). g. In one example, a CCP with a single model (such as CCLM or CCCM) and a CCP with multiple models (such as MM-CCLM or MM-CCCM) may have different HTs. h. In one example, a CCP with a single model (such as CCLM or CCCM) and a CCP with multiple models (such as MM-CCLM or MM-CCCM) can share the same HT. i. In one example, the segment in an entry of the HT may reflect the number of models stored in the entry. ii. In one example, a segment in an entry of HT may reflect at least one threshold for classifying samples into different model groups. i. In one example, the first HT is used to store models of CCLM and its variants. i. In one example, CCLM variants may include CCLM-L, CCLM-T, MM-CCLM, MM-CCLM-L, MM-CCLM-T, GLM, and CCLM with slope adjustment. 1) The segment in an entry of HT may reflect the number of models stored in the entry. 2) The segments in the entries of HT may reflect at least one threshold used to classify samples into different model groups. 3) The segment in the HT entry may reflect whether GLM is applied. 4) The segments in the entries of HT may reflect the downsampling filters of GLM. j. In one example, the second HT is used to store models of CCCM and its variants. i. In one example, CCCM variants may include CCCM-L, CCCM-T, MM-CCCM, MM-CCCM-L, MM-CCCM-T. 1) The segment in an entry of HT may reflect the number of models stored in the entry. 2) The segments in the entries of HT may reflect at least one threshold used to classify samples into different model groups. 2. It is proposed that a block can be encoded and decoded using the history-based CCP (H-CCP) mode, in which at least one CCP model used by the current block is obtained or derived from the HT. a. In one example, at least one syntax element (SE) may be signaled to indicate whether H-CCP is applied. i. In one example, SE can be conditionally signaled. For example, SE is signaled only when a specific mode (such as CCCM or CCLM) is used. 1) For example, SE is signaled only when the current mode is CCCM or CCLM. b. In one example, at least one syntax element (SE) may be signaled to indicate which entry in the HT is retrieved to derive model(s) for cross-component prediction. i.SE can reflect the index in HT. 1) In one example, SE may be set equal to f(k), where k is an index and f is a function. 2) In one example, SE may be set equal to f(k, M), where k is an index, M is the number of valid entries in the HT, and f is a function. a) In another example, M is the size of HT. 3) In one example, SE may be set equal to k, where k is an index. 4) In one example, SE may be set equal to M-1-k, where k is the index and M is the number of valid entries in the HT. a) In another example, M is the size of HT. ii. SE can reflect the index of the list, and the list can be constructed based on HT. 1) In one example, the list L is constructed by reversing HT. For example, L[i]=HT[M-1-i], where M is the number of valid entries in HT. a) In another example, M is the size of HT. b) In one example, L may have a fixed size. c) In one example, if L is not full, the empty entry is filled with a default entry. iii. In one example, SE can be signaled conditionally, for example, SE is signaled only when H-CCP is applicable. iv. SE can be signaled only when more than one entry in HT can be selected. The maximum value of v.SE (denoted as V) is determined by the number of entries to be selected. 1) For example, V=K, or V=K-1, or V=K+1, or V=K-2, or V=K+2. c. In one example, at least one syntax element (SE) may be signaled to indicate which HT is used. i. In one example, SE can be signaled conditionally, for example, SE is signaled only when H-CCP is applicable. ii. SE can be signaled only when more than one HT can be selected. d. In one example, it can be inferred at the encoder / decoder which HT is used. i. In one example, if the current mode is CCLM, the first HT storing the model of CCLM and its variants is used. ii. In one example, if the current mode is CCCM, the second HT storing the model of CCCM and its variants is used. e. In one example, the current block may be predicted using a CCP model obtained from the determined entry of the determined HT. f. In one example, the current block may be predicted using CCCM or CCLM based on whether the first HT or the second HT is applied. g. In one example, the current block can be predicted using multiple models. i. Whether a single model or multiple models is applied can be deduced / obtained from the determined entries of the determined HT. ii. At least one threshold value for classifying samples into different model groups may be obtained / derived from the determined entries of the determined HT. HT Maintenance 3. The maximum size of HT can be predetermined, such as 5 or 6. a. Alternatively, the maximum size of the HT can be signaled as SE at the block level / sequence level / picture group level / picture level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB or sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. b. Alternatively, the maximum size of the HT can be derived using the following encoding / decoding information: i. The mode of the current block; ii. Patterns of neighboring blocks; iii. The pattern of the luminance blocks in the same region as the current block; iv. The pattern of luminance blocks in the same region as the neighboring blocks; v.QP; vi. Strip / image type; vii. Image width / height; viii.Block width / height; ix. Reconstructed sample points. 4. HT can be refreshed at the beginning of a coding / decoding sequence / picture / slice / slice / sub-picture / CTU row / CTU. a. For example, HT can be refreshed by clearing the table. b. For example, the HT can be refreshed by completing the table with default entries. 5. After encoding / decoding a block (such as a CU), the HT may be updated. a. For example, when dual-tree coding is applied, the CU must be a chroma CU. b. For example, the CU must be a CU with CCP mode. c. For example, which HT to be updated may depend on the coding mode of the CU. i. For example, if the CU is encoded using a CCLM mode (such as CCLM, CCLM-L, CCLM-T, MM-CCLM, MM-CCLM-L, MM-CCLM-T, GLM, and CCLM with slope adjustment), then (multiple) models and related information (such as (multiple) thresholds for classifying samples into different model groups) are stored in the first HT. ii. For example, if the CU is encoded and decoded using a CCCM mode (such as CCCM, CCCM-L, CCCM-T, MM-CCCM, MM-CCCM-L, MM-CCCM-T), then (multiple) models and related information (such as (multiple) thresholds for classifying samples into different model groups) are stored in the first HT. d. For example, a set of information about the CCP model(s) used by the current block may be placed in the HT. i. The group may include one or more CCP models. ii. The number of models that the group may include. iii. The set may include threshold(s) for classifying samples into different model groups. iv. The group may include slope adjustments. e. In one example, if the current block is encoded using CCLM with slope adjustment, the CCP model can be adjusted before being used to update the HT. 6. How a new set of information related to the CCP model(s) is placed into the HT may depend on whether the HT is full. a. For example, if the HT is not full, the new set may be placed into the first empty entry of the HT. i. For example, the first slot entry is the slot entry with the smallest index. ii. For example, the first slot entry is the slot entry with the largest index. iii. After being placed in the HT, the new group may be placed as the last occupied entry in the HT. 1) The last occupied entry may be the occupied entry with the largest index. 2) The last occupied entry may be the occupied entry with the smallest index. b. For example, if the HT is full, an existing entry in the HT may be removed. i. In one example, HT can be managed in a first-in-first-out manner. ii. The existing entry with the smallest index can be removed. 1) The updated HT' may be set to: HT'[i]=HT[i+1], for 0<=i<=N-2, and HT'[N-1]=new group, where N is the size of the HT. iii. The existing entry with the largest index can be removed. 1) The updated HT' may be set to: HT'[i]=HT[i-1], for 1<=i<=N-1, and HT'[0]=new group, where N is the size of the HT. 7. In one example, the new group may be compared to at least one existing entry in the HT to determine whether to place the new group and / or how to update the HT. 8. In one example, if a new group is identical or similar to one of the existing entries in the HT, the new group is not placed in the HT. It is assumed that the new group is identical or similar to a particular entry of the HT. a. For example, in this case, the special entry can be put at the first of the HT, and the entry that originally preceded the special entry is pushed back one position. i. For example, assuming that the entry is HT[i] (where i=0, 1, ...), and the special entry is HT[k], the updated HT' will be as follows: HT'[0]=HT[k]; HT'[i]=HT[i-1] (for 1<=i<=k); HT'[i]=HT[i] (for i>k). b. For example, in this case, the special entry may be placed at the end of the HT, and the entry that originally preceded the special entry is pushed forward one position. i. For example, assuming that the entry is HT[i] (where i=0, 1, ...), and the special entry is HT[k], the updated HT' will be as follows: HT'[N-1]=HT[k]; HT'[i]=HT[i+1] (for k<=i<=N-2); HT'[i]=HT[i] (for i <k)。 9. In one example, whether to put a new group and / or how to update the HT may depend on the codec information of the CU with the new group. 10. In one example, if the new group is for a CU coded in H-CCP mode, the new group is not placed in the HT. It is assumed that the special entry in the HT is used by the CU coded in H-CCP mode. a. For example, in this case, the special entry can be put at the first of the HT, and the entry that originally preceded the special entry is pushed back one position. i. For example, assuming that the entry is HT[i] (where i=0, 1, ...), and the special entry is HT[k], the updated HT' will be as follows: HT'[0]=HT[k]; HT'[i]=HT[i-1] (for 1<=i<=k); HT'[i]=HT[i] (for i>k). b. For example, in this case, the special entry may be placed at the end of the HT, and the entry that originally preceded the special entry is pushed forward one position. i. For example, assuming that the entry is HT[i] (where i=0, 1, ...), and the special entry is HT[k], the updated HT' will be as follows: HT'[N-1]=HT[k]; HT'[i]=HT[i+1] (for k<=i<=N-2); HT'[i]=HT[i] (for i <k)。 11. It is proposed that the entries of HT may include models for more than one chroma component (such as Cb and Cr). a. If the entry is selected, the models for components Cb and Cr are applied to the two components separately. 12. It is proposed that the entry for HT may include a model for only one component (such as Cb or Cr). a. If the entry is selected, the model for a specific component such as Cb or Cr is applied to the specific component. b. In one example, different HTs may be constructed for different components. List Mode 13. It is proposed that at least one list with a CCP model can be constructed. a. In one example, chroma blocks can be predicted using the CCP model in the list in "list mode". b. In one example, list L may be populated with one type of CCP model, such as CCCM. c. In one example, the list may be populated with multiple types of CCP models, such as both CCCM and CCLM. i. In one example, the type of CCP model will be stored in a list along with the CCP model. d. In one example, at least one syntax element (SE) may be signaled to indicate whether a CCP model in the list is used. i. In one example, SE can be conditionally signaled. For example, SE is signaled only when a specific mode (such as CCCM or CCLM) is used. 1) For example, SE is signaled only when the current mode is CCCM or CCLM. 2) For example, SE is signaled only when "list mode" is applicable. e. In one example, at least one syntax element (SE) may be signaled to indicate which entry in the list is used to derive the model(s) for cross-component prediction. i.SE can reflect the index in the list. 1) In one example, SE may be set equal to f(k), where k is an index and f is a function. 2) In one example, SE may be set equal to f(k, M), where k is the index, M is the number of valid entries in the list, and f is a function. a) In another example, M is the size of the list. 3) In one example, SE may be set equal to k, where k is an index. 4) In one example, SE may be set equal to M-1-k, where k is the index and M is the number of valid entries in the list. a) In another example, M is the size of the list. f. In one example, L may have a fixed size. g. In one example, multiple lists can be constructed. i. For example, at least one syntax element (SE) may be signaled to indicate which list is used. ii. In one example, SE can be signaled conditionally, for example, only when "list mode" is applicable. iii. SE can be signaled only when more than one list can be selected. h. In one example, which list to use can be inferred at the encoder / decoder. i. In one example, if the current mode is CCLM, a first list of models storing CCLM and its variants is used. ii. In one example, if the current mode is CCCM, a second list of models storing CCCM and its variants is used. 14. It is proposed that an entry of a list may include models for more than one chroma component (such as Cb and Cr). a. If the entry is selected, the models for components Cb and Cr are applied to the two components separately. 15. It is proposed that an entry of a list may include a model for only one component (such as Cb or Cr). a. If the entry is selected, the model for a specific component such as Cb or Cr is applied to the specific component. 16. Multiple candidates can be put into the list, including: a. CCP model of adjacent neighboring blocks. b. CCP model of non-adjacent neighboring blocks. c. CCP model of the same-position block in the reference image. d. CCP model of the reference block in the reference image. e. CCP model in the history table. f. CCP model derived from non-adjacent sample points. g. Default CCP mode. 17. In one example, the list can be constructed by examining the possible candidates in order. a. For example, the order can be adjacent neighboring blocks, non-adjacent neighboring blocks, models in the history table, and models derived from non-adjacent samples. b. For example, if the number of candidates in the list reaches the maximum allowed size of the list (such as 5 or 6), then list building is completed. c. For example, if the number of candidates in the list reaches f(d), then the list construction is completed, where d is the index of the selected candidate and f is a function. For example, f(d)=d+1. d. For example, if all possible candidates have been examined and the build is not complete, a default model can be put into the list. 18. In one example, if a potential candidate is placed in a list, it may be compared to at least one existing candidate in the list. a. For example, if a potential candidate is the same as or similar to an existing candidate, the potential candidate is not put into the list. b. In one example, if a potential entry of CCP information is placed into a history-based table, it may be compared to at least one existing entry in the list. i. For example, if a potential entry is identical or similar to an existing entry, the potential entry is not placed in the list. c. In one example, two CCP candidates or entries are determined to be different if: i. Different CCP types. ii. The number of models is different. iii. If the CCP has multiple models, the threshold is different. iv. At least one model is different. v. Luma sample offset is different. (Applicable only when the type is CCCM, GL-CCCM, GLM, or CCCM using unsubsampled luma samples). vi. Sample point position shifts are different. (Applicable only when the type is GL-CCCM) 19. For example, the CCP information of an entry in a history-based table or the candidate CCP information in a CCP candidate list may include: a. Type of CCP method, such as CCLM, or CCCM, or GLM, or GLM with luma, or GL-CCCM, or CCCM using non-downsampled luma samples. i. In one example, GLM methods using different downsampling filters can be considered as different types. ii. In one example, GLM methods with luminance using different downsampling filters can be considered as different types. iii. In one example, the types may be CCCM, CCLM, 4 types of GLM using different downsampling filters, 4 types of GLM with luma using different downsampling filters, GL-CCCM, and CCCM using non-downsampled luma samples. iv. “Not using CCP codec” (denoted as NonCCP) can also be considered as a type. b. Position (x, y). c. Number of models. i. For example, the number of models can be 1 or 2. ii. In one example, the number of models can be considered as part of the CCP type. For example, CCLM and MM-CCLM can be considered as two types. d. At least one threshold for classifying samples for different models. i. The threshold is only used when the number of models is at least 2. e. At least one luma sample value offset. i. When luma sample value offsets are used to derive chroma prediction values, the luma sample value offsets may be added to or subtracted from the luma samples (which may be downsampled). ii. Luma sample value offset can be used only for certain types, such as CCCM, GLM with luma, GL-CCCM, and CCCM using non-downsampled luma samples. f. At least one chroma sample value is offset. i. Chroma sample value offsets can be added to or subtracted from the chroma prediction values ​​derived from the CCP model to generate the final prediction. g. At least one model for at least one chroma component. i. For example, it may include different models for the Cb component and the Cr component. ii. For example, the number of models for each component may be included as part of the information. iii. The model can be represented by a CCLM or CCCM or GLM or GLM with luma or GL-CCCM or CCCM using non-downsampled luma samples. h. At least one sample point position displacement expressed as (dX, dY). i. When chroma sample position offsets are used to derive chroma prediction values, the chroma sample position offsets can be added to or subtracted from the sample position (x, y). ii. Chroma sample position shifting can only be used for certain types, such as GL-CCCM. 20. For example, the CCP codec information of a chroma block after being encoded / decoded can be stored in a history-based table or in a CCP candidate list. a. In one example, CCP codec information may be stored only when the chroma block is coded using CCP mode. i. In one example, if the chroma block is coded with at least one CCP mode, such as fusion with chroma intra prediction mode, then the CCP codec information may be stored. 1) The stored type may be set as the CCP type used in fusion of chroma intra prediction modes. b. In one example, CCP codec information can be stored for any chroma block. i. If the chroma block is not coded using CCP mode, the type is stored as "NonCCP". c. If the chroma block is coded using CCP mode, the type of information may be stored depending on the coding mode. i. If the mode is CCCM, or CCCM-T, or CCCM-L, or MM-CCCM, or MM-CCCM-T, or MM-CCCM-L, the type is set to "CCCM". ii. If the mode is CCLM, or CCLM-T, or CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, the type is set to "CCLM". iii. If the mode is CCLM with slope adjustment, or CCLM-T, or CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, the type is set to "CCLM". iv. If the mode is GLM with filter X, the type is set to "GLM with filter X". v. If the mode is GLM with Luminance using filter X, then the type is set to "GLM with Luminance using filter X". vi. If the mode is GL-CCCM, the type is set to "GL-CCCM". vii. If the mode is to use CCCM without downsampling, the type is set to "Use CCCM without downsampling". viii. If the mode is a fusion of chroma intra prediction modes, the type is set to "CCLM". d. The number of models can be stored as the number of models of the chroma block. i. For example, if the mode is MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, or any other multi-model CCP mode (such as GLM, or GL-CCCM, or CCCM with multiple models using non-downsampled luma samples), the number of models is set to 2. e. Information such as thresholds, luma / chroma sample value offsets, and sample position displacements can be stored as information used by chroma blocks. f. The CCP model for a component can be stored as the model used by the chroma blocks. i. The model may be derived by any CCP method such as CCLM, or CCLM-T, or CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, or CCCM, or CCCM-T, or CCCM-L, or MM-CCCM, or MM-CCCM-T, or MM-CCCM-L, or GLM using different downsampling filters, or GLM with luma using different downsampling filters, or GL-CCCM, or CCCM using non-downsampled luma samples. ii. The stored model may be the model of the final application, such as after being modified by slope adjustment. 21. In one example, a history table of CCP information following a coding / decoding region (such as a CU / CTU / CTU row) may be stored, referred to as a storage table. a. A history table of CCP information maintained for the current block (referred to as an online table) can be used together with a stored history table of CCP information. b. In one example, entries in the storage table and the online table may be checked sequentially to generate new candidates. i. In one example, entries in the online table may be checked before all entries in the storage table. ii. In one example, entries in the storage table may be checked before all entries in the online table. iii. For example, the kth entry in the storage table may be checked after the kth entry in the online table. iv. For example, the kth entry in the online table may be checked after the kth entry in the storage table. v. For example, the kth entry in the line table may be checked after the mth entry (m=0...S, where S is an integer) in the storage table. vi. For example, the kth entry in the storage table may be checked after the mth entry (m=0...S, where S is an integer) in the line table. vii. For example, the kth entry in the line table may be checked after the mth entry in the storage table (m=S...maxT, where S is an integer and maxT is the last entry). viii. For example, the kth entry in the storage table may be checked after the mth entry in the line table (m=S...maxT, where S is an integer and maxT is the last entry). c. In one example, the storage table(s) to be used may depend on the dimensions and / or position of the current block. i. For example, a table stored in a CTU above the current CTU may be used. ii. For example, a table stored in the CTU to the upper left of the current CTU may be used. iii. For example, a table stored in the CTU to the upper right of the current CTU may be used. d. In one example, whether and / or how to use the storage table may depend on the dimensions and / or position of the current block. i. In one example, whether to use a storage table and / or how to use the storage table may depend on whether the current CU is located at the top boundary of a CTU and whether an upper neighboring CTU is available. 1) For example, the storage table can be used only when the current CU is located at the top boundary of a CTU and the upper neighboring CTU is available. 2) For example, if the current CU is located at the top boundary of a CTU and an upper adjacent CTU is available, at least one entry in the storage table may be placed at a front position. e. In one example, entries in two storage tables may be examined sequentially to generate new candidates. i. For example, the first (or second) storage table stored in the CTU above the current CTU may be used. ii. For example, the first (or second) storage table stored in the CTU to the upper left of the current CTU may be used. iii. For example, the first (or second) storage table stored in the CTU above and to the right of the current CTU may be used. 2.28. Non-adjacent cross-component prediction (NA-CCP) 1. It is proposed that the model(s) for cross-component prediction (such as CCLM or CCCM) in a block can be derived based on a set of samples that are not adjacent to the current block, which is called non-adjacent cross-component prediction (NA-CCP). a. In one example, a group of samples is non-adjacent to the current block only if no samples in the group are adjacent to the current block (such as adjacent above or adjacent to the left of the current block). b. In one example, a set of samples is reconstructed before encoding / decoding the current block. c. The samples may include chroma samples and / or their corresponding luma samples. If the color format is 4:2:0 or 4:2:2, the luma samples may be generated by downsampling. 2. In one example, at least one syntax element (SE) may be signaled to indicate whether non-adjacent cross-component prediction is applied. In one example, SE can be signaled conditionally. For example, SE can be signaled only when a specific mode (such as CCCM or CCLM) is used. 3. In one example, more than one set of samples that are not adjacent to the current block can be used to derive model(s) for cross-component prediction. a. In one example, samples in more than one group can be jointly used to derive model(s) for cross-component prediction. b. In one example, one of multiple sets of candidates may be selected to derive model(s) for cross-component prediction. 4. In one example, at least one syntax element (SE) may be signaled to indicate which group of non-adjacent samples is used to derive the model(s) for cross-component prediction. In one example, SE can be signaled conditionally, for example, only when NA-CCP is applicable. b. SE can be signaled only when more than one set of non-adjacent samples can be selected. c. The maximum value of SE (denoted as V) is determined by the number of non-adjacent samples to be selected (denoted as K). i. For example, V=K, or V=K-1, or V=K+1, or V=K-2, or V=K+2. 5. Whether / how to apply NA-CCP may be the same for more than one color component (such as Cb and Cr). a. Alternatively, whether / how to apply NA-CCP may be different for different components (such as Cb and Cr). 6. Whether NA-CCP is applicable may depend on the dimension / position of the current block. 7. In one example, a group of non-adjacent samples may include samples in a region. a. In one example, a region may be a codec block (eg, a CU). b. In one example, a region may be represented by a position relative to the region. c. In one example, the region may be an M×N rectangle (eg, M=N=8). d. In one example, a rectangular region may be represented by a position relative to the region (such as the top left position of the region (x, y)) and dimensions M×N. e. In one example, regions of non-adjacent samples from different groups may share the same shape and size. f. In one example, regions of non-adjacent samples in different groups may have different shapes or sizes. g. Sample points in the area must be reconstructed. i. Alternatively, if the samples in the region are not reconstructed, they should be filled. 8. In one example, luma samples corresponding to a set of non-adjacent chroma samples may be prepared or generated to be used for training a cross-component model. a. In one example, if the color format is 4:2:0 or 4:2:2, downsampling may be applied to generate corresponding luma samples. b. In one example, the generated luma samples may correspond to a larger area than the area of ​​non-adjacent chroma samples. i. In one example, assuming that the area of ​​non-adjacent chroma samples is an M×N rectangle, the generated luma samples may correspond to a (M+T+B)×(N+L+R) chroma rectangle, such as Figure 33 shown. Figure 33 An example of luminance samples to be prepared is shown. 1) In one example, T=B=L=R=1. c. In one example, if the luma sample to be generated is not available (e.g., it is outside the picture boundary, or it has not been reconstructed, or it is in a different CTU that has not been reconstructed, etc.), then the luma sample can be handled specially. i. In one example, it may be filled, such as repeatedly filled with nearby available generated brightness values. ii. In one example, it may not be generated and marked as "unavailable". 1) The dimension of the brightness area can be set to the available area. 9. In one example, whether a region comprising non-adjacent sample points is a valid set of sample points for deriving model(s) may be determined by the availability of at least one sample point of the region. a. For example, the region is a rectangle. b. For example, a region is determined to be valid only when both the upper left reconstructed sample points and the lower right reconstructed sample points of the region are available. c. For example, a region is determined to be valid only when both the upper right reconstructed sample points and the lower left reconstructed sample points of the region are available. 10. In one example, a region list can be constructed to record multiple groups of non-adjacent sample points. a. In one example, the index of the list can be signaled as SE to indicate which set of non-adjacent samples is used to derive the model(s) for cross-component prediction. i. For example, SE can be binarized into a truncated unary code. ii. In one example, SE can be conditionally signaled, for example, SE is signaled only when NA-CCP is applied. iii. SE can be signaled only when more than one set of non-adjacent samples can be selected. iv. The maximum value of SE (denoted as V) is determined by the number of non-adjacent samples to be selected (denoted as K). 1) For example, V=K, or V=K-1, or V=K+1, or V=K-2, or V=K+2. b. In one example, a list may be constructed by examining multiple potential candidate regions in sequence. i. The list is initialized to empty. ii. If the number of candidate regions in the list is equal to the maximum size of the list (such as 6), the list construction is completed. iii. If all potential candidate regions have been examined, list building is complete. iv. If the region is determined to be valid, the potential candidate can be placed in a list. v. Deduplication may be applied to construct the list. 1) A potential candidate may not be put into the list if it is a "duplicate" of an existing candidate in the list. a) If the samples of the candidate region are the same as (or similar to) the samples of another region, then the candidate region "overlaps" with the other region. b) A candidate region “duplicates” another region if the same or similar models can be derived from samples in both regions. 11. In one example, the position and / or dimensions of the region including non-adjacent samples may depend on codec information, such as the width / height of the current block. a. This area may be a potential candidate area for the list. b. The distance between the region and the current block may depend on the width / height of the current block. 12. In one example, the potential candidate region may be an M×N rectangle (eg, M=N=8) non-adjacent to the left / lower left / upper left / above / upper right of the current block. Figure 34 Examples of potential candidate regions (shared blocks) are shown. 13. In one example, the potential candidate region is an M×N (e.g., M=N=8) rectangle, and its top left position (x0, y0) can be described as (assuming the top left position of the current block with dimensions W×H is (0, 0)): a. (x0, y0) = (s*f(W, H), t*g(W, H)), where f and g are functions. s and t are scaling factors, such as 0.5, 1, or 2. b. (x0, y0) = (s*f(W), t*g(H)), where f and g are functions. s and t are scaling factors, such as 0.5, 1, or 2. 14. In one example, the potential candidate regions are M×N (e.g., M=N=8) rectangles, and their top-left positions in order are as follows (assuming the top-left position of the current block of dimension W×H is (0, 0)): (-xStep,0), (0,-yStep), (xStep,-yStep), (-xStep,yStep), (-xStep,-yStep), (-2*xStep,0), (0,-2*yStep), (-2*xStep,2*yStep), (2*xStep,-2*yStep), (-2*xStep,yStep), (xStep,-2*yStep), (-2*xStep,-yStep), (-xStep,-2*yStep), (-2*xStep,-2*yStep), (-xStep / 2,0), (0,-yStep / 2), (xStep / 2,-yStep / 2), (-xStep / 2,yStep / 2), (-xStep / 2,-yStep / 2), Where xStep and yStep are integers. a. The order of inspection can be changed. b. In one example, xStep=Max(W, K1), yStep=Max(H, K2), where K1 and K2 are integers, for example, K1=K2=16. 15. In one example, whether to apply NA-CCP and / or how to apply NA-CCP can be signaled from the encoder to the decoder. a. Alternatively, whether to apply NA-CCP and / or how to apply NA-CCP can be derived at the encoder and decoder based on the encoded / decoded information without signaling. b. “How to apply NA-CCP” may include: i. Which CCP (such as CCLM or CCCM) model is derived from NA-CCP; ii. Shape / size / position of (potential) candidate regions; iii. Size of the region list; iv. the number of (potential) candidate regions; v. Color components to which NA-CCP is applied. c. “Encoded / decoded information” may include: i. The mode of the current block; ii. Patterns of neighboring blocks; iii. The pattern of the luminance blocks in the same region as the current block; iv. The pattern of luminance blocks in the same region as the neighboring blocks; v.QP; vi. Strip / image type; vii. Image width / height; viii.Block width / height; ix. Reconstructed sample points. 16. In one example, the CCP codec information of spatial or temporal neighboring blocks can be used by the current block. a. For example, the spatially adjacent blocks may be adjacent to or not adjacent to the current block. b. For example, CCP codec information may include: i. Type of CCP method, such as CCLM, or CCCM, or GLM, or GLM with luma, or GL-CCCM, or CCCM using non-downsampled luma samples. 1) In one example, GLM methods using different downsampling filters can be considered as different types. 2) In one example, GLM methods with luminance using different downsampling filters can be considered different types. 3) In one example, the types may be CCCM, CCLM, 4 types of GLM using different downsampling filters, 4 types of GLM with luma using different downsampling filters, GL-CCCM, and CCCM using non-downsampled luma samples. 4) “Not using CCP codec” (denoted as NonCCP) can also be considered as a type. ii. Position (x, y). iii. Number of models. 1) For example, the number of models can be 1 or 2. 2) In one example, the number of models can be considered as part of the CCP type. For example, CCLM and MM-CCLM can be considered as two types. iv. At least one threshold for classifying samples for different models. 1) The threshold is only used when the number of models is at least 2. v. At least one luma sample value offset. 1) When luma sample value offsets are used to derive chroma prediction values, the luma sample value offsets may be added to or subtracted from the luma samples (which may be downsampled). 2) Luma sample value offset can be used only for certain types, such as CCCM, GLM with luma, GL-CCCM, and CCCM using non-downsampled luma samples. vi. At least one chroma sample value offset. 1) Chroma sample value offsets can be added to or subtracted from the chroma prediction values ​​derived from the CCP model to generate the final prediction. vii. At least one model for at least one chroma component. 1) For example, it may include different models for the Cb component and the Cr component. 2) For example, the number of models for each component may be included as part of the information. 3) The model can be represented by a CCLM or CCCM or GLM or GLM with luma or GL-CCCM or CCCM using luma samples that are not downsampled. viii. At least one sample point position shift expressed as (dX, dY). 1) When chroma sample position offset is used to derive chroma prediction value, the chroma sample position offset can be added to or subtracted from the sample position (x, y). 2) Chroma sample position shifting can only be used for certain types, such as GL-CCCM. c. For example, CCP codec information can be stored after the chroma block is encoded / decoded. i. In one example, CCP codec information may be stored only when the chroma block is coded using CCP mode. 1) In one example, if the chroma block is coded with at least one CCP mode, such as fusion with chroma intra prediction mode, the CCP codec information may be stored. a) The stored type may be set as the CCP type used in fusion of chroma intra prediction modes. ii. In one example, for any chroma block, CCP codec information can be stored. 1) If the chroma block is not coded using CCP mode, the type is stored as "NonCCP". iii. If the chroma block is coded using CCP mode, the type of information may be stored depending on the coding mode. 1) If the mode is CCCM, or CCCM-T, or CCCM-L, or MM-CCCM, or MM-CCCM-T, or MM-CCCM-L, the type is set to "CCCM". 2) If the mode is CCLM, or CCLM-T, or CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, the type is set to "CCLM". 3) If the mode is CCLM with slope adjustment, or CCLM-T, or CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, the type is set to "CCLM". 4) If the mode is GLM with filter X, the type is set to "GLM with filter X". 5) If the mode is GLM with Luma using filter X, the type is set to "GLM with Luma using filter X". 6) If the mode is GL-CCCM, the type is set to "GL-CCCM". 7) If the mode is to use CCCM without downsampling, the type is set to "Use CCCM without downsampling". 8) If the mode is a fusion of chroma intra prediction modes, the type is set to "CCLM". iv. The number of models can be stored as the number of models of the chroma block. 1) For example, if the mode is MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, or any other multi-model CCP mode (such as GLM, or GL-CCCM, or CCCM with multiple models using non-downsampled luma samples), the number of models is set to 2. v. Information such as thresholds, luma / chroma sample value offsets, and sample position displacements can be stored as information used by chroma blocks. vi. The CCP model of a component can be stored as the model used by the chroma blocks. 1) The model can be derived by any CCP method (such as CCLM, or CCLM-T, or CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, or CCCM, or CCCM-T, or CCCM-L, or MM-CCCM, or MM-CCCM-T, or MM-CCCM-L, or GLM using different downsampling filters, or GLM with luma using different downsampling filters, or GL-CCCM, or CCCM using non-downsampled luma samples). 2) The stored model may be the model of the final application, such as the model after being modified by slope adjustment. d. For example, CCP codec information can be stored at an M×N granularity. i. For example, M=N=2. ii. For example, CCP codec information of a specific chroma block covered by, covering, or overlapping an M×N area may be stored in the M×N area. 1) For example, the CCP codec information of the first coding / decoding block having CCP information covered by an M×N area, covering an M×N area, or overlapping an M×N area may be stored. 2) For example, the CCP codec information of the last encoding / decoding block having CCP information covered by, covering, or overlapping an M×N area may be stored. 3) For example, CCP codec information of a coding / decoding block having CCP information of a specific position covered by an M×N area, covering an M×N area, or overlapping an M×N area may be stored. a) The specific position may be the upper left / lower right / upper right / lower left / center position of the M×N area. 17. In one example, a CCP candidate list may be constructed for chroma blocks. a. In one example, a first syntax element (SE) may be signaled to indicate whether a CCP candidate in the list is applied to the current chroma block. (This may be expressed as "the block is coded using the CCP candidate list mode"). i. For example, SE can be a logo. ii. For example, SE can be encoded and decoded by context. b. For example, the first SE may be signaled in a conditional manner. i. For example, the first SE may be signaled only when CCP is applied. ii. For example, the first SE may be signaled only when CCP is applied and a specific mode is applied. 1) The specific mode may be CCLM. 2) The specific mode may be CCCM. c. In one example, a second syntax element (SE) may be signaled to indicate which CCP candidate is applied. i. For example, SE can be an index. ii. For example, SE can be binarized into a truncated unary code. 1) For example, the maximum value of SE may be S-1, where S is the maximum size of the candidate list. iii. For example, the first binary bit of SE can be encoded and decoded by the context. d. For example, the second SE may be signaled in a conditional manner. i. For example, the second SE may be signaled only when the first SE indicates that a CCP candidate in the list is applied. e. In one example, whether the CCP candidate list mode is applicable can be signaled in the VPS / DPS / SPS / PPS / picture header / slice header / etc. f. In one example, the maximum size / length of the CCP candidate list may be signaled in the VPS / DPS / SPS / PPS / picture header / slice header / etc. 18. In one example, the CCP candidate list may include at least one CCP candidate stored in a spatially neighboring block, which may be adjacent to or not adjacent to the current block (assuming that the upper left position of the current block is (Xt, Yt), and the width and height of the current block are W and H, respectively). a. In one example, a set of locations are checked to find stored CCP information. i. For example, if the type of stored CCP information associated with a location is NonCCP, the location is skipped. 1) Alternatively, if the type of stored CCP information associated with the location is NonCCP, then the location is placed in a backup location list. ii. For example, if the type of the stored CCP information associated with the location is not NonCCP, then the stored CCP information is attempted to be appended to the list. b. In one example, a set of positions (Xi, Yi) to be checked in sequence can be derived from positions near the current block to positions farther away from the current block. i. For example, the position can be checked in a cycle-by-cycle manner. For one cycle, several positions are checked, and the next cycle is executed. ii. In one example, the position to be checked in the cycle is: (Xt-NDHor-1,Yt+H+NDVer-1),(Xt+W+NDHor-1,Yt-NDVer-1),(Xt+(W>>1),Yt-NDVer-1),(Xt-NDHor-1,Yt+(H>>1)),(Xt-NDHor-1,Yt-NDVer-1) NDHor and NDVer are different for different cycles. iii. In one example, the position to be checked for period k is derived as: NDHor=(k==0?W / 2:W*k); NDVer=(k==0?H / 2:H*k). iv. In one example, the location to be checked may be different for different cycles. c. In one example, the set of positions (Xi, Yi) to be checked can be the same as the set of positions checked when building the Merge list. d. In one example, the set of positions (Xi, Yi) to be checked may be the same as the set of positions checked when constructing the sub-block based Merge list. 19. In one example, when trying to put the stored CCP information as a candidate (referred to as a potential candidate) into the CCP candidate list, it can be compared with at least one candidate already in the CCP candidate list. a. In one example, all candidates in the list can be compared to potential candidates. b. In one example, if a candidate already in the CCP candidate list is the same as or similar to the potential candidate, the potential candidate cannot be put into the CCP candidate list. c. In one example, two CCP candidates are determined to be different if: i. Different CCP types. ii. The number of models is different. iii. If the CCP has multiple models, the threshold is different. iv. At least one model is different. v. Luma sample offset is different (applicable only when type is CCCM or GL-CCCM or GLM or CCCM using un-downsampled luma samples). vi. Sample point position displacement is different (applicable only when the type is GL-CCCM). 20. In one example, when a CCP candidate in the list is used to generate a prediction for the current block, the CCP will be executed according to the CCP information. a. CCCM, CCLM, 4 types of GLM using different downsampling filters, 4 types of GLM with luma using different downsampling filters, GL-CCCM, and CCCM using non-downsampled luma samples can be applied to the current block based on the candidate CCP type. b. Based on the number of candidate models and the threshold, one model or multiple models with at least one threshold can be used. c. Candidate luma sample value offsets can be added to or subtracted from luma samples to be put into the CCP model (which may be downsampled). i. This process is applicable only when the type is CCCM, or GL-CCCM, or GLM, or CCCM using non-downsampled luma samples. d. Sample point position displacement(s) can be added to or subtracted from the position coordinates to be placed in the CCP model. i. This procedure is applicable only when the type is GL-CCCM. e. How to obtain the downsampled brightness samples can be based on the CCP type. i. The downsampled luma samples may be obtained following the downsampling method required by the CCP mode corresponding to the type. 21. In an example, the prediction values ​​generated by the CCP candidate may be modified before being used to obtain the reconstructed sample values. a. In one example, an offset D may be added to or subtracted from the predicted value. b. In one example, the offset may be derived based on luma / chroma samples of a template calculated using reconstructed samples neighboring the current block (referred to as a "template"). Figures 35A to 35C Possible templates are shown separately. i. In one example, if the reconstructed samples on the left side of the current block are available, the template may be composed of the reconstructed samples on the left side of the current block. ii. In one example, if the reconstructed samples above the current block are available, the template may be composed of the reconstructed samples above the current block. iii. In one example, if the reconstructed samples above / left of the current block are available, the template may be composed of the reconstructed samples above or to the left of the current block. iv. The corresponding luma samples of the template can be downsampled in the same way as the luma samples inside the current block. c. In one example, if the CCP type requires N models (such as two models), then N offsets (denoted as {D 0 ,…,D N-1}) can be derived. i. Offset D i can be added to or subtracted from the predicted values ​​generated by model i. d. In one example, a CCP method indicated by the type of CCP candidate may be applied to the template. i. For example, for the kth sample point of the template, S k =R k -P k is calculated, where R k and P k They represent the reconstructed sample value of the kth sample point and the predicted value using CCP respectively. 1) For example, D is calculated as {S k}average value. 2) For example, suppose S k The number of is M, then D is calculated as D = sign (sum) × ((|sum| + off) >> W), where and ii. For example, for the kth sample point of model i using the template, S i k =R i k -P i k is calculated, where R i k and P i k They represent the reconstructed sample value of the k-th sample point using model i and the predicted value using CCP, respectively. 1) For example, D i is calculated as {S i k}average value. 2) For example, suppose S i k If the number is M, then D is calculated as D i =sign(sum)×((|sum|+off)>>W), where and iii. In one example, no division operation is used to calculate D or D i . 1) For example, a lookup table can be used to calculate D or D i . e. For example, only certain types of CCPs can apply modifications, such as CCLM and CCCM with multiple models. i. For example, the types of CCLM, CCLM with multiple models, CCCM with multiple models, and GLM can apply modifications. 22. In one example, candidates with type "non-adjacent" may be put into the candidate list. a. Information includes location (x, y). b. If such a candidate is used to predict the current block, then the CCP model(s) can be derived using the samples referenced by (x, y), as required by items 1 to 15. c. In one example, the locations stored in the backup location list disclosed in item 18 may be checked to place valid locations into a candidate list. 23. In one example, construction of the candidate list may be terminated if the number of candidates in the list is M and M=D+1, where D is an index indicating the selected candidate. 24. In one example, if all possible potential candidates are checked and the size of the candidate list is less than S (where S is the maximum number of candidates), a default candidate may be put into the list to complete the list. 25. In one example, the CCP candidate list may include at least one candidate obtained from a history-based table. a. The history table can be an online table. b. The history table can be a storage table. c. To construct the CCP candidate list, potential candidates may be examined in order. i. For example, the order can be (1) CCP information stored in spatially adjacent / non-adjacent blocks; (2) CCP candidates with type "non-adjacent"; (3) history-based candidates from the online table; (4) history-based candidates from the stored table; (5) default candidates. ii. For example, the order can be (1) CCP information stored in spatially adjacent blocks; (2) CCP information stored in spatially non-adjacent blocks; (3) CCP candidates with type "non-adjacent"; (4) history-based candidates from the online table; (5) history-based candidates from the stored table; (6) default candidates. iii. For example, the order can be (1) CCP information stored in spatially adjacent blocks; (2) CCP information stored in spatially non-adjacent blocks; (3) history-based candidates from the online table; (4) history-based candidates from the stored table; (5) CCP candidates with type "non-adjacent"; (6) default candidates. iv. For example, the order can be (1) CCP information stored in spatially adjacent blocks; (2) history-based candidates from the online table; (3) CCP information stored in spatially non-adjacent blocks; (4) CCP candidates with type "non-adjacent"; (5) history-based candidates from the stored table; (6) default candidates. v. Any type of candidate in the example order can be removed from it. vi. Any other order of potential candidates of these categories. 26. In one example, if a chroma block is encoded and decoded by using at least one CCP candidate, CCP information of the CCP candidate may be stored. a. The storage method can follow the method disclosed in item 16. 27. In one example, if a chroma block is coded by using at least one CCP candidate, the CCP information of the CCP candidate may be put into a history-based table. a. The process of placing CCP information into history-based tables can follow the process described in Section 2.27. 2.29. CCCM using multiple downsampling filters It is proposed to apply multiple downsampling filters to a set of reconstructed luma samples in CCCM. A linear combination of these downsampled reconstructed samples is multiplied by the derived filter coefficients to form the final chroma prediction value. The horizontal or vertical position of the center luma sample can also be considered in the proposed model. The coefficients are derived by Gaussian elimination as currently used in the CCCM mode in ECM. The following cross-component model is tested as an additional CCCM mode, where the mode index is signaled in the bitstream: (1) Model 1: predChroma=c0*H(C)+c1*G1(C)+c2*G2(C)+c3*G3(C)+c4*P(H(C))+c5*P(G1(C))+c6*P(G2(C))+c7*X+c8*Y+c9*B, (2) Model 2: predChroma=c0*H(C)+c1*H(W)+c2*H(E)+c3*G1(C)+c4*G1(W)+c5*G1(E)+c6*P(H(C))+c7*P(H(W))+c8*P(H(E))+c9*X+c10*B, (3) Model 3: predChroma=c0*H(C)+c1*H(NE)+c2*H(SW)+c3*G3(C)+c4*G3(NE)+c5*G3(SW)+c6*P(H(C))+c7*P(H(NE))+c8*P(H(SW))+c9*Y+c10*B, Where H(·), G1(·), G2(·), G3(·) are Figure 36 The various downsampling filters shown in FIG. 4 , C represents the current chroma sample position, and N, S, W, E, NE, SW are as follows Figure 37 The position around C shown, c i Are filter coefficients, P and B are nonlinear terms and bias terms, and X and Y are the horizontal and vertical positions of the center luma sample relative to the upper left coordinate of the block. Note that model 1 is a 1x1 prediction shape using only the current chroma sample, while the other models are unidirectional prediction models using 3 chroma samples. Figure 36 Various downsampling filters used in the proposed cross-component model are shown. Figure 37 The positions of the chroma samples are shown. 2.30. Non-local cross-component prediction in ECM The CCP candidate mode disclosed in 2.27 and 2.28 may also be referred to as "cross-component Merge mode". The CCP candidates disclosed in 2.27 and 2.28 may also be referred to as cross-component Merge candidates. The CCP candidate list disclosed in 2.27 and 2.28 may also be referred to as a cross-component Merge candidate list. In this document, CCP Candidate Mode, Cross-Component Merge Mode, and CCP Merge Mode may have the same meaning. In this document, CCP candidate, cross-component merge candidate, and CCP merge candidate may have the same meaning. In this document, CCP candidate list, cross-component merge candidate list, and CCP merge candidate list may have the same meaning. In ECM-9.0, the CCP Merge candidate list is constructed from spatially adjacent, spatially non-adjacent, or history-based candidates. After including these candidates, default models are further included to fill the remaining empty positions in the Merge list. To remove redundant CCP models from the list, a deduplication operation is applied. After the list is constructed, the CCP models in the list are reordered based on the SAD costs, which are obtained using the neighboring templates of the current block. More details are described below. Spatially adjacent and non-adjacent candidates The positions and inclusion order of spatially adjacent and non-adjacent candidates are the same as those defined for conventional inter-frame Merge prediction candidates in ECM. History-based candidates A history-based table is maintained to include the most recently used CCP models, and the table is reset at the beginning of each CTU row. If the current list is not full after including spatially adjacent and non-adjacent candidates, a CCP model in the history-based table is added to the list. Default Candidate CCLM candidates with default scaling parameters are considered only when the list is not full after including spatially adjacent, spatially non-adjacent, or history-based candidates. If there are no candidates with single-model CCLM mode in the current list, the default scaling parameters are {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, 4 / 8, -4 / 8, 5 / 8, -5 / 8, 6 / 8}. Otherwise, the default scaling parameters are {0, scaling parameters of the first CCLM candidate + {1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, 4 / 8, -4 / 8, 5 / 8, -5 / 8, 6 / 8}}. The offset parameters are derived based on the default scaling parameters, the average neighboring reconstructed luma sample values, and the average neighboring reconstructed Cb / Cr sample values. A flag is signaled to indicate whether CCP Merge mode is applied. If CCP Merge mode is applied, an index is signaled to indicate which candidate model is used by the current block. In addition, when the current CU is encoded by intra sub-partitioning (ISP) with a single tree, or when the current chroma codec block size is less than or equal to 16, CCP Merge mode is not allowed for the current chroma codec block. 2.31. Local Enhanced Cross-Component Prediction (LB-CCP) in ECM The LB-CCP method is proposed to enhance CCP by taking more advantage of local information including neighboring samples and predicted samples. LB-CCP involves three aspects. Aspect #1: The prediction samples of MM-CCLM / MM-CCCM can be filtered using neighboring samples. Figure 38 As shown, a 3×3 low-pass filter is applied to filter the prediction samples generated by MM-CCLM / MM-CCCM. For samples at the top / left boundary, the filtering window can include neighboring reconstructed samples. For internal samples, the filtering window only includes prediction samples, which can be padded. A flag is signaled to indicate whether filtering is applied for blocks encoded using MM-CCLM / MM-CCCM. Aspect #2: Template costs are calculated to implicitly determine the use of CCP. Figure 39 As shown, the template cost is derived by applying the candidate CCP to the template and calculating the SAD between the predicted samples and the reconstructed samples in the template. For CCCM / CCCM-L / CCCM-T modes, two template costs are calculated by using 6 rows or 2 rows of neighboring samples as training samples, respectively. For MM-CCCM / MM-CCCM-L / MM-CCCM-T modes, two template costs are calculated by deriving the average value as the threshold to separate the two models using the neighboring luma samples or the luma samples co-located with the current chroma block, respectively. For intra chroma fusion mode, two template costs are calculated by fusing the angular chroma prediction with MM-CCLM or MM-CCCM, respectively, and the weighted value depends on the minimum template cost. In all cases, the option with the lower template cost is selected as the final CCP method on the current block. Figure 38 Filters on the samples of MM-CCLM or MM-CCCM are shown. Figure 39 A template for a block is shown. 3. Question 1. The candidates in the CCP candidate list may not be in the best order. It is useful to put better candidates at the front with shorter index bits. 2. How to coordinate CCCM using multiple downsampling filters (MF-CCCM) and the CCP candidate list remains unclear. 3. More information can be inherited by CCP candidates. 4. Detailed solution The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way. In the following discussion, CCCM may refer to the original CCCM mode, or it may refer to variants of CCCM, such as CCCM-L, CCCM-T, MM-CCCM, MM-CCCM-L, MM-CCCM-T. In the following discussion, CCLM may refer to the original CCLM mode, or it may refer to variants of CCLM, such as CCLM-L, CCLM-T, MM-CCLM, MM-CCLM-L, MM-CCLM-T, etc. In the following discussion, cross-component prediction (CCP) may refer to any cross-component prediction, such as CCLM or CCCM or GLM or CCLM with sliding offset. 1. CCP candidates in the CCP candidate list can be reordered. a. The CCP candidate list may include different kinds of candidates, such as candidates having CCP information stored in adjacent / non-adjacent neighboring blocks and / or candidates having CCP information stored in a history-based table and / or CCP information derived from non-adjacent samples. b. In one example, whether to reorder and / or how to reorder can be signaled from the encoder to the decoder, such as at a block level / sequence level / picture group level / picture level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB or sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. c. In one example, the CCP candidate list can be reordered at the encoder and the decoder based on the same rules. d. In one example, assuming that the CCP candidate list before reordering is denoted as L, and the CCP candidate list after reordering is denoted as L', and the index of the CCP candidate transmitted to the decoder or derived by the decoder is denoted as k, the CCP candidate L'[k] can be used to decode the current block. e. In one example, whether and / or how to reorder candidates in the CCP candidate list may depend on the position of the candidate in the list. i. For example, only the first M candidates in the candidate list L can be reordered, where M is not greater than the size of L. f. In one example, whether and / or how to reorder the candidates in the CCP candidate list may depend on the type of CCP candidate and / or CCP information. i. In one example, candidates with certain characteristics may be placed forward in the re-ranking. ii. In one example, candidates with certain characteristics may be placed backward in the re-ranking. iii. In one example, candidates with certain characteristics may not be involved in the re-ranking. iv. Specific characteristics may be: 1) It is the default CCP candidate that populates the candidate list. 2) It has CCP information stored in at least one adjacent neighboring block. 3) It has CCP information stored in at least one non-adjacent neighboring block. 4) It has CCP information stored in a block at a specific location. 5) It has CCP information stored in history based tables. 6) It has CCP information stored in a history-based table that is updated online. 7) It has CCP information stored in a stored history based table. 8) It has CCP information stored in a history based table at a specific entry. 9) It has CCP information and / or CCP information derived from non-adjacent samples. 10) It has CCP information and / or CCP information derived from non-adjacent samples at a specific location. 11) It is associated with a specific CCP method, such as CCLM, or CCLM-T, or CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, or CCCM, or CCCM-T, or CCCM-L, or MM-CCCM, or MM-CCCM-T, or MM-CCCM-L, or GLM using a different downsampling filter, or GLM with luma using a different downsampling filter, or GL-CCCM, or CCCM using non-downsampled luma samples. g. In one example, the reordering process can be performed in a conditional manner. i. For example, if the number of candidates in the CCP candidate list is less than a threshold, the reordering process may be skipped. ii. For example, if W>=Tw, and / or W<=Tw, and / or H>=Th, and / or H<=Th, and / or W*H>=Ts, and / or W*H<=Ts, the reordering process may be skipped. iii. For example, the reordering process can be skipped based on codec information such as mode / QP / neighborhood information / color component / color format. h. In one example, the reordering process may be performed once for at least two components (such as Cb and Cr). i. In one example, the reordering process can be performed separately for different components (such as Cb and Cr). 2. The CCP candidates in the CCP candidate list may be reordered based on cost comparison. a. In one example, for each CCP candidate involved in the reordering process, a cost may be calculated in association with the candidate. b. In one example, the involved CCP candidates may be placed in ascending order based on the cost associated with the candidate. c. In one example, the involved CCP candidates may be placed in descending order based on the cost associated with the candidate. d. In one example, the cost may be a template cost, which is calculated using reconstructed samples adjacent to the current block, called a "template." Figures 40A to 40C Possible templates are shown separately. 1) In one example, if reconstructed samples on the left side of the current block are available, the template may be composed of the reconstructed samples on the left side of the current block. 2) In one example, if the reconstructed samples above the current block are available, the template may be composed of the reconstructed samples above the current block. 3) In one example, if reconstructed samples above / left of the current block are available, the template may be composed of the reconstructed samples above or to the left of the current block. e. The cost of the CCP candidate can be calculated in a process. The process can include at least one of the following two steps: i. Step 1: Cross-component predictions are derived on the template samples. 1) Cross-component prediction is applied to the template in the same / similar manner as the CCP method associated with the CCP candidate is applied on the current block. a) In one example, the luma samples corresponding to the template region may be obtained by a downsampling method required by the CCP candidate. 2) In one example, a cross-component prediction model of a CCP candidate that may be used to generate a prediction of the current block may be used to derive prediction samples of a template. a) In one example, when the prediction samples of the template are derived using the CCP model, the prediction samples may be modified. i. For example, an offset can be added to the prediction samples. ii. For example, the offset can be derived as disclosed in item 21 of Section 2.28. 3) In one example, a threshold used to separate at least two models in the current block (such as in MM-CCCM and MM-CCLM modes) may be used to separate models in the template. ii. Step 2: The distortion between the predicted and reconstructed samples of the template is calculated as the cost. 1) Distortion can be SAD, SSD, mean removed SAD, SATD, etc. f. Costs can be derived separately for different components (such as Cb component and Cr component). i. The cost on one component (such as Cb) can be used to reorder the candidates. ii. The total cost over multiple components (such as Cb and Cr) can be used to re-rank the candidates. 3. When building the CCP candidate list, potential CCP candidates can be reordered. a. For example, a potential CCP candidate may have characteristics. i. Features are default CCP candidates that populate the candidate list. ii. The feature has CCP information stored in at least one adjacent neighboring block. iii. The feature has CCP information stored in at least one non-adjacent neighboring block. iv. A feature has CCP information stored in a block at a specific location. v. Features have CCP information stored in history-based tables. vi. Features have CCP information stored in a history-based table that is updated online. vii. Features have CCP information stored in stored history-based tables. viii. Features have CCP information and / or CCP information derived from non-adjacent samples. b. In one example, all or some of the potential candidates may be checked and re-ranked. The top N potential candidates may be put into a candidate list. i. In one example, when the top N potential candidates are in the candidate list, the top N potential candidates may be kept in order. ii. In one example, as disclosed in items 1 and 2, reordering may be performed based on cost comparison. 4. In one example, when building the CCP candidate list, adjacent neighboring blocks may be checked before non-adjacent neighboring blocks. 5. In one example, whether adjacent neighboring blocks and / or non-adjacent neighboring blocks can be checked to construct the CCP candidate list may depend on the locations of the adjacent neighboring blocks and / or non-adjacent neighboring blocks. a. In one example, if the adjacent neighboring block and / or the non-adjacent neighboring block is not in the current CTU row, it may not be allowed to be checked to construct a CCP candidate. b. In one example, if the adjacent neighboring block and / or the non-adjacent neighboring block is not in the current CTU, it may not be allowed to be checked to construct the CCP candidate. c. In one example, if |y0-Y0|>Ts, the adjacent neighboring blocks and / or non-adjacent neighboring blocks may not be allowed to be checked to construct CCP candidates, where (x0, y0) is the upper left position of the adjacent block and (x0, y0) is the upper left position of the current CTU. d. In one example, if |x0-X0|>Ts, the adjacent neighboring blocks and / or non-adjacent neighboring blocks may not be allowed to be checked to construct CCP candidates, where (x0, y0) is the upper left position of the adjacent block and (X0, Y0) is the upper left position of the current CTU. e. In one example, if |y0-Y0|>Ts, the adjacent neighboring blocks and / or non-adjacent neighboring blocks may not be allowed to be checked to construct CCP candidates, where (x0, y0) is the upper left position of the neighboring block and (x0, y0) is the upper left position of the current block. f. In one example, if |x0-X0|>Ts, the adjacent neighboring blocks and / or non-adjacent neighboring blocks may not be allowed to be checked to construct CCP candidates, where (x0, y0) is the upper left position of the neighboring block and (x0, y0) is the upper left position of the current block. g. In the above items, "CTU" can be replaced by any other area unit (such as "VPDU"). 6. In one example, whether adjacent neighboring blocks and / or non-adjacent neighboring blocks can be checked to construct the CCP candidate list may depend on whether the dual-tree structure is applied. 7. In one example, whether luma samples can be used to derive the CCP model may depend on whether the dual-tree structure is applied. a. For example, when dual-tree is applied, luma samples are available as long as they are in the current CTU. b. For example, when dual tree is not applied, luma samples are available only if the block containing luma samples has been decoded. 8. In one example, a first syntax element (SE) may be signaled to indicate whether a CCP candidate in the list is applied to the current chroma block. (This may be expressed as "block is coded using CCP candidate list mode"), and a second SE may be signaled conditionally depending on the first SE. b. For example, if the first SE indicates that the block is coded using the CCP candidate list mode, the second SE may not be signaled. i. The second SE may indicate whether the type of CCLM mode is applied. ii. The second SE may indicate whether the type of CCCM mode is applied. c. For example, if the first SE indicates that the block is coded using the CCP candidate list mode, an index may be signaled to indicate the CCP candidate, and other information may not be signaled to indicate any other CCP mode. 9. In one example, whether the CCP candidate list mode is applicable to a block may depend on the width W and / or height H of the block. a. In one example, if the CCP candidate list mode is not applicable to a block, the SE indicating whether the CCP candidates in the list are applied to the current chroma block is not signaled. b. In one example, if W×H<=T, then the CCP candidate list mode is not applicable, where T is an integer, such as 8, or 16, or 32. c. In one example, if W×H>=T, then the CCP candidate list mode is not applicable, where T is an integer, such as 1024, or 2048, or 4096. d. In one example, if min{W, H}<=T, then the CCP candidate list mode is not applicable, where T is an integer, such as 4, or 8, or 16, or 32, or 64. e. In one example, if min{W, H}>=T, then the CCP candidate list mode is not applicable, where T is an integer, such as 4, or 8, or 16, or 32, or 64. f. In one example, if max{W, H}<=T, then the CCP candidate list mode is not applicable, where T is an integer, such as 4, or 8, or 16, or 32, or 64. g. In one example, if max{W, H}>=T, then the CCP candidate list mode is not applicable, where T is an integer, such as 4, or 8, or 16, or 32, or 64. 10. In one example, to populate the CCP candidate list, the check order can be a. Adjacent neighboring blocks. b. Non-adjacent neighboring blocks. c. History-based neighboring blocks. d.Default candidate. 11. In one example, Figure 41 As shown, specific adjacent neighboring blocks may be checked in sequence to populate the CCP candidate list. Figure 41 Adjacent neighboring blocks are shown. a. In one example, the specific adjacent neighboring blocks may be A1, A3, A4, A6, A7, and the order may be: i.A3, A6, A4, A7, A1; ii.A3, A6, A7, A4, A1; iii.A6, A3, A7, A4, A1; iv.A6, A3, A4, A7, A1. b. In one example, the specific adjacent neighboring blocks may be A1, A2, A3, A4, A5, A6, A7, and the order may be: i.A3, A6, A4, A7, A1, A2, A5; ii.A3, A6, A7, A4, A1, A2, A5; iii.A6, A3, A7, A4, A1, A2, A5; iv.A6, A3, A4, A7, A1, A2, A5; v.A3, A6, A4, A7, A1, A5, A2; vi.A3, A6, A7, A4, A1, A5, A2; vii.A6, A3, A7, A4, A1, A5, A2; viii.A6, A3, A4, A7, A1, A5, A2; ix.A3, A6, A4, A7, A2, A5, A1; x.A3, A6, A7, A4, A2, A5, A1; xi.A6, A3, A7, A4, A2, A5, A1; xii.A6, A3, A4, A7, A2, A5, A1; xiii.A3, A6, A4, A7, A5, A2, A1. xiv.A3, A6, A7, A4, A5, A2, A1; xv.A6, A3, A7, A4, A5, A2, A1; xvi.A6, A3, A4, A7, A5, A2, A1. 12. In one example, Figure 41 As shown, specific adjacent neighboring blocks may be checked in sequence to populate the CCP candidate list. 13. In one example, two CCP candidates are determined to be the same if: i. CCP types are CCLM or GLM. ii. The number of models is the same. iii. The parameters of the model in the first CCP candidate do not differ from the corresponding parameters of the corresponding model in the second CCP candidate. 1) The parameter can be the offset parameter of the linear model. In other words, assuming the model is of the form y=ax+b, where a and b are parameters, then b is the offset parameter. 14. In one example, when deriving the average value D as disclosed in item 21 of Section 2.28, a process without division operation may be applied. a. For example, D is calculated as {S k}. Assume that S k The number of is M, and Then sum and M can be the inputs of the process, and D can be the output. b. For example, the use of a procedure may depend on whether sum < 0. i. For example, if sum < 0, then -sum can be input to the process and the output can be -D, which will be converted to D through the negation operation. c. For example, a process may include at least one predefined table. i. For example, the table can be divTable

[16] = {0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}. d. For example, at least one log2 operation may be involved in the process. i. For example, the process can be derived e. In one example, the index NormNum of the entry in the reference divTable may be derived. i. The derivation can depend on M. ii. The derivation can depend on x. iii. In one example, NormNum=(M< <4> >x)&15. f. In one example, by setting a specific bit to 1, the value v can be calculated using the selected entry divTable[NormNum] in divTable. i. For example, v = 8 | divTable[NormNum]. g. In one example, x can be modified based on NormNum. i. For example, if NormNum is not equal to 0, then x=x+1. h. In one example, the shift S can be derived based on x. i. For example, S = 13 - x. i. In one example, the value retVal can be derived based on whether S is less than 0. i. In one example, if S<0, then retValue=(sum×v+(1<<(−S−1))>>(−S). ii. In one example, if S>=0, then retValue=(sum×v)>>S. j. In one example, D can be derived from retVal through a shift operation. i. For example, D=retVal>>16. k. In one example, the process or part of the process can also be used as a replacement for a division operation in other codec tools. i. For example, when deriving a cross-component model, it can be used in CCLM or CCCM. ii. For example, it can be used in MM-CCLM or MM-CCCM to derive the threshold for classifying samples. iii. For example, it can be used in affine mode to derive an affine model or corner motion vectors (CPMV). 15. In one example, the prediction generated by the first CCP candidate in the list can be fused with the second prediction to obtain a prediction that is used in another step. a. In one example, two predictions are fused by performing a weighted sum. i. For example, weighting values ​​can be position-dependent. ii. For example, the weighting value may be indicated through signaling. iii. For example, the weighted value may be a fixed value. b. In one example, a second prediction may be generated from a second CCP candidate. c. In one example, the second prediction may be a specific CCP prediction, such as CCLM or CCCM. d. In one example, the second prediction may be a specific angular prediction mode, such as DC mode. e. In one example, the second prediction may be a specific angular prediction mode depending on the luma component, such as DM mode. f. In one example, the second prediction may be a specific angular prediction mode depending on neighboring samples, such as DIMD or TIMD mode. g. In one example, SE may be signaled to indicate whether such fusion is applied. h. In one example, SE may be signaled to indicate the first prediction and / or the second prediction to be fused. i. In one example, the index of the CCP candidate list may indicate whether such fusion is applied. j. In one example, the index of the CCP candidate list may indicate the first prediction and / or the second prediction to be fused. 16. In one example, whether the CCP candidate list mode is applicable to a block may depend on the codec information of the current chroma block and / or the co-located luma block. a. Codec information may include codec mode, CCP model, reconstructed samples, QP, partitioning method, dual-tree structure, etc. b. In one example, the CCP candidate list mode may be excluded from the second mode. i. In one example, if the CCP candidate list mode is used, the second mode does not apply. ii. In one example, if the second mode is used, the CCP candidate list mode is not applicable. iii. In one example, if a mode is not applicable, the SE of the current block related to the mode may not be signaled. c. In one example, if ISP mode is used, CCP candidate list mode is not applicable. d. In one example, if the dual-tree structure is not applied and the ISP mode is used, the CCP candidate list mode is not applicable. 17. In one example, when ISP mode or any other sub-TU / sub-PU / sub-CU (all denoted as sub-block) method is applied, CCP candidate list mode may be used. a. In one example, whether to apply CCP candidate list mode may be determined for the entire block, and each sub-block may share the determination. i. A single SE may be signaled for the entire block to indicate whether the CCP candidate list mode is applied for all sub-blocks. b. In one example, whether to apply the CCP candidate list mode may be determined individually for each sub-block. i.SE may be signaled for a sub-block to indicate whether the CCP candidate list mode is applied for the sub-block. c. In one example, if the CCP candidate list mode is used, a single CCP candidate list can be derived for the entire block, and each sub-block can share a single CCP candidate. d. In one example, if the CCP candidate list mode is used, a separate CCP candidate list can be derived for each sub-block. i. In one example, the codec information of the first sub-block can be used to derive the CCP candidate list of the second sub-block. ii. In one example, the codec information of the first sub-block may be used to reorder the CCP candidate list of the second sub-block. iii. In one example, the codec information of the first sub-block may be used to derive an offset of a CCP model in a CCP candidate list for the second sub-block. iv. Codec information may include codec mode, CCP model, reconstructed samples, etc. 18. In one example, the CCP information may include information about the MF-CCCM. a. In one example, the CCP information may include whether the associated CCP model uses MF-CCCM. b. In one example, the CCP information may include whether the mode is MF-CCCM. c. In one example, the CCP information may include parameters associated with a CCP model using MF-CCCM. d. In one example, information about MF-CCCM may be stored in a block together with the CCP information. i. In one example, the block is coded using MF-CCCM mode. ii. In one example, the block is coded using the CCP candidate list mode, and the selected candidate is using the MF-CCCM mode. e. In one example, information about MF-CCCM may be stored in a history-based table together with CCP information after a block is encoded / decoded. i. In one example, the block is coded using MF-CCCM mode. ii. In one example, the block is coded using the CCP candidate list mode, and the selected candidate is using the MF-CCCM mode. f. In one example, information about the MF-CCCM may be stored in the CCP candidate. g. In one example, when trying to put potential candidates into a history-based table, information about MF-CCCM can be used to compare whether two candidates are identical or similar. h. In one example, when trying to put potential candidates into a candidate list, information about the MF-CCCM can be used to compare whether two candidates are identical or similar. i. In one example, information about the MF-CCCM may be used to reorder candidates in the CCP candidate list. i. In one example, as indicated by the use of MF-CCCM associated with the CCP candidate, the MF-CCCM may be used to generate prediction values ​​for the template. j. In one example, as indicated by the use of MF-CCCM associated with the CCP candidate, information about the MF-CCCM may be used to generate a prediction value for the current block. k. All items in Sections 2.27 and 2.28 can be applied as MF-CCCM together with the CCP method. 1. All claims disclosed in this document can be applied together with the CCP method as MF-CCCM. 19. In one example, a first prediction P0 generated by a CCP candidate may be fused with a second prediction P1 to generate a prediction P2 to be used in further processing. a. In one example, P2 = W0 × P0 + W1 × P1, where W0 and W1 are weighted values, for example, W0 = W1 = 0.5. b. In one example, P2 = (W0 × P0 + W1 × P1 + offset) >> shift, where W0 and W1 are integer weights. Offset and shift are integers. In one example, offset = 1 << (shift - 1). Some examples of (W0, W1, Shift) are as follows: (1, 1, 1), (2, 2, 2), (1, 3, 2), (3, 1, 2), (1, 7, 3), (7, 1, 3), (3, 5, 3), (5, 3, 3), etc. c. In one example, the weight value and / or displacement may depend on the width and / or height of the current block. d. In one example, the weighting value and / or the displacement may depend on the location of the prediction sample. e. In one example, the weighting values ​​and / or displacements may depend on the method of generating P1. f. In one example, the weighting values ​​and / or displacements may depend on the type of CCP candidate that generated P0. g. In one example, the weighting values ​​and / or displacements may depend on the color format (such as 4:4:4 or 4:2:0 or 4:2:2). h. In one example, the weighting values ​​and / or displacements may depend on the color components. i. In one example, P1 can be generated by a predefined prediction mode. i. The predefined prediction mode may be cross-component prediction, such as CCLM, or CCLM-T, or CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, or CCCM, or CCCM-T, or CCCM-L, or MM-CCCM, or MM-CCCM-T, or MM-CCCM-L, or GLM using different downsampling filters, or GLM with luma using different downsampling filters, or GL-CCCM, or CCCM using non-downsampled luma samples, or MF-CCCM. ii. The predefined prediction mode may be non-cross-component prediction, such as DM mode, DIMD mode, TIMD mode, DC mode, planar mode, horizontal mode, vertical mode, etc. j. In one example, P1 may be generated by a prediction mode determined individually for a block / sample. i. For example, the prediction mode can be determined according to the index of the CCP candidate. ii. For example, the prediction mode may be determined according to the width and / or height of the current block. iii. For example, the prediction mode may be determined according to the type of CCP candidate. iv. In one example, the prediction mode may depend on the color format (such as 4:4:4, or 4:2:0, or 4:2:2). v. In one example, the prediction mode may depend on the color component. k. Whether and / or how to apply the fusion operation may depend on the CCP candidate. i. For example, the fusion operation may be applied only when the index of the CCP candidate is less than T (eg, T=1, or 2, or 3, or 4, etc.). ii. For example, the fusion operation can be applied only when the type of the CCP candidate is in a predefined set. 1. In one example, a flag may be signaled to indicate whether fusion is to be applied. i. For example, a flag can be conditionally signaled. 1) For example, the flag is signaled only when the CCP candidate mode is applied. 2) For example, the flag is signaled only when the index of the CCP candidate is less than T (eg, T=1, or 2, or 3, or 4, etc.). m. In one example, whether fusion is to be applied may be indicated by a CCP candidate index. i. For example, a candidate with index K in the CCP list may indicate a fusion mode, where P0 is generated by a CCP candidate not having index K in the list, or a CCP candidate not in the list except for use in the fusion operation. n. In one example, a model and associated information for generating the second prediction may be stored. i. For example, the model may be a CCLM, or CCLM-T, or CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L, or CCCM, or CCCM-T, or CCCM-L, or MM-CCCM, or MM-CCCM-T, or MM-CCCM-L, or a GLM using a different downsampling filter, or a GLM with luma using a different downsampling filter, or a GL-CCCM, or a CCCM using non-downsampled luma samples, or a MF-CCCM model. ii. In one example, the models and associated information may be processed in a similar manner as the models and associated information of CCP-encoded blocks. 1) In one example, the type of model and associated information may depend on the method of generating the second prediction, such as MM-CCCM. 2) In one example, the model and associated information may be stored in a cell block. 3) In one example, the models and associated information may be stored in a history-based table. 4) In one example, the model and associated information may be used to generate CCP candidates in a CCP candidate list. 5) In one example, when a CCP candidate list is stored or generated, the model and associated information may be compared to other models and associated information. 6) In one example, when generating a CCP candidate list, the models and associated information may be involved in the re-ranking process. 20. In one example, rows and / or columns (referred to as lines) that are not adjacent to the current block can be used to obtain a model of CCLM, or CCLM-T, or CCLM-L, or MM-CCLM, or MM-CCLM-T, or MM-CCLM-L. a. In one example, an index may be signaled to indicate the row and / or column of the reconstructed chroma samples and corresponding luma samples (which may be obtained by downsampling) adjacent to the current block. b. Figure 42 and Figure 43 An example of possible rows that may be used to derive CCLM model(s) is shown, where rows 1-4 are not adjacent to the current block. Figure 42 The specific row of sample points used to derive the CCLM model is shown. Figure 43 The particular row of samples used to derive the CCLM model including the top left sample is shown. 21. In one example, whether and / or how filtering is applied to the CCP-generated prediction can be inherited from the CCP candidate. In the following discussion, the information "whether and / or how filtering is applied to the CCP-generated prediction" is referred to as "filtering information." a. For example, the filtering can be the 3x3 filtering introduced by LB-CCP. i. For example, the filter flags of the 3x3 filter introduced by LB-CCP can be inherited. b. In one example, the filtering information can be considered as information associated with the CCP-encoded block. i. In one example, the filtering information may be stored in unit blocks. ii. In one example, the filtering information may be stored in a history-based table. iii. In one example, the filtering information may be used to generate CCP candidates in a CCP candidate list. iv. In one example, the filtering information of one CCP candidate or potential candidate may be compared with the filtering information of another CCP candidate or potential candidate. 1) In one example, if the filtering information of two CCP candidates or potential candidates is different, then the two CCP candidates or potential candidates may be determined to be different. 2) Alternatively, when comparing two CCP candidates or potential candidates, the filtering information may be ignored. c. In one example, only when the CCP candidate is of a specific type, the associated filtering information may be used. i. For example, the type can be a multi-model CCP method. d. In one example, whether and / or how filtering is applied to CCP-generated predictions may depend on inherited filtering information. i. For example, if the inherited filtering flag is true, the CCP-generated prediction may be filtered, and if the inherited filtering flag is false, the CCP-generated prediction may not be filtered. e. In one example, the inherited filtering information of a CCP candidate can be stored and used by subsequent blocks. 22. In one example, decoder-derived determinations may be applied to CCP candidates. a. For example, at least two options may be applied to the template along with the CCP candidate to obtain a cost, and the option with the lower cost may be selected. i. For example, for CCP candidates of a specific type (such as MM-CCCM), two template costs are calculated by using the mean or inherited mean derived from neighboring luma samples as a threshold to separate the two models. Overview 23. The syntax elements disclosed above can be binarized as flags, fixed length codes, EG(x) codes, unary codes, truncated unary codes, truncated binary codes, etc. It can be signed or unsigned. 24. The syntax elements disclosed above can be encoded or decoded using at least one context model, or they can be bypassed. 25. The syntax elements disclosed above may be signaled in a conditional manner. a. SE is signaled only if the corresponding function is applicable. b. SE is signaled only if the dimensions (width and / or height) of the block meet the conditions. 26. The syntax elements disclosed above can be transmitted by signal at block level / sequence level / picture group level / picture level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB or sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 27. Whether and / or how to apply the method disclosed above can be transmitted through signals at block level / sequence level / picture group level / picture level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB or sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 28. Whether and / or how to apply the above disclosed method may depend on coded information such as block size, color format, single / dual tree partitioning, color component, slice / picture type. 29. The proposed method disclosed in this document can be used in other codecs that require chroma fusion.

[0104] Additional details will be described below. Figure 44 FIG4 is a flowchart of a method 4400 for video processing according to an embodiment of the present disclosure. The method 4400 is implemented for converting between a current video block of a video and a bitstream of the video.

[0105] At block 4410, a cross-component prediction (CCP) model for a current video block is determined based on at least one row of samples that is non-adjacent to the current video block, the at least one row comprising at least one of: rows of non-adjacent samples or columns of non-adjacent samples.

[0106] At block 4420, conversion is performed based on the CCP model. In some embodiments, conversion may include encoding the current video block into a bitstream. Alternatively or additionally, conversion may include decoding the current video block from the bitstream.

[0107] Method 4400 enables determining a CCP model based on non-adjacent samples. In this way, codec efficiency and codec effectiveness can be improved.

[0108] In some embodiments, the CCP model includes at least one of the following: a cross-component linear model (CCLM), a CCLM based on top neighboring samples of the current video block (CCLM-T), a CCLM based on left neighboring samples of the current video block (CCLM-L), a multi-model based CCLM (MM-CCLM), a multi-model based CCLM-T (MM-CCCM-T), or a multi-model based CCLM-L (MM-CCCM-L).

[0109] In some embodiments, at least one index of at least one row is included in the bitstream, the at least one index indicating at least one of the following: reconstructed chroma samples and corresponding luma samples of at least one row adjacent to the current video block, or reconstructed chroma samples and corresponding luma samples of at least one column adjacent to the current video block. For example, the corresponding luma samples may be obtained by downsampling.

[0110] In some embodiments, the at least one row includes a plurality of non-adjacent rows above the current video block. For example, the number of the plurality of non-adjacent rows is predetermined, such as 4 or any other positive integer.

[0111] In some embodiments, the widths of the plurality of non-adjacent rows are the same as the width of the current video block. Figure 42 In the example of , the current video block is video unit 4200. The plurality of non-adjacent rows may include Figure 42 4212, 4222, 4232, and 4242. It should be understood that adjacent rows 4202 may or may not be used to determine the CCP model.

[0112] In some embodiments, the widths of the plurality of non-adjacent rows are greater than the width of the current video block, and the plurality of non-adjacent rows include at least one sample above and to the left of the current video block. Figure 43 In the example of , the current video block is video unit 4300. Multiple non-adjacent rows can be Figure 434300. It should be understood that adjacent rows 4302 may or may not be used to determine the CCP model. The width of adjacent rows 4302 may be the same as the width of video unit 4300. In some other embodiments, the width of adjacent rows may be greater than the width of video unit 4300.

[0113] In some embodiments, the at least one row includes a plurality of non-adjacent columns to the left of the current video block. For example, the number of the plurality of non-adjacent columns is predetermined, such as 4 or any other positive integer.

[0114] In some embodiments, the heights of the plurality of non-adjacent columns are the same as the height of the current video block. Figure 42 As shown in the example of , the plurality of non-adjacent columns may include column 4214, column 4224, column 4234, and column 4244. It should be understood that adjacent column 4204 may or may not be used to determine the CCP model.

[0115] In some embodiments, the heights of the plurality of non-adjacent columns are greater than the height of the current video block, and the plurality of non-adjacent columns include at least one sample above and to the left of the current video block. Figure 43 As shown in the example of , the plurality of non-adjacent columns may include column 4314, column 4324, column 4334, and column 4344. It should be understood that adjacent column 4304 may or may not be used to determine the CCP model. The height of adjacent column 4304 may be greater than the height of video unit 4300. Alternatively, in other embodiments, the height of the adjacent column may be the same as the height of video unit 4300.

[0116] As used herein, row 4202, row 4212, row 4222, row 4242, row 4302, row 4312, row 4322, row 4332, and row 4342 may be referred to as rows of samples. Similarly, columns 4204, 4214, 4224, 4234, 4244, 4304, 4314, 4324, 4334, and 4334 may be referred to as rows of samples. For example, row 4212 and / or column 4214 may be referred to as row 1 of video unit 4200. Row 4312 and / or column 4314 may be referred to as row 1 of video unit 4300.

[0117] In some embodiments, if the row index is indicated as 2, for example, the rows and / or columns used to obtain the CCP model may include Figure 42 Alternatively, the rows and / or columns used to obtain the CCP model may include Figure 43 4322 and / or column 4324 in .

[0118] Already referenced Figure 42and Figure 43 Several example non-adjacent rows and non-adjacent columns are described. It should be understood that Figure 42 and Figure 43 The number of non-adjacent rows or columns and the heights and / or widths of these rows and / or columns shown in FIG are for illustrative purposes only and do not imply any limitation. Different numbers of non-adjacent rows and / or columns with different widths and / or heights can be used to obtain CCP models such as (multiple) CCLM models. The scope of the embodiments of the present disclosure is not limited in this regard.

[0119] In some embodiments, syntax elements in the bitstream are binarized into at least one of: a flag, a fixed-length code, an Exponential Golomb (x) (EG(x)) code, a unary code, a truncated unary code, or a truncated binary code.

[0120] In some embodiments, syntax elements are signed or unsigned.

[0121] In some embodiments, syntax elements in a bitstream are encoded or decoded using at least one context model, or are bypass encoded or decoded.

[0122] In some embodiments, a syntax element is included in the bitstream based on a condition that a function associated with the syntax element is applicable.

[0123] In some embodiments, if the dimensions of the current video block meet a condition, the syntax element is included in the bitstream.

[0124] In some embodiments, the syntax element is at at least one of: block level, sequence level, group of pictures level, picture level, slice level, or slice group level.

[0125] In some embodiments, the syntax element is in at least one of the following codec structures: codec tree unit (CTU), codec unit (CU), transform unit (TU), prediction unit (PU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, or slice group header.

[0126] In some embodiments, information about whether to apply method 4400 and / or how to apply method 4400 is included in the bitstream.

[0127] In some embodiments, the information is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.

[0128] In some embodiments, the information is indicated in at least one of the following codec structures: codec tree unit (CTU), codec unit (CU), transform unit (TU), prediction unit (PU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, or slice group header.

[0129] In some embodiments, the information is based on coded information, wherein the coded information includes at least one of: block size, color format, single-tree partitioning or dual-tree partitioning, color component, slice type, or picture type.

[0130] In some embodiments, method 4400 is used in a codec that requires chroma fusion.

[0131] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. In the method, a cross-component prediction (CCP) model for a current video block of the video is determined based on at least one row of samples that is non-adjacent to the current video block, where the at least one row includes at least one of the following: rows of non-adjacent samples or columns of non-adjacent samples. The bitstream is generated based on the CCP model.

[0132] According to further embodiments of the present disclosure, a method for storing a video bitstream is provided. In this method, a cross-component prediction (CCP) model for a current video block of the video is determined based on at least one row of samples that is non-adjacent to the current video block, where the at least one row includes at least one of the following: rows of non-adjacent samples or columns of non-adjacent samples. A bitstream is generated based on the CCP model. The bitstream is stored in a non-transitory computer-readable recording medium.

[0133] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.

[0134] Item 1. A method for video processing, comprising: for conversion between a current video block of a video and a bitstream of the video, determining a cross-component prediction (CCP) model for the current video block based on at least one row of samples that are not adjacent to the current video block, the at least one row comprising at least one of the following: rows of non-adjacent samples, or columns of non-adjacent samples; and performing the conversion based on the CCP model.

[0135] Item 2. The method of Item 1, wherein the CCP model comprises at least one of the following: a cross-component linear model (CCLM), a CCLM based on top neighboring samples of the current video block (CCLM-T), a CCLM based on left neighboring samples of the current video block (CCLM-L), a multi-model based CCLM (MM-CCLM), a multi-model based CCLM-T (MM-CCCM-T), or a multi-model based CCLM-L (MM-CCCM-L).

[0136] Item 3. A method according to Item 1 or 2, wherein at least one index for at least one row is included in the bitstream, the at least one index indicating at least one of the following: reconstructed chroma samples and corresponding luma samples for at least one row adjacent to the current video block, or reconstructed chroma samples and corresponding luma samples for at least one column adjacent to the current video block.

[0137] Item 4. The method of Item 3, wherein the corresponding luma samples are obtained by downsampling.

[0138] Item 5. The method of any one of Items 1 to 4, wherein the at least one row comprises a plurality of non-adjacent rows above the current video block.

[0139] Clause 6. The method of clause 5, wherein the number of the plurality of non-adjacent rows is predetermined.

[0140] Item 7. The method of Item 5 or 6, wherein the plurality of widths of the plurality of non-adjacent rows is the same as the width of the current video block.

[0141] Item 8. The method of Item 5 or 6, wherein the plurality of widths of the plurality of non-adjacent rows is greater than a width of the current video block, and the plurality of non-adjacent rows includes at least one sample above and to the left of the current video block.

[0142] Item 9. The method of any one of Items 1 to 8, wherein the at least one row comprises a plurality of non-adjacent columns to the left of the current video block.

[0143] Clause 10. The method of clause 9, wherein the number of the plurality of non-adjacent columns is predetermined.

[0144] Item 11. The method of Item 9 or 10, wherein the plurality of heights of the plurality of non-adjacent columns is the same as the height of the current video block.

[0145] Item 12. The method of Item 9 or 10, wherein a plurality of heights of the plurality of non-adjacent columns are greater than a height of the current video block, and the plurality of non-adjacent columns include at least one sample above and to the left of the current video block.

[0146] Item 13. A method according to any one of items 1 to 12, wherein the syntax elements in the bitstream are binarized into at least one of the following: a flag, a fixed-length code, an Exponential Golomb (x) (EG(x)) code, a unary code, a truncated unary code or a truncated binary code.

[0147] Clause 14. The method of clause 13, wherein the syntax elements are signed or unsigned.

[0148] Item 15. A method according to any one of Items 1 to 14, wherein syntax elements in the bitstream are encoded or decoded using at least one context model, or are bypass encoded or decoded.

[0149] Clause 16. A method according to any of clauses 13 to 15, wherein the syntax element is included in the bitstream based on a condition that a function associated with the syntax element is applicable.

[0150] Clause 17. The method of any one of clauses 13 to 15, wherein the syntax element is included in the bitstream if the dimensions of the current video block satisfy a condition.

[0151] Clause 18. The method of any one of clauses 13 to 17, wherein the syntax element is at at least one of: a block level, a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level.

[0152] Item 19. A method according to any one of Items 13 to 18, wherein the syntax element is in at least one of the following codec structures: a codec tree unit (CTU), a codec unit (CU), a transform unit (TU), a prediction unit (PU), a codec tree block (CTB), a codec block (CB), a transform block (TB), a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

[0153] Clause 20. A method according to any of clauses 1 to 19, wherein information on whether and / or how the method is applied is included in the bitstream.

[0154] Clause 21. The method of clause 20, wherein the information is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.

[0155] Item 22. A method according to Item 20 or 21, wherein the information is indicated in at least one of the following codec structures: a codec tree unit (CTU), a codec unit (CU), a transform unit (TU), a prediction unit (PU), a codec tree block (CTB), a codec block (CB), a transform block (TB), a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

[0156] Item 23. A method according to any one of Items 20 to 22, wherein the information is based on coded information, wherein the coded information includes at least one of the following: block size, color format, single tree partitioning or dual tree partitioning, color component, slice type, or picture type.

[0157] Item 24. A method according to any one of Items 1 to 23, wherein the method is used in a codec requiring chroma fusion.

[0158] Item 25. The method of any one of Items 1 to 24, wherein converting comprises encoding the current video block into a bitstream.

[0159] Item 26. The method of any one of Items 1 to 24, wherein converting comprises decoding the current video block from a bitstream.

[0160] Item 27. An apparatus for video processing, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 26.

[0161] Item 28. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 26.

[0162] Item 29. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a cross-component prediction (CCP) model for a current video block of the video based on at least one row of samples that are not adjacent to the current video block of the video, the at least one row comprising at least one of the following: rows of non-adjacent samples, or columns of non-adjacent samples; and generating a bitstream based on the CCP model.

[0163] Item 30. A method for storing a bitstream of a video, comprising: determining a cross-component prediction (CCP) model for a current video block of the video based on at least one row of samples that are non-adjacent to the current video block of the video, the at least one row comprising at least one of: rows of non-adjacent samples, or columns of non-adjacent samples; generating a bitstream based on the CCP model; and storing the bitstream in a non-transitory computer-readable recording medium. Example device

[0164] Figure 45 A block diagram of a computing device 4500 in which various embodiments of the present disclosure may be implemented is shown. The computing device 4500 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0165] It should be understood that Figure 45 The computing device 4500 shown in FIG. 4 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the disclosed embodiments.

[0166] like Figure 45 As shown, computing device 4500 comprises a general computing device 4500. Computing device 4500 may include at least one or more processors or processing units 4510, memory 4520, storage unit 4530, one or more communication units 4540, one or more input devices 4550, and one or more output devices 4560.

[0167] In some embodiments, computing device 4500 can be implemented as any user terminal or server terminal with computing capability. A server terminal can be a server, a large computing device, etc. provided by a service provider. A user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that computing device 4500 can support any type of interface to the user (such as a "wearable" circuit device, etc.).

[0168] The processing unit 4510 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 4520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 4500. The processing unit 4510 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0169] The computing device 4500 typically includes various computer storage media. Such media can be any media accessible by the computing device 4500, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 4520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 4530 can be any removable or non-removable medium and can include machine-readable media, such as memory, flash drive, disk or other media that can be used to store information and / or data and can be accessed in the computing device 4500.

[0170] The computing device 4500 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 45 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0171] The communication unit 4540 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 4500 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 4500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0172] Input device 4550 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. Output device 4560 may be one or more of various output devices, such as a display, speaker, printer, and the like. Computing device 4500 may also communicate with one or more external devices (not shown) via communication unit 4540, such as storage devices and display devices, one or more devices that enable a user to interact with computing device 4500, or any device that enables computing device 4500 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like), if desired. Such communication may occur via an input / output (I / O) interface (not shown).

[0173] In some embodiments, some or all components of the computing device 4500 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers in a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein may be provided by a conventional server or installed directly or otherwise on a client device.

[0174] In an embodiment of the present disclosure, the computing device 4500 may be used to implement video encoding / decoding. The memory 4520 may include one or more video encoding / decoding modules 4525 having one or more program instructions. These modules are accessible and executable by the processing unit 4510 to perform the functions of the various embodiments described herein.

[0175] In an example embodiment performing video encoding, an input device 4550 may receive video data as input to be encoded 4570. The video data may be processed, for example, by a video codec module 4525 to generate an encoded bitstream. The encoded bitstream may be provided as output 4580 via an output device 4560.

[0176] In an example embodiment performing video decoding, an input device 4550 may receive an encoded bitstream as input 4570. The encoded bitstream may be processed, for example, by a video codec module 4525 to generate decoded video data. The decoded video data may be provided as output 4580 via an output device 4560.

[0177] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for video processing, comprising: For conversion between a current video block of a video and a bitstream of the video, determining a cross-component prediction (CCP) model for the current video block based on at least one row of samples that are not adjacent to the current video block, the at least one row comprising at least one of the following: rows of non-adjacent samples, or columns of non-adjacent samples; and The conversion is performed based on the CCP model.

2. The method according to claim 1, wherein the CCP model comprises at least one of the following: Cross-Component Linear Model (CCLM), CCLM based on the top neighboring samples of the current video block (CCLM-T), Based on the CCLM of the left neighboring samples of the current video block (CCLM-L), Multi-model based CCLM (MM-CCLM), Multi-model based CCLM-T (MM-CCCM-T), or Multi-model based CCLM-L (MM-CCCM-L).

3. The method according to claim 1 or 2, wherein at least one index of the at least one row is included in the bitstream, the at least one index indicating at least one of the following: reconstructed chroma samples and corresponding luma samples of at least one row adjacent to the current video block, or At least one column of reconstructed chroma samples and corresponding luma samples adjacent to the current video block. The method according to claim 3 , wherein the corresponding luma samples are obtained by downsampling.

5. The method of any one of claims 1 to 4, wherein the at least one row comprises a plurality of non-adjacent rows above the current video block. The method of claim 5 , wherein the number of the plurality of non-adjacent rows is predetermined.

7. The method of claim 5 or 6, wherein the plurality of widths of the plurality of non-adjacent rows is the same as a width of the current video block.

8. The method of claim 5 or 6, wherein widths of the plurality of non-adjacent rows are greater than a width of the current video block, and the plurality of non-adjacent rows include at least one sample above and to the left of the current video block.

9. The method of any one of claims 1 to 8, wherein the at least one row comprises a plurality of non-adjacent columns to the left of the current video block.

10. The method of claim 9, wherein the number of the plurality of non-adjacent columns is predetermined.

11. The method of claim 9 or 10, wherein a plurality of heights of the plurality of non-adjacent columns is the same as a height of the current video block.

12. The method according to claim 9 or 10, wherein heights of the plurality of non-adjacent columns are greater than a height of the current video block, and the plurality of non-adjacent columns include at least one sample above and to the left of the current video block.

13. The method according to any one of claims 1 to 12, wherein the syntax elements in the bitstream are binarized into at least one of the following: a flag, a fixed-length code, an Exponential Golomb (x) (EG(x)) code, a unary code, a truncated unary code or a truncated binary code. The method of claim 13 , wherein the syntax elements are signed or unsigned.

15. The method according to any one of claims 1 to 14, wherein syntax elements in the bitstream are coded using at least one context model, or are bypass coded.

16. The method according to any one of claims 13 to 15, wherein the syntax element is included in the bitstream based on a condition that a function associated with the syntax element is applicable.

17. The method of any one of claims 13 to 15, wherein the syntax element is included in the bitstream if a dimension of the current video block satisfies a condition.

18. The method of any one of claims 13 to 17, wherein the syntax element is at at least one of: block level, sequence level, group of pictures level, picture level, slice level, or slice group level.

19. The method of any one of claims 13 to 18, wherein the syntax element is in at least one of the following codec structures: a codec tree unit (CTU), a codec unit (CU), a transform unit (TU), a prediction unit (PU), a codec tree block (CTB), a codec block (CB), a transform block (TB), a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

20. The method according to any one of claims 1 to 19, wherein information on whether and / or how to apply the method is included in the bitstream.

21. The method of claim 20, wherein the information is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.

22. The method of claim 20 or 21, wherein the information is indicated in at least one of the following codec structures: a codec tree unit (CTU), a codec unit (CU), a transform unit (TU), a prediction unit (PU), a codec tree block (CTB), a codec block (CB), a transform block (TB), a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

23. The method according to any one of claims 20 to 22, wherein the information is based on coded information, wherein the coded information includes at least one of the following: block size, color format, single tree partitioning or dual tree partitioning, color component, slice type, or picture type.

24. The method according to any one of claims 1 to 23, wherein the method is used in a codec requiring chroma fusion.

25. The method of any one of claims 1 to 24, wherein the converting comprises encoding the current video block into the bitstream.

26. The method of any one of claims 1 to 24, wherein the converting comprises decoding the current video block from the bitstream.

27. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 26.

28. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to execute the method according to any one of claims 1 to 26.

29. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: Determining a cross-component prediction (CCP) model for the current video block based on at least one row of samples that are not adjacent to the current video block of the video, the at least one row comprising at least one of: a row of non-adjacent samples, or a column of non-adjacent samples; and The bitstream is generated based on the CCP model.

30. A method for storing a bitstream of a video, comprising: Determining a cross-component prediction (CCP) model for the current video block based on at least one row of samples that are not adjacent to the current video block of the video, the at least one row comprising at least one of the following: a row of non-adjacent samples, or a column of non-adjacent samples; generating the bitstream based on the CCP model; as well as The bitstream is stored in a non-transitory computer-readable recording medium.