Method and device for video processing and medium

Through cross-component prediction technology, the prediction value of the video unit is generated and reconstructed, and the problems of encoding and codec efficiency and performance improvement in the existing video encoding and codec technology are solved, and more efficient video encoding and codec are achieved.

CN120457689APending Publication Date: 2025-08-08DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480006561.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-27
Filing Date
2024-01-02
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

There is room for improvement in the encoding and decoding efficiency and performance of existing video encoding and decoding technologies, especially in cross-component prediction.

Method used

The predicted values of the video unit are generated by cross-component prediction candidates, modified the predicted values, obtained the reconstructed sample point values, and performed conversion based on the reconstructed sample point values to improve encoding and decoding efficiency and performance.

Benefits of technology

It improves the encoding and decoding efficiency and performance of video encoding and decoding, is suitable for existing video encoding and decoding standards such as HEVC and VVC, and supports future standard video codecs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120457689A_ABST
    Figure CN120457689A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises: for a conversion between a video unit of a video and a bitstream of the video unit, generating a prediction value of the video unit based on a cross-component prediction candidate; modifying the predicted value of the video unit; obtaining a reconstructed sample value based on the modified predicted value; and performing a conversion based on the reconstructed sample value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to non-adjacent cross-component prediction. Background Art

[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is provided. The method includes: for conversion between a video unit and a bitstream of the video unit, generating a prediction value for the video unit based on a cross-component prediction candidate; modifying the prediction value for the video unit; obtaining reconstructed sample values based on the modified prediction value; and performing conversion based on the reconstructed sample values. This can improve codec efficiency and codec performance.

[0005] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.

[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, which enable a processor to execute the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: generating a prediction value for a video unit of the video based on a cross-component prediction candidate; modifying the prediction value for the video unit; obtaining a reconstructed sample value based on the modified prediction value; and generating a bitstream based on the reconstructed sample value.

[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: generating a prediction value of a video unit of the video based on a cross-component prediction candidate; modifying the prediction value of the video unit; obtaining a reconstructed sample value based on the modified prediction value; generating a bitstream based on the reconstructed sample value; and storing the bitstream in a non-transitory computer-readable medium.

[0009] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other objects, features and advantages of example embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings, in which like reference numerals generally refer to like components throughout the example embodiments of the present disclosure.

[0011] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;

[0012] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;

[0013] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;

[0014] Figure 4 The nominal vertical and horizontal positions of the 4:2:2 luminance and chrominance samples in the picture are shown;

[0015] Figure 5 An example of an encoder block diagram is shown;

[0016] Figure 6 67 intra prediction modes are shown;

[0017] Figure 7 Reference samples for wide-angle intra prediction are shown;

[0018] Figure 8 The problem of discontinuity is shown in the case of orientations exceeding 45°;

[0019] Figure 9 The positions of the sample points used to derive α and β are shown;

[0020] Figure 10 An example of classifying neighboring points into two groups is shown;

[0021] Figure 11A is a diagram showing the definition of sample points used by PDPC applied to a diagonal upper right mode;

[0022] Figure 11B is a diagram showing the definition of sample points used by PDPC applied to a diagonal bottom-left mode;

[0023] Figure 11C is a diagram showing the definition of sample points used by PDPC applied to an adjacent diagonal upper right pattern;

[0024] Figure 11D is a diagram showing the definition of sample points used by PDPC applied to the adjacent diagonal lower left pattern;

[0025] Figure 12 is a schematic diagram illustrating a gradient scheme for non-vertical / non-horizontal modes;

[0026] Figure 13 is a diagram showing nScale values relative to nTbH and mode number; for all cases where nScale<0, a gradient scheme is used;

[0027] Figure 14 is a schematic diagram showing a flow chart of the current PDPC and the proposed PDPC;

[0028] Figure 15 is a schematic diagram showing neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list;

[0029] Figure 16 is a schematic diagram showing an example of proposed intra reference mapping;

[0030] Figure 17 is a diagram showing an example of four reference rows adjacent to a prediction block;

[0031] Figure 18A is a diagram showing an example of sub-partitioning for 4×8 CU and 8×4 CU;

[0032] Figure 18B is a diagram showing an example of sub-partitioning for CUs other than 4×8, 8×4, and 4×4;

[0033] Figure 19 is a schematic diagram illustrating a matrix-weighted intra prediction process;

[0034] Figure 20 is a schematic diagram showing target points, template points, and reference points of the template used in DIMD;

[0035] Figure 21is a schematic diagram illustrating the proposed intra-block decoding process;

[0036] Figure 22 is a schematic diagram showing the calculation of HoG from a template with a width of 3 pixels;

[0037] Figure 23 is a schematic diagram illustrating prediction fusion by weighted averaging of two HoG modes and a plane;

[0038] Figure 24 is a schematic diagram showing the spatial portion of a convolutional filter;

[0039] Figure 25 is a schematic diagram showing a reference area (and its filling) for deriving filter coefficients;

[0040] Figure 26 is a schematic diagram showing four Sobel-based gradient modes for GLM;

[0041] Figure 27 is a schematic diagram showing spatial sample points for GL-CCCM;

[0042] Figure 28 is a schematic diagram illustrating non-downsampled luminance samples;

[0043] Figure 29 The spatial GPM candidates are shown;

[0044] Figure 30 A GPM template is shown;

[0045] Figure 31 GPM mixing is shown;

[0046] Figure 32 The figure shows the binarization of the cross-component prediction mode in ECM. "CCLM" in the figure can be replaced by "CCCM";

[0047] Figure 33 An example of luminance samples to be prepared is shown;

[0048] Figure 34 Examples of potential candidate regions (shaded blocks) are shown;

[0049] Figure 35 Possible templates are shown;

[0050] Figure 36 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and

[0051] Figure 37 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.

[0052] Throughout the drawings, same or similar reference numbers generally refer to same or similar elements. DETAILED DESCRIPTION

[0053] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.

[0054] In the following description and claims, unless defined otherwise, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0055] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.

[0056] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0057] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms "comprise," "including," "having," "include," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment

[0058] Figure 1is a block diagram illustrating an example video codec system 100 that can utilize the techniques of this disclosure. As shown, video codec system 100 may include a source device 110 and a destination device 120. Source device 110 may also be referred to as a video encoding device, and destination device 120 may also be referred to as a video decoding device. In operation, source device 110 may be configured to generate encoded video data, and destination device 120 may be configured to decode the encoded video data generated by source device 110. Source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0059] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0060] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a coded picture and associated data. The coded picture is a coded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be sent directly to the target device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the target device 120.

[0061] Target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120, or may be external to target device 120, which is configured to interface with an external display device.

[0062] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.

[0063] Figure 2is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.

[0064] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0065] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.

[0066] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in accordance with an IBC mode, wherein at least one reference picture is a picture in which the current video block is located.

[0067] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.

[0068] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0069] The mode selection unit 203 can, for example, select one of a plurality of codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the generated intra-frame codec block or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).

[0070] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.

[0071] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture comprised of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture comprised of macroblocks that are independent of macroblocks in the same picture.

[0072] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to search for a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0073] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate reference indexes indicating the reference pictures in list 0 and list 1 that contain the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0074] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.

[0075] In one example, motion estimation unit 204 may indicate to video decoder 300 a value in a syntax structure associated with the current video block that indicates the current video block has the same motion information as another video block.

[0076] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0077] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0078] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0079] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0080] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0081] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.

[0082] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0083] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0084] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0085] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0086] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.

[0087] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0088] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.

[0089] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-encoded video data, which includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.

[0090] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.

[0091] The motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 based on received syntax information, and the motion compensation unit 302 may use the interpolation filters to generate a prediction block.

[0092] The motion compensation unit 302 may use at least a portion of the syntax information to determine the block sizes used to encode the frame(s) and / or slice(s) of the coded video sequence, partition information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the coded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice may be an entire picture or a region of a picture.

[0093] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0094] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be used to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.

[0095] Some exemplary embodiments of the present disclosure are described in detail below. It should be understood that the section titles used in this document are for ease of understanding and do not limit the embodiments disclosed in the section to only that section. In addition, although certain embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video codec technologies. In addition, although some embodiments describe the video encoding and decoding steps in detail, it will be understood that the corresponding decoding steps of the de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview The present disclosure relates to video coding techniques. Specifically, it relates to cross-component prediction. It can be applied to existing video coding standards such as HEVC or Versatile Video Codec (VVC). It can also be applied to future video coding standards or video codecs. 2. Introduction Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 Visual and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video standard, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standard. As a result of H.262, video codec standards are based on a hybrid video codec architecture that utilizes temporal prediction and transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC. 2.1. Color Space and Chroma Downsampling A color space, also called a color model (or color system), is an abstract mathematical model that simply describes the range of colors as a tuple of numbers, usually 3 or 4 values or color components (e.g., RGB). Basically, a color space is a refinement of a coordinate system and subspace. For video compression, the most commonly used color spaces are YCbCr and RGB. YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luma component, and CB and CR are the blue-difference and red-difference chroma components. Y' (with a prime) is distinguished from Y, which is luma, meaning that light intensity is encoded nonlinearly based on the gamma-corrected RGB primaries. Chroma downsampling is the practice of encoding an image at a lower resolution for chroma information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to differences in color than to luminance. 2.1.1. 4:4:4 Each of the three Y'CbCr components has the same sampling rate, so there is no chroma downsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2. 4:2:2 The two chroma components are sampled at half the sample rate of luma: the horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third with little or no visual difference. Examples of nominal vertical and horizontal positions for a 4:2:2 color format are given in the VVC working draft. Figure 4 Depicted in. 2.1.3. 4:2:0 In 4:2:0, horizontal sampling is doubled compared to 4:1:1, but vertical resolution is halved because the Cb and Cr channels are sampled only on alternate lines. Therefore, the data rate remains the same. Cb and Cr are downsampled by a factor of 2 both horizontally and vertically. There are three variants of the 4:2:0 scheme, with different horizontal and vertical positions. In MPEG-2, Cb and Cr are co-located in the horizontal direction. Cb and Cr are located between pixels in the vertical direction (at interstitial positions). In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located in interstitial positions, in the middle of alternating luma samples. In 4:2:0 DV, Cb and Cr are co-located horizontally. Vertically, they are co-located on alternate lines. Table 2-1 SubWidthC and SubHeightC values derived from chroma_format_idc and separate_colour_plane_flag chroma_format_idc separate_colour_plane_flag Chroma format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1 2.2. Encoding and decoding process of typical video codecs Figure 5 An example of a VVC encoder block diagram is shown, which contains three loop filtering blocks: deblocking filter (DF), sample adaptive offset (SAO), and ALF. Unlike DF, which uses a predefined filter, SAO and ALF use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offset and applying a finite impulse response (FIR) filter, respectively, and using the encoded side information to signal the offset and filter coefficients. ALF is located in the last processing stage of each picture and can be seen as a tool that attempts to collect and repair artifacts caused by previous stages. Intra-mode codec with 67 intra-prediction modes In order to capture arbitrary edge directions present in natural videos, such as Figure 6As shown, the number of directional intra modes is extended from 33 used in HEVC to 65, while planar and DC modes remain unchanged. These more dense directional intra prediction modes are applicable to all block sizes and both luma and chroma intra prediction. In HEVC, each intra-coded block has a square shape, and the length of each side is a power of 2. Therefore, no division operation is required to generate intra prediction values using DC mode. In VVC, blocks can have a rectangular shape, which generally requires a division operation for each block. To avoid division operations for DC prediction, only the longer side is used to calculate the average value of non-square blocks. 2.3.1. Wide-angle intra prediction Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index also depends on the block shape. Conventional angular intra prediction directions are defined as going from 45 degrees to -135 degrees clockwise. In VVC, for non-square blocks, multiple conventional angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced mode is signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e. 67, and the intra mode encoding and decoding method remains unchanged. To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as Figure 7 shown. The number of modes replaced in the wide direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 2-2. Table 2-2 Intra-frame prediction modes replaced by wide-angle mode like Figure 8 As shown in Figure 2, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the increased gap Δp. α negative impact. If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that meet this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted using these modes, the samples in the reference buffer are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of the non-fractional mode in the conventional prediction mode with the wide-angle mode. In VVC, in addition to 4:2:0, the 4:2:2 chroma format and the 4:4:4 chroma format are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, expanding the number of entries from 35 to 67 to align with the expansion of intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values of the entries of the mapping table to more accurately convert the prediction angles for chroma blocks. 2.4. Intra-frame prediction mode encoding and decoding for chroma components For the chroma component of an intra PU, the encoder selects the best chroma prediction mode from five modes, including planar, DC, horizontal, vertical, and direct copy of the intra prediction modes for the luma component. The mapping between the intra prediction direction of chroma and the number of intra prediction modes is shown in Table 2-3. When the number of intra prediction modes for the chroma component is 4, the intra prediction direction of the luma component is used for generating intra prediction samples for the chroma component. When the number of intra prediction modes for the chroma component is not 4 and is the same as the number of intra prediction modes for the luma component, the intra prediction direction of 66 is used for generating intra prediction samples for the chroma component. 2.5. Inter-frame prediction For each inter-frame predicted CU, the motion parameters consist of a motion vector, a reference picture index and a reference picture list usage index, and the new coding features for VVC require additional information for inter-frame prediction sample generation. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU is associated with one PU and has no significant residual coefficients, coded motion vector increments or reference picture indices. A Merge mode is specified, in which the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates and additional lists introduced in VVC. Merge mode can be applied to any inter-frame predicted CU, not just for skip mode. An alternative to Merge mode is the explicit transmission of motion parameters, in which the motion vector, corresponding reference picture index and reference picture list usage flag for each reference picture list, as well as other required information, are explicitly signaled for each CU. 2.6. Intra-block copy (IBC) Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed inside the current picture. The luminance block vector of the CU encoded and decoded by IBC is in integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. The CU encoded and decoded by IBC is regarded as a third prediction mode different from the intra or inter prediction mode. The IBC mode is applicable to CUs with a width and height less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height no greater than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed. In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on a 4×4 sub-block. For larger-sized current blocks, a hash key is determined to match the hash key of a reference block when all hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the smallest cost is selected. In the block matching search, the search range is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using a flag, and it can be signaled as IBC AMVP mode or IBC Skip / Merge mode as shown below: -IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and pairwise candidates. - IBC AMVP mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the top neighbor (if encoded with IBC). When either neighbor is unavailable, the default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index. 2.7. Cross-component linear model prediction To reduce cross-component redundancy, a Cross-Component Linear Model (CCLM) prediction mode is used in VVC, for which chroma samples are predicted based on the reconstructed luma samples of the same CU by using the following linear model: where pred C (i, j) represents the predicted chroma sample in CU, and rec L (i, j) represents the downsampled reconstructed luma samples of the same CU. The CCLM parameters (α and β) are derived using up to four adjacent chroma samples and their corresponding downsampled luma samples. Assuming the current chroma block size is W×H, W' and H' are set to - When LM mode is applied, W'=W, H'=H; - When LM_T mode is applied, W'=W+H; - When LM_L mode is applied, H'=H+W. The upper adjacent position is represented as S[0, -1] ... S[W'-1, -1], and the left adjacent position is represented as S[-1, 0] ... S[-1, H'-1]. Then the four sample points are selected as - When LM mode is applied and both the upper and left neighboring samples are available, S[W' / 4, -1], S[3*W' / 4, -1], S[-1, H' / 4], S[-1, 3*H' / 4]; - When LM_T mode is applied or only upper neighboring samples are available, S[W' / 8, -1], S[3*W' / 8, -1], S[5*W' / 8, -1], S[7*W' / 8, -1]; - When LM_L mode is applied or only left neighbor samples are available, S[-1, H' / 8], S[-1, 3*H' / 8], S[-1, 5*H' / 8], S[-1, 7*H' / 8]. The four adjacent brightness samples at the selected position are downsampled and compared four times to find the two larger values: x 0 A and x 1 A , and two smaller values: x0 B and x 1 B Their corresponding chroma sample values are represented by y 0 A 、y 1 A 、y 0 B and y 1 B Then x A 、x B 、y A and y B is derived as: X a =(x 0 A +x 1 A +1)>>1;X b =(x 0 B +x 1 B +1)>>1;Y a =(y 0 A +y 1 A +1)>>1;Y b =(y 0 B +y 1 B +1)>>1. (2-2) Finally, the linear model parameters α and β are obtained according to the following equations. β=Y b -α·X b (2-4) Figure 9 An example of the positions of left and upper samples involved in the CCLM mode and samples of the current block is shown. The division operation for calculating the parameter α is implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are expressed using exponential notation. For example, diff is approximated using a 4-bit significant part and an exponent. Thus, for 16 values of significant digits, the table of 1 / diff is reduced to 16 elements, as shown below: DivTable[]={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}(2-5) This will have the advantage of reducing the computational complexity as well as the memory size required to store the required tables. In addition to the upper template and the left template being used together to calculate the linear model coefficients, they can also be used alternately in the other two LM modes (called LM_T mode and LM_L mode). In LM_T mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is expanded to (W+H) samples. In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is expanded to (H+W) samples. In LM mode, the left template and the upper template are used to calculate the linear model coefficients. To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters are applied to the luma samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The choice of downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "Type 0" and "Type 2" content, respectively. Note that when the upper reference line is at a CTU boundary, only one luma line (common line buffer in intra prediction) is used to make the downsampled luma samples. This parameter calculation is performed as part of the decoding process, not just as an encoder search operation. Consequently, syntax is not used to convey the α and β values to the decoder. For chroma intra mode encoding and decoding, a total of 8 intra modes are allowed for chroma intra mode encoding and decoding. These modes include five regular intra modes and three cross-component linear model modes (LM, LM_T and LM_L). The chroma modes are transmitted through signals and derivation processes as shown in Table 2-3. The chroma mode encoding and decoding depends directly on the intra prediction mode of the corresponding luminance block. Since separate block partitioning structures for luminance components and chrominance components are enabled in the I stripe, one chroma block can correspond to multiple luminance blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chroma block is directly inherited. Table 2-3 Derivation of chroma prediction mode from luma mode when CCLM is enabled like Figure 2-4 As shown, a single binarization table is used regardless of the value of sps_cclm_enabled_flag. Table 2-4 Unified binarization table for chroma prediction mode In Table 2-4, the first binary bit indicates whether it is normal (0) mode or LM mode (1). If it is LM mode, the next binary bit indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next 1 binary bit indicates whether it is LM_L (0) or LM_T (1). For this case, when sps_cclm_enabled_flag is 0, the first binary bit of the binarization table corresponding to intra_chroma_pred_mode can be discarded before entropy coding. Or, in other words, the first binary bit is inferred to be 0 and therefore not coded. This single binarization table is used for both cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two binary bits in Table 2-4 are context coded using their own context model, and the remaining binary bits are bypass coded. Additionally, to reduce luma-chroma latency in dual trees, when a 64x64 luma codec tree node is split using no split (and ISP is not used for 64x64 CUs) or QT, the chroma CUs in the 32x32 / 32x16 chroma codec tree nodes are allowed to use CCLM in the following manner: – If a 32×32 chroma node is not partitioned or is split by a QT partition, all chroma CUs in the 32×32 node can use CCLM; If a 32×32 chroma node is partitioned using horizontal BT, and the 32×16 child nodes are not partitioned or use vertical BT partitioning, all chroma CUs in the 32×16 chroma node can use CCLM. Under all other luma and chroma codec tree partitioning conditions, CCLM is not allowed for chroma CUs. 2.8. Multi-model Linear Model (MMLM) With MMLM, there can be more than one linear model between the luma samples and chroma samples in a CU. In this method, the neighboring luma samples and chroma samples of the current block are classified into several groups, each of which is used as a training set to derive a linear model (i.e., a specific α and β are derived for a specific group). In addition, the samples of the current luma block are also classified based on the same rules as the classification of the neighboring luma samples. Neighboring samples can be classified into M groups, where M is 2 or 3. In addition to the original LM mode, the MMLM method with M = 2 and M = 3 is designed as two additional chroma prediction modes, called MMLM2 and MMLM3. The encoder selects the best mode during the RDO process and transmits it through the signal. When M is equal to 2, Figure 10An example of classifying neighboring samples into two groups is shown. The threshold is calculated as the average value of the neighboring reconstructed luminance samples. Rec'L[x,y]<=threshold Rec' L Neighboring points with [x, y]≤Threshold are classified into group 1; and Rec′ L Neighboring points where [x,y]>ThresholdRec'L[x,y]>threshold are classified into group 2. Similar to CCLM, there are three modes in MMLM, namely MMLM, MMLM_T, and MMLM_L. The two models are derived as follows. The threshold is the average of neighboring samples of the luminance reconstruction. If enabled, a linear model for each class is derived using the least mean square (LMS) method, or using the min / max method of VVC. 2.9. Position-dependent intra prediction combination In VVC, the results of intra prediction for DC, planar, and multiple angle modes are further modified by the Position Dependent Intra Prediction Combination (PDPC) method. PDPC is an intra prediction method that calls for a combination of boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, intra angles less than or equal to horizontal, and intra angles greater than or equal to vertical and less than or equal to 80. PDPC is not applied if the current block is in BDPCM mode or the MRL index is greater than 0. The prediction sample pred(x', y') is predicted using the intra prediction mode (DC, planar, angular) and a linear combination of the reference samples according to the following equation 2-8: pred(x', y') = limit(0, (1 < <BitDepth)–1,(wL×R -1,y’ +wT×R x’,-1 +(64-wL-wT)×pred(x',y')+32)>>6) (2-9) where R x,-1 、R -1,y Respectively represent the reference sample points located at the upper and left boundaries of the current sample point (x, y). If PDPC is applied to DC, planar, horizontal and vertical intra modes, no additional boundary filters are required, as required in the case of HEVC DC mode boundary filters or horizontal / vertical mode edge filters. The PDPC process is identical for DC mode and planar mode. For angular mode, if the current angular mode is HOR_IDX or VER_IDX, the left reference sample or the top reference sample, respectively, is not used. The PDPC weights and scaling factors depend on the prediction mode and the size of the block. PDPC is applied to blocks with width and height both greater than or equal to 4. 11A to 11D The reference samples (R x,-1 and R -1,y ) definition. The prediction sample point pred(x', y') is located at (x', y') in the prediction block. As an example, the reference sample point R x,-1 The coordinate x of is given by the following formula: x=x'+y'+1, and the reference point R -1,y The coordinate y of is similarly given by the following formula: y = x' + y' + 1 (for diagonal mode). For other angle modes, the reference point R x,-1 and R -1,y Can be located at fractional sample positions. In this case, the sample value at the nearest integer sample position is used. Gradient PDPC like Figure 12 As shown, the gradient-based scheme is extended for non-vertical / non-horizontal modes. Here, the gradient is calculated as r(-1,y)–r(-1+d,-1), where d is the horizontal displacement depending on the angle direction. There are a few points to note here: The gradient term r(-1,y)–r(-1+d,-1) needs to be calculated once for each row since it does not depend on the x position. The calculation of d is already part of the original intra prediction process and can be reused, so there is no need to calculate d separately. Therefore, d has 1 / 32 pixel accuracy. When d is in fractional positions, we use two-tap (linear) filtering, i.e., if dPos is the displacement with 1 / 32 pixel accuracy, dInt is the (rounded-down) integer part (dPos>>5), and dFract is the fractional part with 1 / 32 pixel accuracy (dPos>31), then r(-1+d) is calculated as: r(-1+d)=(32–dFrac)*r(-1+dInt)+dFrac*r(-1+dInt+1). As explained in a, this two-tap filtering is performed once for each row (if necessary). Finally, the prediction signal is calculated as follows: p(x,y)=clipping(((64–wL(x))*p(x,y)+wL(x)*(r(-1,y)-r(-1+d,-1))+32)>>6) Where wL(x)=32>>((x<<1)>>nScale2), and nScale2=(log2(nTbH)+log2(nTbW)–2)>>2, which is the same as the vertical / horizontal mode. In short, the same process is applied compared to the vertical / horizontal mode (in fact, d=0 indicates the vertical / horizontal mode). Second, when (nScale < 0) or when PDPC cannot be applied due to the unavailability of secondary reference samples, we activate the gradient-based scheme for non-vertical / non-horizontal modes. Figure 13 The nScale values with respect to the TB size and angle pattern are shown in , to better visualize the use of the gradient scheme. Figure 14 In , we have shown the flow charts for the current PDPC and the proposed PDPC. 2.11.Secondary MPM The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list includes 16 entries. First, a general MPM list with 22 entries is constructed, and then the first 6 entries in the general MPM list are included in the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the plane mode. The remaining entries are as follows: Figure 15 The shown consists of intra modes for the left (L), above (A), bottom left (BL), upper right (AR), and upper left (AL) neighboring blocks, directional modes with offsets added from the first two available directional modes for the neighboring blocks, and a default mode. If the CU block is vertically oriented, the order of neighboring blocks is A, L, BL, AR, AL; otherwise, it is L, A, BL, AR, AL. The PMPM flag is parsed first, if equal to 1, the PMPM index is parsed to determine which entry of the PMPM list is selected, otherwise the SPMPM flag is parsed to determine whether to parse the SMPM index or the remaining mode. 2.12.6 Tap Intra-frame Interpolation Filter In order to improve the prediction accuracy, it is proposed to replace the 4-tap cubic interpolation filter with a 6-tap interpolation filter. The filter coefficients are derived based on the same polynomial regression model, but the polynomial order is 6. The filter coefficients are listed as follows, {0,0,256,0,0,0}, / / 0 / 32 position {0,-4,253,9,-2,0}, / / 1 / 32 position {1,-7,249,17,-4,0}, / / 2 / 32 position {1,-10,245,25,-6,1}, / / 3 / 32 position {1,-13,241,34,-8,1}, / / 4 / 32 position {2,-16,235,44,-10,1}, / / 5 / 32 position {2,-18,229,53,-12,2}, / / 6 / 32 position {2,-20,223,63,-14,2}, / / 7 / 32 position {2,-22,217,72,-15,2}, / / 8 / 32 position {3,-23,209,82,-17,2}, / / 9 / 32 position {3,-24,202,92,-19,2}, / / 10 / 32 position {3,-25,194,101,-20,3}, / / 11 / 32 position {3,-25,185,111,-21,3}, / / 12 / 32 position {3,-26,178,121,-23,3}, / / 13 / 32 position {3,-25,168,131,-24,3}, / / 14 / 32 position {3,-25,159,141,-25,3}, / / 15 / 32 position {3,-25,150,150,-25,3}, / / half-pixel position. The reference samples used for interpolation come from reconstructed samples or padding samples as in HEVC, so no conditional check on the availability of reference samples is required. It is proposed to use a 4-tap cubic interpolation filter instead of using the nearest integer operation to derive the extended intra-frame reference samples. Figure 16 As shown in the example in , a four-tap interpolation filter is used to derive the value of the reference sample point P, while in JEM-3.0 or HM, P is directly set to X1. 2.13. Multiple Reference Line (MRL) Intra Prediction Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. Figure 17In

[15] , an example of 4 reference lines is depicted, where the samples of segments A and F are not obtained from reconstructed neighboring samples, but are filled with the closest samples from segments B and E, respectively. HEVC intra picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 2) are used. The index of the selected reference row (mrl_idx) is signaled and used to generate intra prediction values. For reference row indices greater than 0, only additional reference row modes are included in the MPM list, and only the MPM index is signaled without signaling the remaining modes. The reference row index is signaled before the intra prediction mode, and if a non-zero reference row index is signaled, planar mode is excluded from the intra prediction mode. MRL is disabled for the first row of blocks within a CTU to prevent the use of extended reference samples outside the current CTU row. In addition, PDPC is disabled when additional rows are used. For MRL mode, the derivation of the DC value in DC intra prediction mode for non-zero reference row index is consistent with the derivation of reference row index 0. MRL requires the storage of 3 neighboring luma reference rows with the CTU to generate the prediction. The Cross Component Linear Model (CCLM) tool also requires 3 neighboring luma reference rows for its downsampling filter. The definition of MRL using the same 3 rows is aligned with CCLM to reduce storage requirements for the decoder. 2.14. Intra-frame sub-segmentation (ISP) Intra sub-partitioning (ISP) divides the luma intra prediction block into 2 or 4 sub-partitions vertically or horizontally depending on the block size. For example, the minimum block size for ISP is 4×8 (or 8×4). If the block size is larger than 4×8 (or 8×4), the corresponding block is divided into 4 sub-partitions. It has been noted that M×128 (where M≤64) and 128×N (where N≤64) ISP blocks may cause potential problems for 64×64 VDPU. For example, an M×128 CU in the single-tree case has an M×128 luma TB and two corresponding Chroma TB. If the CU uses ISP, the luma TB will be divided into four M×32TBs (only horizontal division is possible), each TB is smaller than a 64×64 block. However, in the current ISP design, the chroma blocks are not divided. Therefore, both chroma components will have a size larger than a 32×32 block. Similarly, a 128×NCU using ISP can cause a similar situation. Therefore, these two situations are problems for a 64×64 decoder pipeline. For this reason, the CU size that can use ISP is limited to a maximum of 64×64. 18A to 18B Examples of two possibilities are shown. All sub-segments satisfy the condition of having at least 16 samples. In ISP, dependencies of 1×N / 2×N sub-block predictions on reconstructed values of previously decoded 1×N / 2×N sub-blocks of the codec block are not allowed, resulting in a minimum prediction width of four samples for the sub-block. For example, an 8×N (N>4) codec block coded using ISP with vertical partitioning is partitioned into two prediction regions of size 4×N each and four transforms of size 2×N. Furthermore, a 4×N codec block coded using ISP with vertical partitioning is predicted using a full 4×N block; four transforms of size 1×N are used. Although transform sizes of 1×N and 2×N are allowed, it is asserted that transforms for these blocks in a 4×N region can be performed in parallel. For example, when a 4×N prediction region contains four 1×N transforms, there is no transform in the horizontal direction; the transform in the vertical direction can be performed as a single 4×N transform in the vertical direction. Similarly, when a 4×N prediction region contains two 2×N transform blocks, transform operations for the two 2×N blocks in each direction (horizontally and vertically) can be performed in parallel. Therefore, there is no added delay in processing these smaller blocks compared to processing intra blocks for 4x4 regular codecs. Table 2-5 Entropy encoding and decoding coefficient group size Block size Coefficient group size 1×N,N≥16 1×16 N×1,N≥16 16×1 2×N,N≥8 2×8 N×2,N≥8 8×2 All other possible M×N situations 4×4 For each sub-partition, the reconstructed samples are obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated by processes such as entropy decoding, inverse quantization and inverse transformation. Therefore, the reconstructed sample values of each sub-partition can be used to generate the prediction of the next sub-partition, and each sub-partition is processed repeatedly. In addition, the first sub-partition to be processed is the sub-partition containing the upper left sample of the CU, and then continues downward (horizontal partitioning) or to the right (vertical partitioning). As a result, the reference samples used to generate the sub-partition prediction signal are only located to the left and above the row. All sub-partitions share the same intra mode. The following is a summary of the interaction of ISP with other codec tools. – Multiple Reference Line (MRL): If a block has an MRL index other than 0, the ISP codec mode will be inferred to be 0, so the ISP mode information will not be sent to the decoder. – Entropy coding coefficient group size: The size of the entropy coding sub-blocks has been modified so that they have 16 samples in all possible cases, as shown in Table 2-5. Note that the new size only affects blocks generated by ISP where one dimension is less than 4 samples. In all other cases, the coefficient group remains 4×4 in size. –CBF codec: It is assumed that at least one subpartition has a non-zero CBF. Thus, if n is the number of subpartitions and the first n-1 subpartitions produce zero CBF, the CBF of the nth subpartition is assumed to be 1. – Transform size restriction: All ISP transforms with length greater than 16 points use DCT-II. -MTS flag: If the CU uses ISP codec mode, the MTS CU flag will be set to 0 and will not be sent to the decoder. Therefore, the encoder will not perform RD tests for the different available transforms for each resulting sub-partition. Instead, the transform selection for ISP mode will be fixed and selected based on the utilized intra mode, processing order and block size. Therefore, no signaling is required. For example, making t H and t V are the horizontal and vertical transforms selected for the w×h sub-partition, respectively, where w is the width and h is the height. The transforms are then selected according to the following rules: If w=1 or h=1, there is no horizontal transform or vertical transform, respectively. – If w ≥ 4 and w ≤ 16, t H =DST-VII, otherwise, t H =DCT-II. – If h ≥ 4 and h ≤ 16, t V =DST-VII, otherwise, t V =DCT-II. In ISP mode, all 67 intra prediction modes are allowed. PDPC is also applied if the corresponding width and height are at least 4 samples long. In addition, the conditions for the reference sample filtering process (reference smoothing) and intra interpolation filter selection no longer exist, and the cubic (DCT-IF) filter is always applied to fractional position interpolation in ISP mode. 2.15. Matrix Weighted Intra Prediction (MIP) The matrix weighted intra prediction (MIP) method is a new intra prediction technique in VVC. In order to predict the samples of a rectangular block of width W and height H, the matrix weighted intra prediction (MIP) takes a row of H reconstructed adjacent boundary samples on the left side of the block and a row of W reconstructed adjacent boundary samples above the block as input. If the reconstructed samples are not available, they are generated in the same way as in conventional intra prediction. The generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication and linear interpolation, as shown in Figure 19 shown. 2.15.1. Averaging neighboring points Among the boundary samples, four or eight samples are selected by averaging based on the block size and shape. Specifically, the boundary bdry is input by averaging the adjacent boundary samples according to a predefined rule depending on the block size. top and bdry left Reduced to a smaller boundary and Then, the two reduced bounds and is spliced to the reduced boundary vector bdry red , so for blocks of shape 4×4 its size is 4, and for blocks of all other shapes its size is 8. If mode refers to a MIP mode, the stitching is defined as follows: Matrix Multiplication With the averaged samples as input, a matrix-vector multiplication followed by the addition of an offset is performed. The result is a reduced prediction signal over a subsampled set of samples in the original block. red , the reduced prediction signal pred red , is generated, which is of width W red and height H red Here, W red and H red is defined as: Reduced prediction signal pred red It is calculated by taking the matrix-vector product and adding the offset: pred red =A·bdry red +b (2-13) Here, if w=H=4 columns, and in all other cases W=H=8 columns, then A is a red ·H red rows and 4 columns. b is a matrix of size W red ·H red The matrix A and the offset vector b are from the set S0, S1, S 2. One defines index idx=idx(W,H) as follows: Here, each coefficient of the matrix A is represented with 8 bits of precision. Set S0 consists of 16 matrices and 16 offset vectors Each matrix has 16 rows and 4 columns, and each offset vector has a size of 16. The matrices and offset vectors of this set are used for blocks of size 4×4. Set S1 consists of 8 matrices and 8 offset vectors Each matrix has 16 rows and 8 columns, and each offset vector has a size of 16. Set S2 consists of 6 matrices and 6 offset vectors Each matrix has 64 rows and 8 columns, and each offset vector has a size of 64. 2.15.3. Interpolation The prediction signals at the remaining positions are generated from the prediction signals on the subsample set by linear interpolation, which is a single-step linear interpolation in each direction. Regardless of the block shape or block size, the interpolation is performed first in the horizontal direction and then in the vertical direction. 2.15.4.MIP Mode Signaling and Coordination with Other Codec Tools For each codec unit (CU) in intra mode, a flag indicating whether the MIP mode is to be applied is sent. If the MIP mode is to be applied, the MIP mode (predModeIntra) is signaled. For the MIP mode, the transposed flag (isTransposed) that determines whether the mode is transposed and the MIP mode ID (modeId) that determines which matrix to use for a given MIP mode are derived as follows. isTransposed=predModeIntra&1 modeId=predModeIntra>>1 (2-15). The MIP codec mode is coordinated with other codecs by taking into account the following aspects: – Enable LFNST for MIP on large blocks. Here, the planar LFNST transform is used; – Reference sample derivation for MIP is performed in exactly the same way as for regular intra prediction modes; – For the upsampling step used in MIP prediction, the original reference samples are used instead of the downsampled reference samples; – Clipping is performed before upsampling, rather than after upsampling; – Regardless of the maximum transform size, MIPs are allowed to be up to 64×64. For sizeId=0, the number of MIP patterns is 32, for sizeId=1, the number of MIP patterns is 16, and for sizeId=2, the number of MIP patterns is 12. 2.16. Decoder-side intra-mode derivation In JEM-2.0, intra modes are expanded from 35 in HEVC to 67 modes, and they are derived at the encoder and explicitly signaled to the decoder. In JEM-2.0, a significant amount of overhead is spent on intra mode encoding and decoding. For example, in a full intra codec configuration, the intra mode signaling overhead can be as high as 5% to 10% of the total bitrate. This contribution proposes a decoder-side intra mode derivation scheme to reduce the intra mode encoding and decoding overhead while maintaining prediction accuracy. To reduce the overhead of intra mode signaling, this contribution proposes a decoder-side intra mode derivation (DIMD) scheme. In the proposed scheme, instead of explicitly signaling the intra mode, this information is derived from the neighboring reconstructed samples of the current block at both the encoder and the decoder. The intra mode derived by DIMD is used in two ways: 1) For a 2N×2N CU, when the corresponding CU-level DIMD flag is turned on, the DIMD mode is used as the intra mode for intra prediction; 2) For N×N CU, DIMD mode is used to replace a candidate in the existing MPM list to improve the efficiency of intra mode encoding and decoding. 2.16.1. Template-based intra-mode derivation like Figure 20 As shown, the target represents the current block (block size is N) for which the intra prediction mode is to be estimated. Figure 20 The pattern area indication in ( ) specifies a set of reconstructed samples that are used to derive the intra mode. The template size is expressed as the number of samples in the template that extend above and to the left of the target block, i.e., L. In the current implementation, for 4×4 and 8×8 blocks, a template size of 2 (i.e., L=2) is used, and for 16×16 and larger blocks, a template size of 4 (i.e., L=4) is used. The reference of the template (given by Figure 20 The dashed area in the figure (indicated by the dashed area in the figure) refers to a set of neighboring samples from the top and left of the template as defined by JEM-2.0. Unlike the template samples, which are always from the reconstructed area, the reference samples of the template may not have been reconstructed when the target block is encoded / decoded. In this case, the existing reference sample replacement algorithm of JEM-2.0 is utilized to replace the unavailable reference samples with available reference samples. For each intra prediction mode, DIMD calculates the SAD between the reconstructed template samples and its predicted samples obtained from the template's reference samples. The intra prediction mode that produces the smallest SAD is selected as the final intra prediction mode for the target block. 2.16.2. DIMD for Intra 2N×2N CU For Intra 2Nx2N CUs, DIMD is used as an additional Intra mode, which is adaptively selected by comparing the DIMD Intra mode with the best normal Intra mode (i.e., explicitly signaled). For each Intra 2Nx2N CU, a flag is signaled to indicate the use of DIMD. If the flag is 1, the Intra mode derived by DIMD is used to predict the CU; otherwise, DIMD is not applied and the Intra mode explicitly signaled in the bitstream is used to predict the CU. When DIMD is enabled, the chroma components always reuse the same Intra mode as the Intra mode derived for the luma component, i.e., DM mode. In addition, for each DIMD-encoded CU, blocks in the CU can adaptively choose to derive their intra mode at the PU level or the TU level. Specifically, when the DIMD flag is 1, another CU-level DIMD control flag is signaled to indicate the level at which DIMD is performed. If the flag is 0, it means that DIMD is performed at the PU level and all TUs in the PU use the same derived intra mode for their intra prediction; otherwise (i.e., the DIMD control flag is 1), it means that DIMD is performed at the TU level and each TU in the PU derives its own intra mode. In addition, when DIMD is enabled, the number of angular directions increases to 129, and DC and planar modes remain the same. To accommodate the increased granularity of angular intra modes, the precision of intra interpolation filtering for DIMD-encoded CUs is increased from 1 / 32 pixel to 1 / 64 pixel. In addition, in order to use the derived intra modes of DIMD-encoded CUs as MPM candidates for neighboring intra blocks, these 129 directions of DIMD-encoded CUs are converted to "normal" intra modes (i.e., 65 angular intra directions) before being used as MPMs. 2.16.3. DIMD for Intra N×N CU In the proposed method, the intra mode of an N×N CU is always signaled. However, to improve the efficiency of intra mode encoding and decoding, the intra mode derived from DIMD is used as an MPM candidate to predict the intra mode of the four PUs in the CU. In order not to increase the overhead of MPM index signaling, the DIMD candidate is always placed at the first position in the MPM list, and the last existing MPM candidate is removed. In addition, deduplication is performed so that if a DIMD candidate is redundant, it will not be added to the MPM list. 2.16.4.DIMD Intra-frame Pattern Search Algorithm To reduce the encoding / decoding complexity, a straightforward fast intra mode search algorithm is used for DIMD. First, an initial estimation process is performed to provide a good starting point for the intra mode search. Specifically, an initial candidate list is created by selecting N fixed modes from the allowed intra modes. Then, the SAD is calculated for all candidate intra modes, and the candidate intra mode that minimizes the SAD is selected as the starting intra mode. To achieve a good complexity / performance trade-off, the initial candidate list consists of 11 intra modes, including DC, planar, and every 4th mode of the 33 angular intra directions as defined in HEVC, i.e., intra modes 0, 1, 2, 6, 10…30, 34. If the starting Intra mode is DC or Planar, it is used as the DIMD mode. Otherwise, based on the starting Intra mode, a refinement process is then applied, where the best Intra mode is identified through an iterative search. It works by comparing the SAD values of three Intra modes separated by a given search interval at each iteration and maintaining the Intra mode that minimizes the SAD. The search interval is then reduced to half, and the selected Intra mode from the previous iteration will be used as the center Intra mode for the current iteration. For the current DIMD implementation with 129 angular Intra directions, a maximum of 4 iterations are used in the refinement process to find the best DIMD Intra mode. 2.17. Decoder-side Intra-mode Derivation by Computing Gradients of Neighboring Samples The three angle modes are selected from the Histogram of Gradients (HoG) calculated from the neighboring pixels of the current block. Once the three modes are selected, their prediction values are calculated normally, and then their weighted average is used as the final prediction value for the block. To determine the weights, the corresponding amplitude in the HoG is used for each of the three modes. The DIMD mode is used as an alternative prediction mode and is always checked in FullRD mode. The current version of DIMD has modified some aspects of signaling, HoG calculation, and prediction fusion. The purpose of these modifications is to improve codec performance and address the complexity issues raised during the last meeting (i.e., throughput of 4x4 blocks). The following sections describe the modifications made to each aspect. 2.17.1. Signaling Figure 21 Shown is the order of parsing flags / indexes in VTM5 integrated with the proposed DIMD. As can be seen, the DIMD flag of the block is first parsed using a single CABAC context, which is initialized to a default value of 154. If flag == 0, parsing continues normally. Otherwise (if flag == 1), only the ISP index is parsed and the following flags / indexes are assumed to be zero: BDPCM flag, MIP flag, MRL index. In this case, the entire IPM parsing is also skipped. During the parsing phase, when a regular non-DIMD block queries its DIMD neighbor's IPM, the pattern PLANAR_IDX is used as a virtual IPM for the DIMD block. 2.17.2. Texture Analysis DIMD's texture analysis includes the calculation of the Histogram of Gradients (HoG) ( Figure 22 ). The HoG calculation is performed by applying horizontal and vertical Sobel filters to the pixels in a template of width 3 around the block. Unless the above template pixels fall into a different CTU, they will not be used in the texture analysis. Once calculated, the IPMs corresponding to the two highest histogram bins are selected for the block. In previous versions, all pixels in the middle row of the template participated in the HoG calculation [1]. However, the current version improves the throughput of this process by applying the Sobel filter more sparsely on 4×4 blocks. For this purpose, only one pixel from the left and one pixel from above are used. This Figure 22 is shown in . Besides reducing the number of operations for gradient computation, this property also simplifies the selection of the best 2 modes from the HoG, since the resulting HoG cannot have more than two non-zero magnitudes. 2.17.3. Prediction Fusion The current method uses a fusion of three prediction values for each block. However, the choice of prediction mode is different and utilizes the combined hypothesis intra prediction method proposed in [2], where planar mode is considered to be combined with other modes when computing intra prediction candidates. In the current version, the two IPMs corresponding to the two highest HoG slices are combined with planar mode. Prediction fusion is applied as a weighted average of the above three prediction values. For this purpose, the weight of the plane is fixed to 21 / 64 (~1 / 3). The remaining 43 / 64 (~2 / 3) weight is then shared between the two HoG IPMs, proportional to the amplitude of their HoG stripes. Figure 23 The process is visualized. 2.18. Template-based Intra Mode Derivation (TIMD) This contribution proposes a Template-Based Intra Mode Derivation (TIMD) method using MPM, where TIMD modes are derived from MPM using neighboring templates. TIMD modes are used as an additional intra prediction method for a CU. 2.18.1. TIMD Mode Derivation For each intra prediction mode in the MPM, the SATD between the template's prediction and the reconstructed samples is calculated. The intra prediction mode with the smallest SATD is selected as the TIMD mode and used for intra prediction of the current CU. Position-dependent intra prediction combining (PDPC) is included in the derivation of the TIMD mode. 2.18.2.TIMD Signaling A flag is signaled in the sequence parameter set (SPS) to enable / disable the proposed method. When the flag is true, a CU-level flag is signaled to indicate whether the proposed TIMD method is used. The TIMD flag is signaled immediately after the MIP flag. If the TIMD flag is true, the remaining syntax elements related to luma intra prediction mode, including MRL, ISP, and the normal parsing phase of luma intra prediction mode, are skipped. 2.18.3. Interaction with new codec tools The DIMD method with prediction fusion using planes is integrated in EE2. When the EE2 DIMD flag is equal to true, the proposed TIMD flag is not signaled and is set equal to false. Similar to PDPC, gradient PDPC is also included in the derivation of TIMD mode. When the secondary MPM is enabled, both the primary MPM and the secondary MPM are used to derive the TIMD mode. The 6-tap interpolation filter is not used for the derivation of TIMD mode. 2.18.4. Modification of MPM List Construction in TIMD Mode Derivation During the construction of the MPM list, the intra prediction mode of the neighboring blocks is derived as a plane when they are inter-coded. To improve the accuracy of the MPM list, when the neighboring blocks are inter-coded, the propagated intra prediction mode is derived using the motion vector and reference picture and used in the construction of the MPM list. This modification is only applied to the derivation of TIMD mode. 2.18.5. TIMD with Fusion Instead of selecting only one mode with minimum SATD cost, this contribution proposes to select the top two modes with minimum SATD cost for intra modes derived using TIMD method, then fuse them with weights, and such weighted intra prediction is used to encode and decode the current CU. The costs of the two selected modes are compared with the threshold, and a cost factor of 2 is applied in the test as follows: Cost mode 2<2′cost mode 1. If this condition is true, fusion is applied, otherwise only mode 1 is used. The weight of a pattern is calculated from its SATD cost as follows: Weight 1 = Cost Mode 2 / (Cost Mode 1 + Cost Mode 2) Weight 2 = 1 – Weight 1. 2.19. Convolutional Cross-Component Model (CCCM) for Intra Prediction It is proposed to apply a convolutional cross-component model (CCCM) to predict chroma samples from reconstructed luma samples in a similar spirit to what the current CCLM mode does. Like CCLM, when chroma downsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid. In addition, similar to CCLM, there is an option to use a single model or a multi-model variant of CCCM. The multi-model variant uses two models, one model is derived for samples above the average luminance reference value, and the other model is derived for the remaining samples (following the spirit of CCLM design). Multi-model CCCM mode can be selected for PUs with at least 128 available reference samples. 2.19.1. Convolutional Filters The proposed convolutional 7-tap filter consists of a 5-tap plus sign-shaped spatial component, a nonlinear term, and a bias term. The input of the spatial 5-tap component of the filter consists of the center (C) luminance sample co-located with the chrominance sample to be predicted and its upper / northern (N) neighbor, lower / south (S) neighbor, left / west (W) neighbor, and right / east (E) neighbor, as shown in Figure 24 shown. The nonlinear term P is expressed as the square of the center luminance sample C and is scaled to the sample value range of the content: P=(C*C+midVal)>>bitDepth. That is, for 10-bit content, it is calculated as: P=(C*C+512)>>10. The bias term B represents a scalar offset between the input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content). The output of the filter is calculated as the filter coefficient c i Convolution with the input value and clipped to the range of valid chroma samples: predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B. 2.19.2. Calculation of filter coefficients Filter coefficient c i is calculated by minimizing the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region. Figure 25A reference region consisting of six rows of chroma samples above and to the left of the PU is shown. The reference region extends one PU width to the right of the PU boundary and one PU height below the PU boundary. The region is adjusted to include only available samples. The extension of the region shown in blue is required to support the "side samples" of the shaped spatial filter and is filled in when not in the available region. MSE minimization is performed by computing the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is LDL-decomposed, and the final filter coefficients are calculated using inverse substitution. This process roughly follows the calculation of the ALF filter coefficients in ECM, however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. The proposed method uses only integer arithmetic. 2.19.3. Bitstream Signaling The use of the mode is signaled via a PU-level flag for the CABAC codec. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered a sub-mode of CCLM. That is, the CCCM flag is signaled only when the intra prediction mode is LM_CHROMA_IDX (to enable single-mode CCCM) or MMLM_CHROMA_IDX (to enable multi-mode CCCM). 2.20. Gradient Linear Model (GLM) Compared to CCLM, GLM uses the gradient of luma samples to infer the linear model instead of the downsampled luma values. Specifically, when applying GLM, the input of CCLM process (i.e., the downsampled luma samples L) is replaced by the gradient of luma samples G. The other parts of CCLM (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged. C=α·G+β For signaling, when CCLM mode is enabled for the current CU, two flags are transmitted separately for the Cb component and the Cr component to indicate whether GLM is enabled for each component; if GLM is enabled for a component, a syntax element is further transmitted by signal to select one of the four gradient filters for gradient calculation. ·like Figure 26 As shown, four gradient filters are enabled for the GLM. GLM with brightness In ECM-6.0[1], GLM uses the gradient of luma samples to predict chroma samples as follows: pred C (i, j) = α·G(i, j) + β where pred C(i, j) represents the predicted value of the chrominance sample, G(i, j) represents the gradient of the corresponding reconstructed luminance sample, and the linear model parameters α and β are derived as CCLM based on the linear minimum mean square error (LMMSE) method from adjacent reconstructed samples. A new GLM model is proposed, in which the chrominance samples are based on the gradient G(i,j) of the luma samples and the reconstructed value rec′ of the downsampled luma samples. L (i,j) is predicted using different parameters: pred C (i,j)=α0·G(i,j)+α1·rec′ L (i,j)+α2·midValue The model parameters α0, α1, and α2 are derived from six rows and six columns of adjacent sample points based on the LDL decomposition method as the CCCM mode in ECM-6.0. 2.21. Gradient and Position-Based Convolutional Cross-Component Model (GL-CCCM) for Intra Prediction The proposed GL-CCCM method uses gradient and position information to replace the four spatial neighboring samples in the CCCM filter. The GL-CCCM filter used for prediction is: predChromaVal=c0C+c1G y +c2G x +c3Y+c4X+c5P+c6B Among them G y and G x are the vertical and horizontal gradients respectively, and are calculated as: G y =(2N+NW+NE)–(2S+SW+SE) G x =(2W+NW+SW)–(2E+NE+SE). In addition, the Y parameter and the X parameter are the vertical position and the horizontal position of the center luma sample, and they are calculated relative to the top left coordinate of the block. The remaining parameters are the same as those of the CCCM tool. The reference area used for parameter calculation is the same as that of the CCCM method. Figure 27 The spatial domain samples used for GL-CCCM are shown. Bitstream signaling The use of this mode is signaled via a PU-level flag for the CABAC codec. A new CABAC context is included to support this. When it comes to signaling, GL-CCCM is considered a submode of CCCM. That is, the GL-CCCM flag is signaled only when the original CCCM flag is true. Encoder Operation The encoder performs two new RD checks in the chroma prediction mode loop, one for single-model GL-CCCM mode and one for multi-model GL-CCCM mode. 2.22. CCCM using non-subsampled luma samples 2.22.1. Block Level In this contribution, CCCM using non-subsampled luma samples is proposed, where the chroma samples are predicted directly from the original reconstructed luma samples, i.e., without downsampling. Figure 28 As shown in Figure 2, the proposed CCCM filter consists of a 6-tap spatial term, two nonlinear terms, and a bias term. The 6-tap spatial term is related to the chrominance samples to be predicted (i.e., The chroma samples of up to 6 rows / columns on the left are used to derive the filter coefficients. The filter coefficients are derived based on the same LDL decomposition method used in CCCM. In this contribution, the proposed method is signaled as an additional CCCM model in addition to the existing CCCM model. For signaling, when CCCM is selected, a single flag is signaled and used for both chroma components to indicate whether the default CCCM model or the proposed CCCM model is applied. 2.22.2. High-level control For content with sharp details (such as SCC content), downsampling of the luma component may not be optimal for CCCM model derivation. In this contribution, it is proposed to disable luma downsampling and derive and apply the model directly on non-downsampled luma samples. If downsampling is not applied, the CCCM model shape is a diamond 5x5. The SPS flag is signaled to indicate whether luma downsampling is applied to the CCCM. 2.23. Airspace GPM (SGPM) In spatial GPM, a candidate list is constructed that includes partitioning and two intra prediction modes. Up to 11 MPMs of intra prediction modes are used to form a combination, and the length of the candidate list is set to equal 16. The selected candidate index is transmitted through the signal. Figure 29 Spatial GPM candidates are shown. The list is reordered using the template shown in the figure above. The GPM blending process is not used in the template, and the SAD between the prediction and reconstruction of the template is used for sorting. Figure 30 A GPM template is shown. The SGPM mode is applied to blocks whose width and height satisfy the same restrictions as in inter-frame GPM. The following projects are considered: Airspace GPM segmentation mode: 26 predefined modes; An adaptive inference algorithm based on the ratio of horizontal gradient to vertical gradient. Intra-frame prediction mode selection: List of IPMs with and without TIMD: For each segmentation mode, an IPM list is derived for each part using the intra-inter GPM list derivation. The IPM list size is 3. In the list, the TIMD-derived pattern is replaced by 2 derived patterns with horizontal and vertical directions (using the upper template or the left template), or the TIMD-derived pattern is excluded. MPM List: A unified MPM list (maximum 11 elements) is used for all segmentation modes. Template size (left and top): 1 or 4. Extended block size: Spatial GPM is extended to be further applied to 4x8 blocks, 8x4 blocks, 4x16 blocks and 16x4 blocks, which can be described as 4<=width<=64, 4<=height<=64, width<height*8, height<width*8, width*height>=32. Adaptive Mixing: Adaptive mixing is tested for spatial GPM, where the mixing depth τ is derived as follows: ■If min(width, height) == 4, then 1 / 2τ is selected, ■ Otherwise, if min(width, height) == 8, then τ is selected, ■ Otherwise, if min(width, height) == 16, then 2τ is selected, ■ Otherwise, if min(width, height) == 32, then 4τ is selected, ■Otherwise, 8τ is selected. Figure 31 GPM mixing is shown. 2.24. Signaling of cross-component prediction modes in ECM In ECM-7, cross-component modes include CCLM, CCLM-L, CCLM-T, MM-CCLM, MM-CCLM-L, MM-CCLM-T and CCCM, CCCM-L, CCCM-T, MM-CCCM, MM-CCCM-L, MM-CCCM-T. A flag is transmitted through the signal to determine whether it is a CCCM mode or a CCLM mode. The truncation unary code is used to indicate Figure 32 CCLM mode or CCCM mode shown. CCLM or CCCM: 0; MM-CCLM or MM-CCCM: 10; CCLM-L or CCCM-L: 110; CCLM-T or CCCM-T: 1110; MM-CCLM-L or MM-CCCM-L: 11110; MM-CCLM-T or MM-CCCM-T: 11110. 2.25. Slope Adjustment for CCLM CCLM uses a 2-parameter model to map luma values to chroma values. The slope parameter "a" and the bias parameter "b" define the mapping as follows: Chromaticity value = a*luminance value + b. It is proposed to adjust the slope parameter “u” via signal transmission to update the model to the following form: Chromaticity value = a'*luminance value + b' in a'=a+u b'=bu*y r . With this choice, the mapping function is centered around the brightness value y r The points are tilted or rotated. It is proposed to use the average value of the reference brightness samples used in model creation as y r , in order to provide meaningful modifications to the model. The following figure illustrates this process. 2.26. Fusion of Chroma Intra Prediction Modes In test 1.2b, it is proposed that the DM mode and the four default modes can be fused with the MMLM_LT mode as follows: pred=(w0*pred0+w1*pred1+(1<<(shift-1)))>>shift Where pred0 is the prediction value obtained by applying the non-LM mode, pred1 is the prediction value obtained by applying the MMLM_LT mode, and pred is the final prediction value of the current chroma block. The two weights w0 and w1 are determined by the intra prediction mode of the adjacent chroma blocks, and shift is set to be equal to 2. Specifically, when the upper adjacent block and the left adjacent block are both encoded and decoded in the LM mode, {w0, w1} = {1, 3}; when the upper adjacent block and the left adjacent block are both encoded and decoded in the non-LM mode, {w0, w1} = {3, 1}; otherwise, {w0, w1} = {2, 2}. For syntax design, if non-LM mode is selected, a flag is signaled to indicate whether fusion is applied, and the proposed fusion is only applied to I slices. 2.27. History-based Cross-Component Prediction (H-CCP) 1. It is proposed that the model(s) (such as CCLM or CCCM) of cross-component prediction (CCP) in a block can be stored into a history table (HT). a.HT is a list with ordered entries. i. Each entry has an index. For example, the first entry has an index of 0, and subsequent entries have indices of 1, 2, 3, ... b. The model parameters of CCLM and its variants can include a, b, and displacement to control the calculation accuracy. c. The model parameters of CCLM and its variants may include a linear part (such as c0-c4) and a nonlinear part (such as c5). d. Models may include models for different color components such as Cb and Cr. i. For example, the models for Cb and Cr can be coupled in the entry. e. In one example, different CCPs such as CCLM and CCCM can share the same HT. i. In one example, a segment in an entry of an HT may reflect the type of CCP model(s) stored in the entry. f. In one example, different CCPs such as CCLM and CCCM may have different HTs. i. In one example, a CCLM_HT can store models of CCLM and its variants such as CCLM-L or CCLM-T. ii. In one example, a CCCM_HT can store models of CCCM and its variants such as CCCM-T or CCCM-T. g. In one example, a CCP with a single model (such as CCLM or CCCM) and a CCP with multiple models (such as MM-CCLM or MM-CCCM) may have different HTs. h. In one example, a CCP with a single model (such as CCLM or CCCM) and a CCP with multiple models (such as MM-CCLM or MM-CCCM) can share the same HT. i. In one example, the segment in an entry of the HT may reflect the number of models stored in the entry. ii. In one example, a segment in an entry of HT may reflect at least one threshold of a model used to classify samples into different groups. i. In one example, the first HT is used to store models of CCLM and its variants. i. In one example, CCLM variants may include CCLM-L, CCLM-T, MM-CCLM, MM-CCLM-L, MM-CCLM-T, GLM, and CCLM with slope adjustment. 1) The segment in an entry of HT may reflect the number of models stored in the entry. 2) The segments in the entries of HT may reflect at least one threshold of the model used to classify samples into different groups. 3) The segment in the HT entry may reflect whether GLM is applied. 4) The segments in the entries of HT may reflect the downsampling filters of GLM. j. In one example, the second HT is used to store models of CCCM and its variants. i. In one example, CCCM variants may include CCCM-L, CCCM-T, MM-CCCM, MM-CCCM-L, MM-CCCM-T. 1) The segment in an entry of HT may reflect the number of models stored in the entry. 2) The segments in the entries of HT may reflect at least one threshold of the model used to classify samples into different groups. 2. It is proposed that a block can be encoded and decoded in a history-based CCP (H-CCP) mode, in which at least one CCP model used by the current block is obtained or derived from the HT. a. In one example, at least one syntax element (SE) may be signaled to indicate whether H-CCP is applied. i. In one example, SE can be conditionally signaled. For example, SE is signaled only when a specific mode (such as CCCM or CCLM) is used. 1) For example, SE is signaled only when the current mode is CCCM or CCLM. b. In one example, at least one syntax element (SE) may be signaled to indicate which entry in the HT is retrieved to derive model(s) for cross-component prediction. i.SE can reflect the index in HT. 1) In one example, SE may be set equal to f(k), where k is an index and f is a function. 2) In one example, SE may be set equal to f(k, M), where k is an index, M is the number of valid entries in the HT, and f is a function. a) In another example, M is the size of HT. 3) In one example, SE may be set equal to k, where k is an index. 4) In one example, SE may be set equal to M-1-k, where k is the index and M is the number of valid entries in the HT. a) In another example, M is the size of HT. ii. SE can reflect the index of the list, and the list can be constructed based on HT. 1) In one example, the list L is constructed by reversing HT. For example, L[i]=HT[M-1-i], where M is the number of valid entries in HT. a) In another example, M is the size of HT. b) In one example, L may have a fixed size. c) In one example, if L is not full, the vacant entry is filled with a default entry. iii. In one example, SE can be signaled conditionally, for example, SE is signaled only when H-CCP is applicable. iv. SE can be signaled only when more than one entry in HT can be selected. The maximum value of v.SE (denoted as V) is determined by the number of entries to be selected. 1) For example, V=K, or V=K-1, or V=K+1, or V=K-2, or V=K+2. c. In one example, at least one syntax element (SE) may be signaled to indicate which HT is used. i. In one example, SE can be signaled conditionally, for example, SE is signaled only when H-CCP is applicable. ii. SE can be signaled only when more than one HT can be selected. d. In one example, which HT to use can be derived at the encoder / decoder. i. In one example, if the current mode is CCLM, the first HT storing the model of CCLM and its variants is used. ii. In one example, if the current mode is CCCM, the second HT storing the model of CCCM and its variants is used. e. In one example, the current block may be predicted with a CCP model obtained from the determined entry of the determined HT. f. In one example, the current block may be predicted using CCCM or CCLM based on whether the first HT or the second HT is applied. g. In one example, the current block can be predicted using multiple models. i. Whether a single model or multiple models are applied can be deduced / obtained from the determined entries of the determined HT. ii. At least one threshold value of the model for classifying the samples into different groups may be obtained / derived from the determined entries of the determined HT. HT Maintenance 3. The maximum size of HT can be predetermined, such as 5 or 6. a. Alternatively, the maximum size of the HT can be signaled as SE at the block level / sequence level / picture group level / picture level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB or sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. b. Alternatively, the maximum size of the HT can be derived using encoding / decoding information such as: i. The mode of the current block; ii. Patterns of neighboring blocks; iii. The pattern of the luminance blocks in the same region as the current block; iv. The pattern of luminance blocks in the same region as the neighboring blocks; v.QP; vi. Strip / image type; vii. Image width / height; viii.Block width / height; ix. Reconstructed sample points. 4. HT can be refreshed at the beginning of a coding / decoding sequence / picture / slice / slice / sub-picture / CTU row / CTU. a. For example, HT can be refreshed by clearing the table. b. For example, the HT can be refreshed by populating the table with default entries. 5. After encoding / decoding a block (such as a CU), the HT may be updated. a. For example, when dual-tree coding is applied, the CU must be a chroma CU. b. For example, the CU must be a CU with CCP mode. c. For example, which HT to be updated may depend on the coding mode of the CU. i. For example, if the CU is encoded and decoded in CCLM mode (such as CCLM, CCLM-L, CCLM-T, MM-CCLM, MM-CCLM-L, MM-CCLM-T, GLM, and CCLM with slope adjustment), then (multiple) models and related information (such as (multiple) thresholds of the model for classifying samples into different groups) are stored in the first HT. ii. For example, if the CU is encoded and decoded in CCCM mode (such as CCCM, CCCM-L, CCCM-T, MM-CCCM, MM-CCCM-L, MM-CCCM-T), (multiple) models and related information (such as (multiple) thresholds of the model for classifying samples into different groups) are stored in the first HT. d. For example, a set of information about the CCP model(s) used by the current block can be placed in the HT. i. The group may include one or more CCP models. ii. The number of models that the group may include. iii. The set may include the threshold(s) of the model used to classify the samples into different groups. iv. The group may include slope adjustments. e. In one example, if the current block is encoded using CCLM with slope adjustment, the CCP model can be adjusted before being used to update the HT. 6. How a new set of information related to the CCP model(s) is placed into the HT may depend on whether the HT is full or not. a. For example, if the HT is not full, the new group can be placed into the first empty entry of the HT. i. For example, the first missing entry is the missing entry with the smallest index. ii. For example, the first missing entry is the missing entry with the largest index. iii. After being placed in the HT, the new group may be placed as the last occupied entry in the HT. 1) The last occupied entry may be the occupied entry with the largest index. 2) The last occupied entry may be the occupied entry with the smallest index. b. For example, if the HT is full, an existing entry in the HT may be removed. i. In one example, HT can be managed in a first-in-first-out manner. ii. The existing entry with the smallest index may be removed. 1) The updated HT' may be set to: HT'[i]=HT[i+1], for 0<=i<=N-2, and HT'[N-1]=new group, where N is the size of the HT. iii. The existing entry with the largest index may be removed. 1) The updated HT' may be set to: HT'[i]=HT[i-1], for 1<=i<=N-1, and HT'[0]=new group, where N is the size of the HT. 7. In one example, the new group may be compared to at least one existing entry in the HT to determine whether to place the new group and / or how to update the HT. 8. In one example, if the new group is identical or similar to one of the existing entries in the HT, the new group is not placed in the HT. It is assumed that the new group is identical or similar to a particular entry of the HT. a. For example, in this case, the special entry may be placed in the first HT of the HT, and the entry that originally preceded the special entry is pushed back one position. i. For example, assuming the entry is HT[i] (i=0,1…), and the special entry is HT[k], the updated HT' will be as follows: HT'[0]=HT[k]; HT'[i]=HT[i-1] (for 1<=i<=k); HT'[i]=HT[i] (for i>k). b. For example, in this case, the special entry may be placed at the end of the HT, and the entry that originally preceded the special entry may be pushed forward one position. i. For example, assuming that the entry is HT[i] (i=0, 1…), and the special entry is HT[k], the updated HT' will be as follows: HT'[N-1]=HT[k]; HT'[i]=HT[i+1] (for k<=i<=N-2); HT'[i]=HT[i] (for i <k)。 9. In one example, whether to place a new group and / or how to update the HT may depend on the codec information of the CU with the new group. 10. In one example, if the new group is a new group of CUs coded in H-CCP mode, the new group is not placed in the HT. It is assumed that the special entry in the HT is used by CUs coded with H-CCP. a. For example, in this case, the special entry may be placed in the first HT of the HT, and the entry that originally preceded the special entry is pushed back one position. i. For example, assuming the entry is HT[i] (i=0,1…), and the special entry is HT[k], the updated HT' will be as follows: HT'[0]=HT[k]; HT'[i]=HT[i-1] (for 1<=i<=k); HT'[i]=HT[i] (for i>k). b. For example, in this case, the special entry may be placed at the end of the HT, and the entry that originally preceded the special entry may be pushed forward one position. i. For example, if the entry is HT[i] (i=0,1…), and the special entry is HT[k], the updated HT' will be as follows: HT'[N-1]=HT[k]; HT'[i]=HT[i+1] (for k<=i<=N-2); HT'[i]=HT[i] (for i <k)。 11. It is proposed that the entry for HT may include models for more than one chroma component (such as Cb and Cr). a. If the entry is selected, the models for components Cb and Cr are applied to the two components separately. 12. It is proposed that the entry for HT may include a model for only one component (such as Cb or Cr). a. If the entry is selected, the model for a specific component such as Cb or Cr is applied to the specific component. b. In one example, different HTs can be constructed for different components. List Mode 13. It is proposed that at least one list with a CCP model can be constructed. a. In one example, chroma blocks can be predicted in "list mode" using the CCP model in the list. b. In one example, list L may be populated with one type of CCP model, such as CCCM. c. In one example, the list may be populated with multiple types of CCP models, such as both CCCM and CCLM. i. In one example, the type of CCP model will be stored in a list along with the CCP model. d. In one example, at least one syntax element (SE) may be signaled to indicate whether a CCP model in the list is used. i. In one example, SE can be conditionally signaled. For example, SE is signaled only when a specific mode (such as CCCM or CCLM) is used. 1) For example, SE is signaled only when the current mode is CCCM or CCLM. 2) For example, SE is signaled only when "list mode" is applicable. e. In one example, at least one syntax element (SE) may be signaled to indicate which entry in the list is used to derive the model(s) for cross-component prediction. i.SE can reflect the index in the list. 1) In one example, SE may be set equal to f(k), where k is an index and f is a function. 2) In one example, SE may be set equal to f(k, M), where k is the index, M is the number of valid entries in the list, and f is a function. a) In another example, M is the size of the list. 3) In one example, SE may be set equal to k, where k is an index. 4) In one example, SE may be set equal to M-1-k, where k is the index and M is the number of valid entries in the list. a) In another example, M is the size of the list. f. In one example, L may have a fixed size. g. In one example, multiple lists can be constructed. i. For example, at least one syntax element (SE) may be signaled to indicate which list is used. ii. In one example, SE can be signaled conditionally, for example, SE is signaled only when "list mode" is applicable. iii. SE can be signaled only when more than one list can be selected. h. In one example, which list to use can be derived at the encoder / decoder. i. In one example, if the current mode is CCLM, a first list of models storing CCLM and its variants is used. ii. In one example, if the current mode is CCCM, a second list of models storing CCCM and its variants is used. 14. It is proposed that the entries of the list may include models for more than one chroma component (such as Cb and Cr). a. If the entry is selected, the models for components Cb and Cr are applied to the two components separately. 15. An entry in the proposed list may include a model for only one component (such as Cb or Cr). a. If the entry is selected, the model for a specific component such as Cb or Cr is applied to the specific component. 16. Multiple candidates can be placed in the list, including: a. CCP model of adjacent neighboring blocks. b. CCP model of non-adjacent neighboring blocks. c. CCP model of the same-position block in the reference image. d. CCP model of the reference block in the reference image. e. CCP model in the history table. f. CCP model derived from non-adjacent sample points. g. Default CCP mode. 17. In one example, the list can be constructed by examining the possible candidates in order. a. For example, the order can be adjacent neighboring blocks, non-adjacent neighboring blocks, models in the history table, and models derived from non-adjacent samples. b. For example, if the number of candidates in the list reaches the maximum allowed size of the list (such as 5 or 6), then list building is completed. c. For example, if the number of candidates in the list reaches f(d), then the list construction is completed, where d is the index of the selected candidate and f is a function. For example, f(d)=d+1. d. For example, if all possible candidates have been examined and the build is not complete, a default model may be placed in the list. 18. In one example, if a potential candidate is placed in a list, it may be compared to at least one existing candidate in the list. a. For example, if a potential candidate is the same as or similar to an existing candidate, the potential candidate is not placed in the list. b. In one example, if a potential entry of CCP information is placed into a history-based table, it may be compared to at least one existing entry in the list. i. For example, if a potential entry is identical or similar to an existing entry, the potential entry is not placed in the list. c. In one example, two CCP candidates or entries are determined to be different if the following conditions are met: i. CCP types are different. ii. The number of models is different. iii. If the CCP has multiple models, the threshold is different. iv. At least one model is different. v. Luma sample offset is different. (Applicable only when type is CCCM or GL-CCCM or GPM or CCCM using non-subsampled luma samples.) vi. Sample point position displacement is different. (Applicable only when the type is GL-CCCM) 19. For example, the CCP information of an entry in a history-based table or the candidate CCP information in a CCP candidate list may include: a. Type of CCP method, such as CCLM or CCCM or GLM or GLM with luma or GL-CCCM or CCCM using non-subsampled luma samples. i. In one example, GLM methods using different downsampling filters can be considered as different types. ii. In one example, GLM methods with luminance using different downsampling filters can be considered as different types. iii. In one example, the types may be CCCM, CCLM, 4 types of GLM using different downsampling filters, 4 types of GLM with luma using different downsampling filters, GL-CCCM, and CCCM using non-downsampled luma samples. iv. “Not using CCP codec” (denoted as NonCCP) can also be considered as a type. b. Position (x, y). c. Number of models. i. For example, the number of models can be 1 or 2. ii. In one example, the number of models can be considered as part of the CCP type. For example, CCLM and MM-CCLM can be considered as two types. d. At least one threshold used to classify samples for different models. i. Threshold can be used only when the number of models is at least 2. e. At least one luma sample value offset. i. Luma Sample Value Offsets can be added to or subtracted from luma samples (which may be downsampled) when used to derive chroma prediction values. ii. Luma sample value offset can be used only for certain types, such as CCCM, GLM with luma, GL-CCCM, and CCCM using non-subsampled luma samples. f. At least one chroma sample value is offset. i. Chroma sample value offsets can be added to or subtracted from the chroma prediction values derived from the CCP model to generate the final prediction. g. At least one model for at least one chroma component. i. For example, it may include different models for the Cb component and the Cr component. ii. For example, the number of models for each component may be included as part of the information. iii. The model can be represented by a model form of CCLM or a model form of CCCM or a model form of GLM or a model form of GLM with luma or a model form of GL-CCCM or a model form of CCCM using non-subsampled luma samples. h. At least one sample point position displacement expressed as (dX, dY). i. Chroma Sample Positions A chroma sample position offset can be added to or subtracted from the sample position (x, y) when used to derive the chroma prediction value. ii. Chroma sample position shifting can only be used for certain types, such as GL-CCCM. 20. For example, the CCP codec information of a chroma block after being encoded / decoded can be stored in a history-based table or in a CCP candidate list. a. In one example, the CCP codec information can be stored only when the chroma block is coded in CCP mode. i. In one example, if the chroma block is coded in at least one CCP mode (such as fusion in chroma intra prediction mode), the CCP codec information may be stored. 1) The stored type may be set as the CCP type used in fusion of chroma intra prediction modes. b. In one example, CCP codec information can be stored for any chroma block. i. If the chroma block is not coded in CCP mode, the type is stored as "NonCCP". c. If the chroma block is encoded and decoded in CCP mode, the type of information can be stored according to the encoding and decoding mode. i. If the mode is CCCM or CCCM-T or CCCM-L or MM-CCCM or MM-CCCM-T or MM-CCCM-L, the type is set to "CCCM". ii. If the mode is CCLM or CCLM-T or CCLM-L or MM-CCLM or MM-CCLM-T or MM-CCLM-L, the type is set to "CCLM". iii. If the mode is CCLM or CCLM-T or CCLM-L or MM-CCLM or MM-CCLM-T or MM-CCLM-L with slope adjustment, the type is set to "CCLM". iv. If the mode is GLM with filter X, the type is set to "GLM with filter X". v. If the mode is GLM with Luminance using filter X, then the type is set to "GLM with Luminance using filter X". vi. If the mode is GL-CCCM, the type is set to "GL-CCCM". vii. If the mode is CCCM with non-subsampling, the type is set to "CCCM with non-subsampling". viii. If the mode is a fusion of chroma intra prediction modes, the type is set to "CCLM". d. The number of models can be stored as the number of models of the chroma block. i. For example, if the mode is MM-CCLM or MM-CCLM-T or MM-CCLM-L or MM-CCLM or MM-CCLM-T or MM-CCLM-L or any other multi-model CCP mode (such as GLM or GL-CCCM or CCCM with multiple models using non-subsampled luma samples), the number of models is set to 2. e. Information such as thresholds, luma / chroma sample value offsets, and sample position displacements may be stored as information used by the chroma block. f. The CCP model for a component can be stored as the model used by the chroma block. i. The model can be derived by any CCP method, such as CCLM or CCLM-T or CCLM-L or MM-CCLM or MM-CCLM-T or MM-CCLM-L or CCCM or CCCM-T or CCCM-L or MM-CCCM or MM-CCCM-T or MM-CCCM-L or GLM using different downsampling filters or GLM with luma using different downsampling filters or GL-CCCM or CCCM using non-downsampled luma samples. ii. The stored model may be the one that is ultimately applied, such as after being modified by slope adjustment. 21. In one example, a history table of CCP information after encoding / decoding a region (such as a CU / CTU / CTU row) may be stored, referred to as a storage table. a. A history table of CCP information maintained for the current block (referred to as an online table) may be used together with a stored history table of CCP information. b. In one example, entries in the storage table and the online table may be checked sequentially to generate new candidates. i. In one example, entries in the online table may be checked before all entries in the storage table. ii. In one example, entries in the storage table may be checked before all entries in the online table. iii. For example, the kth entry in the storage table may be checked after the kth entry in the online table. iv. For example, the kth entry in the online table may be checked after the kth entry in the storage table. v. For example, the kth entry in the online table may be checked after storing all mth entries in the table, for m = 0...S, where S is an integer. vi. For example, the kth entry in the storage table may be checked after all mth entries in the online table, for m=0...S, where S is an integer. vii. For example, the kth entry in the online table may be checked after storing all mth entries in the table, for m = S...maxT, where S is an integer and maxT is the last entry. viii. For example, the kth entry in the storage table may be checked after all mth entries in the online table, for m=S...maxT, where S is an integer and maxT is the last entry. c. In one example, which storage table(s) to use may depend on the dimensions and / or position of the current block. i. For example, a table stored in a CTU above the current CTU may be used. ii. For example, a table stored in the CTU to the upper left of the current CTU may be used. iii. For example, a table stored in the CTU to the upper right of the current CTU may be used. d. In one example, whether and / or how to use the storage table may depend on the dimensions and / or location of the current block. i. In one example, whether and / or how to use the storage table may depend on whether the current CU is at the top boundary of a CTU and whether an upper neighboring CTU is available. 1) For example, the storage table can be used only when the current CU is at the top boundary of a CTU and an upper adjacent CTU is available. 2) For example, if the current CU is at the top boundary of a CTU and an upper adjacent CTU is available, at least one entry in the storage table may be placed at a more forward position. e. In one example, entries in two storage tables may be examined sequentially to generate new candidates. i. For example, the first (or second) storage table stored in the CTU above the current CTU may be used. ii. For example, the first (or second) storage table stored in the CTU to the upper left of the current CTU may be used. iii. For example, the first (or second) storage table stored in the CTU to the upper right of the current CTU may be used. 3. Question 1. The model for cross-component prediction is trained using adjacent neighboring samples, which can be inefficient. 4. Specific Implementation Methods The following specific embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow manner. In addition, these embodiments can be combined in any way. In the following discussion, CCCM may refer to the original CCCM mode, or it may refer to variants of CCCM, such as CCCM-L, CCCM-T, MM-CCCM, MM-CCCM-L, MM-CCCM-T. In the following discussion, CCLM may refer to the original CCLM mode, or it may refer to variants of CCLM, such as CCLM-L, CCLM-T, MM-CCLM, MM-CCLM-L, MM-CCLM-T, etc. In the following discussion, cross-component prediction (CCP) may refer to any cross-component prediction, such as CCLM or CCCM or GLM or CCLM with slip compensation. 1. It is proposed that the model(s) for cross-component prediction in a block (such as CCLM or CCCM) can be derived based on a set of samples that are not adjacent to the current block, which is called non-adjacent cross-component prediction (NA-CCP). a. In one example, a group of samples is non-adjacent to the current block only if none of the samples in the group is adjacent to the current block (such as adjacent above or adjacent to the left of the current block). b. In one example, a set of samples is reconstructed before encoding / decoding the current block. c. The samples may include chroma samples and / or their corresponding luma samples. If the color format is 4:2:0 or 4:2:2, the luma samples may be generated by downsampling. 2. In one example, at least one syntax element (SE) may be signaled to indicate whether non-adjacent cross-component prediction is applied. In one example, SE can be signaled conditionally. For example, SE can be signaled only when a specific mode (such as CCCM or CCLM) is used. 3. In one example, more than one set of samples that are not adjacent to the current block can be used to derive model(s) for cross-component prediction. a. In one example, samples in more than one group can be used jointly to derive model(s) for cross-component prediction. b. In one example, one of multiple candidate groups may be selected to derive model(s) for cross-component prediction. 4. In one example, at least one syntax element (SE) may be signaled to indicate which set of non-adjacent samples is used to derive the model(s) for cross-component prediction. In one example, SE can be signaled conditionally, for example, only when NA-CCP is applicable. b. SE can be signaled only when more than one set of non-adjacent samples can be selected. c. The maximum value of SE (denoted as V) is determined by the number of groups of non-adjacent samples to be selected (denoted as K). i. For example, V=K, or V=K-1, or V=K+1, or V=K-2, or V=K+2. 5. Whether / how NA-CCP is applied may be the same for more than one color component (such as Cb and Cr). a. Alternatively, whether / how to apply NA-CCP may be different for different components (such as Cb and Cr). 6. Whether NA-CCP is applicable may depend on the dimension / position of the current block. 7. In one example, a group of non-adjacent samples may include samples in a region. a. In one example, a region may be a codec block (eg, a CU). b. In one example, a region may be represented by a position relative to the region. c. In one example, the region may be an M×N rectangle (eg, M=N=8). d. In one example, a rectangular region may be represented by a position relative to the region (such as the top left position of the region (x, y)) and dimensions M×N. e. In one example, regions of non-adjacent samples from different groups may share the same shape and size. f. In one example, regions of non-adjacent samples in different groups may have different shapes or sizes. g. Sample points in the area must be reconstructed. i. Alternatively, if the samples in the region are not reconstructed, they should be filled. 8. In one example, luma samples corresponding to a set of non-adjacent chroma samples may be prepared or generated to be used for training a cross-component model. a. In one example, if the color format is 4:2:0 or 4:2:2, downsampling may be applied to generate corresponding luma samples. b. In one example, the generated luma samples may correspond to a larger area than the area of non-adjacent chroma samples. i. In one example, if the area of non-adjacent chroma samples is an M×N rectangle, the generated luma samples may correspond to a (M+T+B)×(N+L+R) chroma rectangle, such as Figure 33 shown. 1) In one example, T=B=L=R=1. c. In one example, if the luma sample to be generated is not available (e.g., it is outside the picture boundary, or it has not been reconstructed, or it is in a different CTU that has not been reconstructed, etc.), then the luma sample can be handled specially. i. In one example, the luma samples may be padded, such as repeatedly filled with nearby available generated luma values. ii. In one example, it may not be generated and not marked as “unavailable”. 1) The dimension of the brightness area can be set to the available area. 9. In one example, whether a region comprising non-adjacent sample points is a valid set of sample points for deriving model(s) may be determined by the availability of at least one sample point in the region. a. For example, the region is a rectangle. b. For example, a region is determined to be valid only when both the upper left reconstruction sample point and the lower right reconstruction sample point of the region are available. c. For example, a region is determined to be valid only when both the upper right reconstruction sample point and the lower left reconstruction sample point of the region are available. 10. In one example, a region list can be constructed to record multiple groups of non-adjacent sample points. a. In one example, the index of the list can be signaled as SE to indicate which set of non-adjacent samples is used to derive the model(s) for cross-component prediction. i. For example, SE can be binarized into a truncated unary code. ii. In one example, SE can be conditionally signaled, for example, SE is signaled only when NA-CCP is applied. iii. SE can be signaled only when more than one set of non-adjacent samples can be selected. iv. The maximum value of SE (denoted as V) is determined by the number of groups of non-adjacent samples to be selected (denoted as K). 1) For example, V=K, or V=K-1, or V=K+1, or V=K-2, or V=K+2. b. In one example, a list may be constructed by examining multiple potential candidate regions in sequence. i. The list is initialized to empty. ii. If the number of candidate regions in the list is equal to the maximum size of the list (such as 6), then the list building is completed. iii. If all potential candidate regions have been examined, list building is complete. iv. If the region is determined to be valid, the potential candidates can be put into a list. v. Deduplication may be applied to construct the list. 1) If a potential candidate is a "duplicate" of an existing candidate in the list, the potential candidate may not be placed in the list. a) If the samples of a candidate region are the same as (or similar to) the samples of another region, then the candidate region “overlaps” with the other region. b) A candidate region "duplicates" another region if the same or similar model can be derived from samples in those two regions. 11. In one example, the position and / or dimensions of the region including non-adjacent samples may depend on codec information, such as the width / height of the current block. a. This area may be a potential candidate area for the list. b. The distance between the region and the current block may depend on the width / height of the current block. 12. In one example, the potential candidate region may be an M×N rectangle (eg, M=N=8) non-adjacent to the left / lower left / upper left / above / upper right of the current block. Figure 34 An example is shown. 13. In one example, the potential candidate region is an M×N (e.g., M=N=8) rectangle, and its top left position (x0, y0) can be described as (assuming the top left position of the current block with dimensions W×H is (0, 0)): a. (x0, y0) = (s*f(W, H), t*g(W, H)), where f and g are functions. s and t are scaling factors, such as 0.5, 1, or 2. b. (x0, y0) = (s*f(W), t*g(H)), where f and g are functions. s and t are scaling factors, such as 0.5, 1, or 2. 14. In one example, the potential candidate regions are M×N (e.g., M=N=8) rectangles, and their top-left positions in order are as follows (assuming the top-left position of the current block of dimension W×H is (0, 0)): (-xStep,0), (0,-yStep), (xStep,-yStep), (-xStep,yStep), (-xStep,-yStep), (-2*xStep,0), (0,-2*yStep), (-2*xStep,2*yStep), (2*xStep,-2*yStep), (-2*xStep,yStep), (xStep,-2*yStep), (-2*xStep,-yStep), (-xStep,-2*yStep), (-2*xStep,-2*yStep), (-xStep / 2,0), (0,-yStep / 2), (xStep / 2,-yStep / 2), (-xStep / 2,yStep / 2), (-xStep / 2,-yStep / 2), Where xStep and yStep are integers. a. The order of inspection can be changed. b. In one example, xStep=Max(W, K1), yStep=Max(H, K2), where K1 and K2 are integers, for example, K1=K2=16. 15. In one example, whether and / or how to apply NA-CCP can be signaled from the encoder to the decoder. a. Alternatively, whether and / or how to apply NA-CCP can be derived at the encoder and decoder based on coded / decoded information without signaling. b. “How to apply NA-CCP” may include: i. Which CCP (such as CCLM or CCCM) model is derived via NA-CCP; ii. Shape / size / position of (potential) candidate regions; iii. Size of the region list; iv. the number of (potential) candidate regions; v. The color component to which NA-CCP is applied. c. “Encoded / decoded information” may include: i. The mode of the current block; ii. Patterns of neighboring blocks; iii. The pattern of the luminance blocks in the same region as the current block; iv. The pattern of luminance blocks in the same region as the neighboring blocks; v.QP; vi. Strip / image type; vii. Image width / height; viii.Block width / height; ix. Reconstructed sample points. 16. In one example, the CCP coding information of the spatial domain neighboring blocks or the temporal domain neighboring blocks can be used by the current block. a. For example, the spatially adjacent blocks may be adjacent to or not adjacent to the current block. b. For example, CCP codec information may include: i. Type of CCP method, such as CCLM or CCCM or GLM or GLM with luma or GL-CCCM or CCCM using non-subsampled luma samples. 1) In one example, GLM methods using different downsampling filters can be considered as different types. 2) In one example, GLM methods with luminance using different downsampling filters can be considered different types. 3) In one example, the types may be CCCM, CCLM, 4 types of GLM using different downsampling filters, 4 types of GLM with luma using different downsampling filters, GL-CCCM, and CCCM using non-downsampled luma samples. 4) "Not using CCP codec" (denoted as NonCCP) can also be considered as a type. ii. Position (x, y). iii. Number of models. 1) For example, the number of models can be 1 or 2. 2) In one example, the number of models can be considered as part of the CCP type. For example, CCLM and MM-CCLM can be considered as two types. iv. At least one threshold used to classify samples for different models. 1) Thresholding can be used only when the number of models is at least 2. v. At least one luma sample value offset. 1) Luma sample value offsets can be added to or subtracted from luma samples (which may be downsampled) when used to derive chroma prediction values. 2) Luma sample value offset can be used only for certain types, such as CCCM, GLM with luma, GL-CCCM, and CCCM using non-subsampled luma samples. vi. At least one chroma sample value offset. 1) Chroma sample value offsets can be added to or subtracted from the chroma prediction values derived from the CCP model to generate the final prediction. vii. At least one model for at least one chroma component. 1) For example, it may include different models for the Cb component and the Cr component. 2) For example, the number of models for each component may be included as part of the information. 3) The model can be represented by a CCLM model form, a CCCM model form, a GLM model form, a GLM model form with luma, a GL-CCCM model form, or a CCCM model form using non-subsampled luma samples. viii. At least one sample point position shift expressed as (dX, dY). 1) Chroma sample position displacements can be added to or subtracted from the sample position (x, y) when used to derive the chroma prediction value. 2) Chroma sample position shifting can only be used for certain types, such as GL-CCCM. c. For example, CCP codec information can be stored after the chroma block is encoded / decoded. i. In one example, the CCP codec information can be stored only when the chroma block is coded in CCP mode. 1) In one example, if the chroma block is coded in at least one CCP mode, such as coded in a fusion of chroma intra prediction modes, CCP codec information may be stored. a) The stored type may be set as the CCP type used in fusion of chroma intra prediction modes. ii. In one example, CCP codec information can be stored for any chroma block. 1) If the chroma block is not coded in CCP mode, the type is stored as "NonCCP". iii. If the chroma block is encoded and decoded in CCP mode, the type of information can be stored according to the encoding and decoding mode. 1) If the mode is CCCM or CCCM-T or CCCM-L or MM-CCCM or MM-CCCM-T or MM-CCCM-L, the type is set to "CCCM". 2) If the mode is CCLM or CCLM-T or CCLM-L or MM-CCLM or MM-CCLM-T or MM-CCLM-L, the type is set to "CCLM". 3) If the mode is CCLM or CCLM-T or CCLM-L or MM-CCLM or MM-CCLM-T or MM-CCLM-L with slope adjustment, the type is set to "CCLM". 4) If the mode is GLM with filter X, the type is set to "GLM with filter X". 5) If the mode is GLM with Luma using filter X, the type is set to "GLM with Luma using filter X". 6) If the mode is GL-CCCM, the type is set to "GL-CCCM". 7) If the mode is to use non-subsampled CCCM, the type is set to "use non-subsampled CCCM". 8) If the mode is a fusion of chroma intra prediction modes, the type is set to "CCLM". iv. The number of models can be stored as the number of models of the chroma block. 1) For example, if the mode is MM-CCLM or MM-CCLM-T or MM-CCLM-L or MM-CCLM or MM-CCLM-T or MM-CCLM-L or any other multi-model CCP mode (such as GLM or GL-CCCM or CCCM using non-subsampled luma samples with multiple models), the number of models is set to 2. v. Information such as thresholds, luma / chroma sample value offsets, and sample position displacements may be stored as information used by the chroma block. vi. The CCP model of a component can be stored as the model used by the chroma block. 1) The model can be derived by any CCP method, such as CCLM or CCLM-T or CCLM-L or MM-CCLM or MM-CCLM-T or MM-CCLM-L or CCCM or CCCM-T or CCCM-L or MM-CCCM or MM-CCCM-T or MM-CCCM-L or GLM using different downsampling filters or GLM with luma using different downsampling filters or GL-CCCM or CCCM using non-subsampled luma samples. 2) The stored model may be the one that is ultimately applied, such as a model after being modified by slope adjustment. d. For example, CCP codec information can be stored at an M×N granularity. i. For example, M=N=2. ii. For example, CCP codec information of a specific chroma block covered by, covered by, or overlapped with an M×N area may be stored in the M×N area. 1) For example, CCP codec information of a first coded / decoded block having CCP information covered by or with or overlapping an M×N area may be stored. 2) For example, the CCP codec information of the last coded / decoded block having CCP information covered by or with or overlapping an M×N area may be stored. 3) For example, CCP codec information of an encoded / decoded block having CCP information covered by or with or overlapping a specific position of an M×N area may be stored. a) The specific position may be the upper left position / lower right position / upper right position / lower left position / center position of the M×N area. 17. In one example, a CCP candidate list can be constructed for chroma blocks. a. In one example, a first syntax element (SE) may be signaled to indicate whether a CCP candidate in the list is applied to the current chroma block. (This may be expressed as "block is coded in CCP candidate list mode") i. For example, SE can be a logo. ii. For example, SE can be encoded and decoded by context. b. For example, the first SE may be signaled in a conditional manner. i. For example, the first SE may be signaled only when CCP is applied. ii. For example, the first SE may be signaled only when CCP is applied and a specific mode is applied. 1) The specific mode may be CCLM. 2) The specific mode may be CCCM. c. In one example, a second syntax element (SE) may be signaled to indicate which CCP candidate is applied. i. For example, SE can be an index. ii. For example, SE can be binarized into a truncated unary code. 1) For example, the maximum value of SE may be S-1, where S is the maximum size of the candidate list. iii. For example, the first binary bit of SE can be encoded and decoded by the context. d. For example, the second SE may be signaled in a conditional manner. i. For example, the second SE may be signaled only when the first SE indicates that a CCP candidate in the list is applied. e. In one example, whether the CCP candidate list mode is applicable may be signaled in the VPS / DPS / SPS / PPS / picture header / slice header / etc. f. In one example, the maximum size / length of the CCP candidate list may be signaled in the VPS / DPS / SPS / PPS / picture header / slice header / etc. 18. In one example, the CCP candidate list may include at least one CCP candidate stored in a spatially neighboring block, which may be adjacent to or not adjacent to the current block (assuming that the upper left position of the current block is (Xt, Yt), and the width and height of the current block are W and H, respectively). a. In one example, a set of locations are checked to find stored CCP information. i. For example, if the type of stored CCP information associated with a location is NonCCP, the location is skipped. 1) Alternatively, if the type of stored CCP information associated with the location is NonCCP, then the location is placed in a backup location list. ii. For example, if the type of the stored CCP information associated with the location is not NonCCP, then the stored CCP information is attempted to be appended to the list. b. In one example, a set of positions (Xi, Yi) to be checked in sequence can be derived from positions close to the current block to positions far away from the current block. i. For example, the position can be checked in a cycle-by-cycle manner. For one cycle, multiple positions are checked and the next cycle is executed. ii. In one example, the position to be checked in the cycle is: (Xt-NDHor-1,Yt+H+NDVer-1),(Xt+W+NDHor-1,Yt-NDVer-1),(Xt+(W>>1),Yt-NDVer-1),(Xt-NDHor-1,Yt+(H>>1)),(Xt-NDHor-1,Yt-NDVer-1) NDHor and NDVer are different for different cycles. iii. In one example, the position to be checked for period k is derived as: NDHor=(k==0?W / 2:W*k); NDVer=(k==0?H / 2:H*k) iv. In one example, the locations to be checked may be different for different cycles. c. In one example, the set of positions (Xi, Yi) to be checked can be the same as the set of positions checked when building the Merge list. d. In one example, the set of positions (Xi, Yi) to be checked may be the same as the set of positions checked when constructing the sub-block based Merge list. 19. In one example, when attempting to place stored CCP information as a candidate (referred to as a potential candidate) into the CCP candidate list, it may be compared with at least one candidate already in the CCP candidate list. a. In one example, all candidates in the list can be compared to potential candidates. b. In one example, if an existing candidate in the CCP candidate list is the same as or similar to the potential candidate, the potential candidate cannot be placed in the CCP candidate list. c. In one example, two CCP candidates are determined to be not identical if the following conditions are met: i. CCP types are different. ii. The number of models is different. iii. If the CCP has multiple models, the threshold is different. iv. At least one model is different. v. Luma sample offset is different. (Applicable only when type is CCCM, GL-CCCM, GPM, or CCCM using non-subsampled luma samples.) vi. Sample point position displacement is different. (Only applicable when the type is GL-CCCM) 20. In one example, when a CCP candidate in the list is used to generate a prediction for the current block, the CCP will be executed according to the CCP information. a. CCCM, CCLM, 4 types of GLM using different downsampling filters, 4 types of GLM with luma using different downsampling filters, GL-CCCM, and CCCM using non-downsampled luma samples can be applied to the current block based on the candidate CCP type. b. Based on the number of candidate models and the threshold, one model or multiple models with at least one threshold can be used. c. Candidate luma sample value offsets can be added to or subtracted from the luma samples to be placed in the CCP model (which may be downsampled). i. This process may be applicable only when the type is CCCM or GL-CCCM or GPM or CCCM using non-subsampled luma samples. d. Sample point position displacement(s) may be added to or subtracted from the position coordinates to be placed in the CCP model. i. This process may be applicable only when the type is GL-CCCM. e. How to obtain the downsampled brightness samples can be based on the CCP type. i. The downsampled luma samples may be obtained following the downsampling method required by the CCP mode corresponding to the type. 21. In an example, the prediction values generated by the CCP candidate may be modified before being used to obtain the reconstructed sample values. a. In one example, an offset D may be added to or subtracted from the predicted value. b. In one example, the offset may be derived based on luma / chroma samples of a template calculated using reconstructed samples neighboring the current block, referred to as a "template." Figure 35 An example of a template is shown. i. In one example, if reconstructed samples on the left side of the current block are available, the template may be composed of the reconstructed samples on the left side of the current block. ii. In one example, if reconstructed samples above the current block are available, the template may be composed of the reconstructed samples above the current block. iii. In one example, if reconstructed samples above / left of the current block are available, the template may be composed of the reconstructed samples above or to the left of the current block. iv. The corresponding luma samples of the template can be downsampled in the same way as the luma samples inside the current block. c. In one example, if there are N models required by the CCP type (such as two models), then for the N models, a representation can be derived as {D 0 ,…,D N-1}N offsets. i. Offset D i can be added to or subtracted from the predicted values generated by model i. d. In one example, a CCP method indicated by the type of CCP candidate may be applied to the template. i. For example, for the kth sample point of the template, S k =R k -P k is calculated, where R k and P k represent the reconstructed sample value of the kth sample point and the predicted value using CCP respectively. 1) For example, D is calculated as {S k}average value. 2) For example, suppose S kIf the number of is M, then D is calculated as D=sign(sup)×((|sum|+off)>>W), where and ii. For example, for the kth sample point of model i using the template, S i k =R i k -P i k is calculated, where R i k and P i k represent the reconstructed sample value of the kth sample using model i and the predicted value using CCP, respectively. 1) For example, D i is calculated as {S i k}average value. 2) For example, suppose S i k If the number is M, then D is calculated as D i =sign(sum)×((|sum|+off)>>W), where and iii. In one example, the division operation is not used to calculate D or D i . 1) For example, a lookup table can be used to calculate D or D i . e. For example, only certain types of CCPs can apply modifications, such as CCLM and CCCM with multiple models. i. For example, the types of CCLM, CCLM with multiple models, CCCM with multiple models, and GPM may apply modifications. 22. In one example, candidates with type "non-adjacent" may be placed in a candidate list. a. Information includes location (x, y). b. If such a candidate is used to predict the current block, then the CCP model(s) can be derived using the samples referenced by (x, y) as stated in items 1 to 15. c. In one example, the locations stored in the backup location list disclosed in item 18 may be checked to place valid locations into the candidate list. 23. In one example, if the number of candidates in the list is M and M=D+1, the construction of the candidate list may be terminated, where D is the index indicating the selected candidate. 24. In one example, if all possible potential candidates are checked and the size of the candidate list is less than S, where S is the maximum number of candidates, a default candidate may be placed into the list to fill the list. 25. In one example, the CCP candidate list may include at least one candidate obtained from a history-based table. a. The history table can be an online table. b. The history table can be a storage table. c. To construct the CCP candidate list, potential candidates may be examined in order. i. For example, the order can be (1) CCP information stored in spatially adjacent blocks / non-adjacent blocks; (2) CCP candidates with type "non-adjacent"; (3) history-based candidates from the online table; (4) history-based candidates from the stored table; (5) default candidates. ii. For example, the order can be (1) CCP information stored in spatially adjacent blocks; (2) CCP information stored in spatially non-adjacent blocks; (3) CCP candidates with type "non-adjacent"; (4) history-based candidates from the online table; (5) history-based candidates from the stored table; (6) default candidates. iii. For example, the order can be (1) CCP information stored in spatially adjacent blocks; (2) CCP information stored in spatially non-adjacent blocks; (3) history-based candidates from the online table; (4) history-based candidates from the stored table; (5) CCP candidates with type "non-adjacent"; (6) default candidates. iv. For example, the order can be (1) CCP information stored in spatially adjacent blocks; (2) history-based candidates from the online table; (3) CCP information stored in spatially non-adjacent blocks; (4) CCP candidates with type "non-adjacent"; (5) history-based candidates from the stored table; (6) default candidates. v. Any type of candidate in the exemplary order may be removed from it. vi. Any other order of potential candidates of these categories. 26. In one example, if a chroma block is encoded and decoded by using at least one CCP candidate, CCP information of the CCP candidate may be stored. a. The storage method can follow the method disclosed in item 16. 27. In one example, if a chroma block is encoded by using at least one CCP candidate, the CCP information of the CCP candidate may be placed into a history-based table. a. The process for placing CCP information into history-based tables may follow the process described in Section 2.27. General 28. The syntax elements disclosed above can be binarized as flags, fixed-length codes, EG(x) codes, unary codes, truncated unary codes, truncated binary codes, etc. The syntax elements can be signed or unsigned. 29. The syntax elements disclosed above may be encoded or decoded using at least one context model, or the syntax elements may be bypassed for encoding or decoding. 30. The syntax elements disclosed above may be signaled in a conditional manner. a. SE is transmitted via the signal only if the corresponding function is applicable. b. SE is signaled only if the block dimensions (width and / or height) meet the conditions. 31. The syntax elements disclosed above can be transmitted through signals at the block level / sequence level / picture group level / picture level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB or sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 32. Whether and / or how to apply the method disclosed above can be transmitted through signals at block level / sequence level / picture group level / picture level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB or sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 33. Whether and / or how to apply the above disclosed methods may depend on coded information such as block size, color format, single / dual tree partitioning, color components, slice / picture type. 34. The proposed method disclosed in this document can be used in other codecs that require chroma fusion.

[0096] The term "video unit" or "codec unit" or "block" may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, or a TB. In the present disclosure, with respect to a "block coded in mode N", "mode N" here may be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a codec technology (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS, and corresponding variants, etc.). The terms "cross-component prediction" and "cross-component prediction mode" may be used interchangeably. As used herein, the term "cross-component prediction model" may refer to a model used in cross-component prediction or cross-component prediction mode. The terms "multi-model cross-component prediction mode" and "multi-model cross-component prediction" may be used interchangeably.

[0097] Figure 36 1. A flow chart of a method 3600 for video processing according to an embodiment of the present invention is shown. The method 3600 is implemented during conversion between a target video block of a video and a bitstream of the video.

[0098] At block 3610, for conversion between a video unit of a video and a bitstream of the video unit, a prediction value for the video unit is generated based on the cross-component prediction candidates. In some embodiments, the video unit is applied with other codec tools that require chroma fusion.

[0099] At block 3620, the prediction value for the video unit is modified. At block 3630, reconstructed sample values are obtained based on the modified prediction value.

[0100] At block 3640, conversion is performed based on the reconstructed sample values. In some embodiments, conversion may include encoding the video unit into a bitstream. Alternatively or additionally, conversion may include decoding the video unit from the bitstream. In this way, this can improve codec efficiency and codec performance.

[0101] In some embodiments, the offset is added to or subtracted from the predicted value. In some embodiments, the offset is derived based on luma samples and / or chroma samples of the template. In some embodiments, the template is calculated using reconstructed samples adjacent to the current block. Figure 35An example of a template is shown.

[0102] In some embodiments, if reconstructed samples to the left of the current block are available, the template includes reconstructed samples to the left of the current block. In some other embodiments, if reconstructed samples above the current block are available, the template includes reconstructed samples above the current block.

[0103] In some embodiments, if reconstructed samples above or to the left of the current block are available, the template includes reconstructed samples above or to the left of the current block. In some embodiments, the corresponding luma samples of the template are downsampled in the same manner as the luma samples inside the current block.

[0104] In some embodiments, if there are a predetermined number of models required for the cross-component prediction type, then a predetermined number of offsets are derived for the predetermined number of models. In one example, if there are N models required for the CCP type (such as two models), then {D 0 ,…,D N-1}N offsets can be derived for N models.

[0105] In some embodiments, the i-th offset from a predetermined number of offsets is added to or subtracted from the predicted value generated by the i-th model from the predetermined number of models, where i is an integer. For example, the offset D i can be added to or subtracted from the predicted values generated by model i.

[0106] In some embodiments, the cross-component prediction method indicated by the type of the cross-component prediction candidate is applied to a template calculated using reconstructed samples adjacent to the current block. In some embodiments, where for the kth sample of the template, S k Calculated as R k -P k , where S k Indicates the kth increment value, R k represents the reconstructed sample value, and P k Represents the predicted value of the kth sample point. For example, for the kth sample point of the template, calculate S k =R k -P k , where R k and P k denote the reconstructed sample values respectively, and the predicted value with CCP of the kth sample is calculated.

[0107] In some embodiments, the offset is determined as the average of the incremental values of the template. For example, D is calculated as {S k}average value.

[0108] In some embodiments, the offset is determined as: D = sign(sum) x ((|sum| + off) >> W), and where D represents the offset, M represents the number of sample points of the template, and S k Indicates the kth increment value.

[0109] In some embodiments, for the kth sample point of the i-th model using the template, S i k Calculated as R i k -P i k , where S i k Indicates the kth increment value using the i-th model, R i k represents the reconstructed sample value of the kth sample using the i-th model, and P i k represents the predicted value of the kth sample point using the i-th model. For example, for the kth sample point of model i using the template, calculate S i k =R i k -P i k , where R i k and P i k denote the reconstructed sample values respectively, and the predicted value with CCP of the k-th sample using model i is calculated.

[0110] In some embodiments, using the i-th model, the offset for the i-th model is determined as the average of the incremental values. For example, D i is calculated as {S i k In some other embodiments, the offset for the i-th model is determined as: D i =sign(sum)×((|sum|+off)>>W), and where D i represents the offset, M represents the number of samples using the i-th model, and Indicates the kth sample point using the i-th model.

[0111] In some embodiments, the division operation is not used to calculate the offset for the template or the offset for the i-th model (ie, to calculate D or D i). For example, the lookup table is used to calculate the offset for the template or the offset for the i-th model.

[0112] In some embodiments, modification of the predicted values is applied across target types of component predictions. For example, the target types may include one or more of: CCLM or CMMM with multiple models.

[0113] In some embodiments, candidates of a type having non-adjacent cross-component prediction information are placed in a cross-component prediction candidate list. In some embodiments, the cross-component prediction information includes a position. In some embodiments, if the candidate is used to predict the current block, a cross-component prediction model is derived using samples related to the position. For example, if such a candidate is used to predict the current block, (multiple) CCP models can be derived using samples related to (x, y), as required by items 1 to 15. In some embodiments, the positions stored in the backup position list are checked to place valid candidates in the cross-component prediction candidate list.

[0114] In some embodiments, if the number of candidates in the cross-component prediction candidate list is M, the construction of the cross-component prediction candidate list is terminated, where M=D+1, and D represents an index indicating the selected candidate. In some embodiments, if all possible potential candidates are checked and the size of the cross-component prediction candidate list is less than a threshold, a default candidate is placed in the cross-component prediction candidate list to enrich the cross-component prediction candidate list, where the threshold is equal to the maximum number of candidates.

[0115] In some embodiments, the cross-component prediction candidate list includes at least one candidate extracted from a history-based table. In some embodiments, the history-based table is an online table. Alternatively, the history-based table is a stored table.

[0116] In some embodiments, to construct a list of cross-component prediction candidates, potential candidates are checked in order. For example, the order is as follows: (1) cross-component prediction information stored in spatially adjacent or non-adjacent blocks; (2) cross-component prediction candidates of non-adjacent types; (3) history-based candidates from an online table; (4) history-based candidates from a storage table; (5) default candidates. In some other embodiments, the order is as follows: (1) cross-component prediction information stored in spatially adjacent blocks; (2) cross-component prediction information stored in spatially non-adjacent blocks; (3) cross-component prediction candidates of non-adjacent types; (4) history-based candidates from an online table; (5) history-based candidates from a storage table; (6) default candidates. In some embodiments, the order is as follows: (1) cross-component prediction information stored in spatially adjacent blocks; (2) cross-component prediction information stored in spatially non-adjacent blocks; (3) History-based candidates from the online table; (4) History-based candidates from the stored table; (5) Cross-component prediction candidates of non-adjacent types; (6) Default candidates. In some other embodiments, the order is as follows: (1) Cross-component prediction information stored in spatially adjacent blocks; (2) History-based candidates from the online table; (3) Cross-component prediction information stored in spatially non-adjacent blocks; (4) Cross-component prediction candidates of non-adjacent types; (5) History-based candidates from the storage table; (6) Default candidates.

[0117] In some embodiments, the types of candidates are removed from the order. In some other embodiments, potential candidates of these categories may be checked in any other order.

[0118] In some embodiments, if a chroma block is encoded by using at least one cross-component prediction candidate, cross-component prediction information of the cross-component prediction candidate is stored. For example, the storage method may follow the method in item 16.

[0119] In some embodiments, the cross-component prediction codec information is stored after the chroma block is encoded or decoded. In some embodiments, the cross-component prediction codec information is stored if the chroma block is encoded or decoded in a cross-component prediction mode. In some embodiments, the cross-component prediction codec information is stored if the chroma block is encoded or decoded in at least one cross-component prediction mode. In some embodiments, the cross-component prediction codec information is stored if the chroma block is encoded or decoded in a fusion of chroma intra prediction modes. In some embodiments, the stored type is set to the cross-component prediction type used in the fusion of chroma intra prediction modes.

[0120] In some embodiments, cross-component prediction codec information is stored for chroma blocks regardless of the codec mode of the chroma blocks. In one example, CCP codec information can be stored for any chroma block. For example, if the chroma block is not coded in cross-component prediction mode, the cross-component prediction type is stored as non-cross-component prediction.

[0121] In some embodiments, if the chroma block is encoded and decoded in cross-component prediction mode, the type of cross-component prediction information is stored according to the encoding and decoding mode. The type of cross-component prediction (CCP) information may refer to the type of data structure used to store CCP information. When a CCP candidate is applied, the type may determine how luma samples are used to predict chroma samples. For example, if the previous block is encoded and decoded in CCLM mode, the type of CCP information may be CCLM. In this case, if the CCP information is used by the current block, the current block may use the stored parameters (such as a and b) to predict chroma samples in the same manner as CCLM. As another example, if the previous block is encoded and decoded in CCLM-L mode, the type of CCP information may also be CCLM. In this case, if the CCP information is used by the current block, the current block may use the stored parameters (such as a and b) to predict chroma samples in the same manner as CCLM.

[0122] In some embodiments, if the codec mode is one of the following: CCCM, CCCM-top (CCCM-T), CCCM-left (CCCM-L), MM-CCCM, MM-CCCM-T, or MM-CCCM-L, then the type is set to CCCM. In some other embodiments, if the codec mode is one of the following: CCLM, CCLM-T, CCLM-L, MM-CCLM, MM-CCLM-T, or MM-CCLM-L, then the type is set to CCLM.

[0123] In some embodiments, if the codec mode is one of the following with slope adjustment: CCLM, CCLM-T, CCLM-L, MM-CCLM, MM-CCLM-T, or MM-CCLM-L, then the type is set to CCLM. In some other embodiments, if the codec mode is GLM with filter, then the type is set to GLM with filter.

[0124] In some embodiments, if the codec mode is GLM with luma using filter, then the type is set to GLM with luma using filter. In some other embodiments, if the codec mode is GL-CCCM, then the type is set to GL-CCCM.

[0125] In some embodiments, if the codec mode is CCCM with non-subsampling, the type is set to CCCM with non-subsampling. In some other embodiments, if the codec mode is a fusion of chroma intra prediction modes, the type is set to CCLM.

[0126] In some embodiments, the number of models is stored as the number of models for the chroma block. In some embodiments, if the codec mode is one of: MM-CCLM, MM-CCLM-T, MM-CCLM-L, MM-CCLM, MM-CCLM-T, MM-CCLM-L, or other multi-model cross-component prediction mode, the number of models is set to 2.

[0127] In some embodiments, at least one of the following cross-component prediction codec information is stored as the cross-component prediction codec information used by the chroma block: a threshold for classifying samples for different models, a luma sample value offset, a chroma sample value offset, or a sample position displacement.

[0128] In some embodiments, the cross-component prediction model of a component is stored as the model used by the chroma block. In some embodiments, the model is used by the cross-component prediction method. For example, the cross-component prediction method includes at least one of the following: CCLM, CCLM-T, CCLM-L, MM-CCLM, MM-CCLM-T, MM-CCLM-L, CCCM, CCCM-T, CCCM-L, MM-CCCM, MM-CCCM-T, MM-CCCM-L, or GLM using different downsampling filters, GLM with luma using different downsampling filters, or GL-CCCM or CCCM using non-downsampled luma samples. In some embodiments, the stored model is the model that is ultimately applied, such as a model after being modified by slope adjustment.

[0129] In some embodiments, cross-component prediction codec information is stored at an M×N granularity, where M and N are integers, for example, M equals 2 and N equals 2.

[0130] In some embodiments, cross-component prediction codec information of a target chroma block covered by, covered by, or overlapped with an M×N area is stored in the M×N area. For example, cross-component prediction codec information of the first coded / decoded block having cross-component prediction information covered by, covered by, or overlapped with an M×N area is stored. As another example, cross-component prediction codec information of the last coded / decoded block having cross-component prediction information covered by, covered by, or overlapped with an M×N area is stored.

[0131] In some embodiments, cross-component prediction codec information of an encoded / decoded block having cross-component prediction information covered by or overlapped with a target position of an M×N area is stored. For example, the target position is one of the following: upper left position, lower right position, upper right position, lower left position, or center position of the M×N area.

[0132] In some embodiments, if a chroma block is encoded using at least one cross-component prediction candidate, the cross-component prediction information of the cross-component prediction candidate is placed in a history-based table. For example, the process for placing CCP information in the history-based table can follow the process described in Section 2.27.

[0133] In some embodiments, an indication of whether and / or how to generate a cross-component prediction candidate list for a chroma block associated with a video unit is indicated at one of the following: sequence level, group of picture level, picture level, slice level, or slice group level. In some embodiments, an indication of whether and / or how to generate a cross-component prediction candidate list for a chroma block associated with a video unit is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header. In some embodiments, an indication of whether and / or how to generate a cross-component prediction candidate list for a chroma block associated with a video unit is included in one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a codec tree block (CTB), or a codec tree unit (CTU).

[0134] In some embodiments, method 3600 further includes determining whether and / or how to generate a cross-component prediction candidate list for a chroma block associated with the video unit based on codec information of the video unit, where the codec information may include at least one of the following: block size, color format, single-tree and / or dual-tree partitioning, color component, slice type, or picture type.

[0135] In some embodiments, the SE is binarized as one of the following: a flag, a fixed-length code, an EG(x) code, a unary code, a truncated unary code, or a truncated binary code. In some embodiments, the SE is signed or unsigned. In some embodiments, the SE is encoded or decoded using at least one context model, or wherein the SE is bypassed for encoding or decoding. In some embodiments, the SE is signaled in a conditional manner.

[0136] In some embodiments, wherein the SE is signaled only when the corresponding function is applicable, or wherein the SE is signaled only when the dimensions of the video unit meet a condition. In some embodiments, the SE is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or slice group level. In some embodiments, the SE is indicated at one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), or codec tree unit (CTU).

[0137] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: generating a prediction value for a video unit of the video based on a cross-component prediction candidate; modifying the prediction value for the video unit; obtaining a reconstructed sample value based on the modified prediction value; and generating a bitstream based on the reconstructed sample value.

[0138] According to further embodiments of the present disclosure, a method for storing a video bitstream is provided. The method includes: generating a prediction value for a video unit of the video based on a cross-component prediction candidate; modifying the prediction value for the video unit; obtaining a reconstructed sample value based on the modified prediction value; generating a bitstream based on the reconstructed sample value; and storing the bitstream in a non-transitory computer-readable medium.

[0139] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.

[0140] Item 1. A method of video processing, comprising: for conversion between a video unit of a video and a bitstream of the video unit, generating a prediction value for the video unit based on a cross-component prediction candidate; modifying the prediction value for the video unit; obtaining a reconstructed sample value based on the modified prediction value; and performing the conversion based on the reconstructed sample value.

[0141] Clause 2. The method of clause 1, wherein an offset is added to or subtracted from the predicted value.

[0142] Clause 3. The method of clause 2, wherein the offset is derived based on luma samples and / or chroma samples of a template.

[0143] Clause 4. The method of clause 3, wherein the template is calculated using reconstructed samples neighboring the current block.

[0144] Item 5. The method of Item 4, wherein if reconstructed samples to the left of the current block are available, the template includes the reconstructed samples to the left of the current block.

[0145] Item 6. The method of Item 4, wherein if reconstructed samples above the current block are available, the template includes the reconstructed samples above the current block.

[0146] Item 7. The method of Item 4, wherein if reconstructed samples above or to the left of the current block are available, the template includes the reconstructed samples above or to the left of the current block.

[0147] Clause 8. The method of clause 4, wherein corresponding luma samples of the template are downsampled in the same manner as luma samples inside the current block.

[0148] Clause 9. The method of clause 1, wherein if there are a predetermined number of models required for the cross-component prediction type, a predetermined number of offsets are derived for the predetermined number of models.

[0149] Item 10. The method of Item 9, wherein an i-th offset from the predetermined number of offsets is added to or subtracted from the predicted value generated by an i-th model from the predetermined number of models, where i is an integer.

[0150] Item 11. The method of Item 1, wherein the cross-component prediction method indicated by the type of the cross-component prediction candidate is applied to a template calculated using reconstructed samples neighboring the current block.

[0151] Item 12. The method according to Item 11, wherein for the kth sample point of the template, S k Calculated as R k -P k , where S k Indicates the kth increment value, R k represents the reconstructed sample value, and P k represents the predicted value of the kth sample point.

[0152] Clause 13. The method of clause 12, wherein the offset is determined as an average of the incremental values of the template.

[0153] Item 14. The method of Item 12, wherein the offset is determined as: D = sign(sum) x ((|sum| + off) >> W), and wherein D represents the offset, and M represents the number of samples in the template.

[0154] Item 15. The method according to Item 11, wherein for the k-th sample point of the i-th model using the template, S i k Calculated as R i k -P i k , where S i k Indicates the kth increment value using the i-th model, R i k represents the reconstructed sample value of the kth sample using the i-th model, and P i k represents the predicted value of the kth sample point using the i-th model.

[0155] Clause 16. The method of clause 15, wherein using the i-th model, the offset for the i-th model is determined as an average of the incremental values.

[0156] Item 17. The method of Item 15, wherein the offset for the i-th model is determined as: D i =sign(sum)×((|sum|+off)>>W), and where D i represents the offset, and M represents the number of samples using the i-th model.

[0157] Item 18. The method of Item 11, wherein a division operation is not used to calculate the offset for the template or the offset for the i-th model, or wherein a lookup table is used to calculate the offset for the template or the offset for the i-th model.

[0158] Clause 19. The method of clause 1, wherein the modification of the predicted value is applied across component predicted target types.

[0159] Item 20. The method of Item 1, wherein candidates of a type having non-adjacent cross-component prediction information are placed into a cross-component prediction candidate list.

[0160] Item 21. The method of Item 20, wherein the cross-component prediction information comprises a position.

[0161] Item 22. The method of Item 21, wherein if the candidate is used to predict the current block, a cross-component prediction model is derived using samples related to the position.

[0162] Item 23. The method of Item 20, wherein positions stored in a backup position list are checked to place valid candidates into the cross-component prediction candidate list.

[0163] Item 24. The method according to Item 1, wherein if the number of candidates in the cross-component prediction candidate list is M, the construction of the cross-component prediction candidate list is terminated, where M=D+1, and D represents an index indicating the selection candidate.

[0164] Item 25. A method according to item 1, wherein if all possible potential candidates are checked and the size of the cross-component prediction candidate list is less than a threshold, a default candidate is placed in the cross-component prediction candidate list to enrich the cross-component prediction candidate list, wherein the threshold is equal to the maximum number of candidates.

[0165] Clause 26. The method of clause 1, wherein the cross-component prediction candidate list comprises at least one candidate extracted from a history-based table.

[0166] Clause 27. The method of clause 26, wherein the history-based table is an online table, or wherein the history-based table is a stored table.

[0167] Item 28. The method of Item 26, wherein to construct the cross-component prediction candidate list, potential candidates are examined in sequence.

[0168] Item 29. A method according to Item 28, wherein the order is as follows: (1) cross-component prediction information stored in spatially adjacent or non-adjacent blocks; (2) cross-component prediction candidates of non-adjacent types; (3) history-based candidates from an online table; (4) history-based candidates from a stored table; (5) default candidates.

[0169] Item 30. A method according to Item 28, wherein the order is as follows: (1) cross-component prediction information stored in spatially adjacent blocks; (2) cross-component prediction information stored in spatially non-adjacent blocks; (3) cross-component prediction candidates of non-adjacent types; (4) history-based candidates from an online table; (5) history-based candidates from a stored table; (6) default candidates.

[0170] Item 31. A method according to Item 28, wherein the order is as follows: (1) cross-component prediction information stored in spatially adjacent blocks; (2) cross-component prediction information stored in spatially non-adjacent blocks; (3) history-based candidates from an online table; (4) history-based candidates from a stored table; (5) cross-component prediction candidates of non-adjacent types; (6) default candidates.

[0171] Item 32. A method according to Item 28, wherein the order is as follows: (1) cross-component prediction information stored in spatially adjacent blocks; (2) history-based candidates from an online table; (3) cross-component prediction information stored in spatially non-adjacent blocks; (4) cross-component prediction candidates of non-adjacent types; (5) history-based candidates from a stored table; and (6) default candidates.

[0172] Clause 33. The method of clause 28, wherein candidate types are removed from the order.

[0173] Item 34. The method of Item 1, wherein if a chroma block is encoded by using at least one cross-component prediction candidate, cross-component prediction information of the cross-component prediction candidate is stored.

[0174] Item 35. The method of Item 1, wherein if a chroma block is encoded by using at least one cross-component prediction candidate, cross-component prediction information of the cross-component prediction candidate is placed in a history-based table.

[0175] Item 36. A method according to any one of Items 1 to 35, wherein an indication of whether and / or how to generate the cross-component prediction candidate list for the chroma block associated with the video unit is indicated at one of: sequence level, picture group level, picture level, slice level, or slice group level.

[0176] Item 37. A method according to any one of Items 1 to 35, wherein an indication of whether and / or how to generate the cross-component prediction candidate list for the chroma block associated with the video unit is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

[0177] Item 38. A method according to any one of Items 1 to 35, wherein the indication of whether and / or how to generate the cross-component prediction candidate list for the chroma block associated with the video unit is included in one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a codec tree block (CTB), or a codec tree unit (CTU).

[0178] Item 39. The method according to any one of Items 1 to 35 further includes: determining whether and / or how to generate the cross-component prediction candidate list of the chroma block associated with the video unit based on the codec information of the video unit, the codec information including at least one of the following: block size, color format, single tree and / or dual tree partitioning, color component, slice type, or picture type.

[0179] Item 40. A method according to any one of items 1 to 35, wherein the SE is binarized into one of the following: a flag, a fixed length code, an EG(x) code, a unary code, a truncated unary code or a truncated binary code.

[0180] Item 41. The method of Item 40, wherein the SE is signed or unsigned.

[0181] Item 42. A method according to any one of items 1 to 41, wherein the SE is encoded or decoded using at least one context model, or wherein the SE is bypass encoded or decoded.

[0182] Item 43. A method according to any one of Items 1 to 41, wherein the SE is signaled in a conditional manner.

[0183] Item 44. The method of Item 43, wherein the SE is transmitted via a signal only when a corresponding function is applicable, or wherein the SE is transmitted via a signal only when a dimension of the video unit satisfies a condition.

[0184] Item 45. The method of any one of Items 1 to 44, wherein the SE is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.

[0185] Item 46. A method according to any one of Items 1 to 44, wherein the SE is indicated at one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a codec tree block (CTB), or a codec tree unit (CTU).

[0186] Item 47. The method of any one of Items 1 to 46, wherein the converting comprises encoding the video unit into the bitstream.

[0187] Item 48. The method of any one of Items 1 to 46, wherein the converting comprises decoding the video unit from the bitstream.

[0188] Item 49. An apparatus for video processing, comprising a processor and non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 48.

[0189] Item 50. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method of any one of Items 1 to 48.

[0190] Item 51. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: generating a prediction value for a video unit of the video based on a cross-component prediction candidate; modifying the prediction value for the video unit; obtaining a reconstructed sample value based on the modified prediction value; and generating the bitstream based on the reconstructed sample value.

[0191] Item 52. A method for storing a bitstream of a video, comprising: generating a prediction value for a video unit of the video based on a cross-component prediction candidate; modifying the prediction value for the video unit; obtaining a reconstructed sample value based on the modified prediction value; generating the bitstream based on the reconstructed sample value; and storing the bitstream in a non-transitory computer-readable medium. Example device

[0192] Figure 37 A block diagram of a computing device 3700 in which various embodiments of the present disclosure may be implemented is shown. The computing device 3700 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0193] It should be understood that Figure 37 The computing device 3700 shown in FIG is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.

[0194] like Figure 37 As shown, computing device 3700 includes a general computing device 3700. Computing device 3700 may include at least one or more processors or processing units 3710, memory 3720, storage unit 3730, one or more communication units 3740, one or more input devices 3750, and one or more output devices 3760.

[0195] In some embodiments, the computing device 3700 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 3700 can support any type of interface to the user (such as a "wearable" circuit device, etc.).

[0196] The processing unit 3710 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 3720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 3700. The processing unit 3710 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0197] The computing device 3700 typically includes various computer storage media. Such media can be any media accessible by the computing device 3700, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 3720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 3730 can be any removable or non-removable medium and can include machine-readable media, such as memory, a flash drive, a disk or other media that can be used to store information and / or data and can be accessed in the computing device 3700.

[0198] The computing device 3700 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 37 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0199] The communication unit 3740 communicates with another computing device via a communication medium. In addition, the functionality of the components in the computing device 3700 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 3700 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0200] Input device 3750 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 3760 may be one or more of various output devices, such as a display, speaker, printer, etc. Computing device 3700 may also communicate with one or more external devices (not shown) via communication unit 3740, such as storage devices and display devices, one or more devices that enable a user to interact with computing device 3700, or any device that enables computing device 3700 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0201] In some embodiments, some or all components of the computing device 3700 may not be integrated into a single device, but may instead be arranged in a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and may work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to be aware of the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (e.g., the Internet) using appropriate protocols. For example, a cloud computing provider provides an application over a wide area network that can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, the components and functionality described herein may be provided from a conventional server or installed directly or otherwise on a client device.

[0202] In an embodiment of the present disclosure, the computing device 3700 may be used to implement video encoding / decoding. The memory 3720 may include one or more video encoding / decoding modules 3725 having one or more program instructions. These modules can be accessed and executed by the processing unit 3710 to perform the functions of the various embodiments described herein.

[0203] In an example embodiment performing video encoding, an input device 3750 may receive video data as input 3770 to be encoded. The video data may be processed, for example, by a video codec module 3725 to generate an encoded bitstream. The encoded bitstream may be provided as output 3780 via an output device 3760.

[0204] In an example embodiment performing video decoding, an input device 3750 may receive an encoded bitstream as input 3770. The encoded bitstream may be processed, for example, by a video codec module 3725 to generate decoded video data. The decoded video data may be provided as output 3780 via an output device 3760.

[0205] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such changes are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A video processing method, comprising: For conversion between a video unit of a video and a bitstream of the video unit, generating a prediction value for the video unit based on cross-component prediction candidates; modifying the predicted value of the video unit; Obtaining a reconstructed sample value based on the modified predicted value; as well as The conversion is performed based on the reconstructed sample values. The method of claim 1 , wherein an offset is added to or subtracted from the predicted value. 3 . The method of claim 2 , wherein the offset is derived based on luma samples and / or chroma samples of a template. The method of claim 3 , wherein the template is calculated using reconstructed samples adjacent to the current block. 5 . The method of claim 4 , wherein if reconstructed samples on the left side of the current block are available, the template includes the reconstructed samples on the left side of the current block. 6 . The method of claim 4 , wherein if reconstructed samples above the current block are available, the template includes the reconstructed samples above the current block. 7 . The method of claim 4 , wherein if reconstructed samples above or to the left of the current block are available, the template includes the reconstructed samples above or to the left of the current block.

8. The method of claim 4, wherein corresponding luma samples of the template are downsampled in the same manner as luma samples inside the current block. 9 . The method of claim 1 , wherein if there are a predetermined number of models required for the cross-component prediction type, a predetermined number of offsets are derived for the predetermined number of models.

10. The method of claim 9, wherein an i-th offset from the predetermined number of offsets is added to or subtracted from the predicted value generated by an i-th model from the predetermined number of models, where i is an integer. 11 . The method of claim 1 , wherein the cross-component prediction method indicated by the type of the cross-component prediction candidate is applied to a template calculated using reconstructed samples adjacent to a current block.

12. The method according to claim 11, wherein for the k-th sample point of the template, S k Calculated as R k -P k , where S k Indicates the kth increment value, R k represents the reconstructed sample value, and P k represents the predicted value of the kth sample point. The method of claim 12 , wherein the offset is determined as an average of the incremental values of the template.

14. The method of claim 12, wherein the offset is determined as: D = sign(sum) × ((|sum| + off) >> W), and in D represents the offset, and M represents the number of samples in the template.

15. The method according to claim 11, wherein for the k-th sample point of the i-th model using the template, S i k Calculated as R i k -P i k , where S i k Indicates the kth increment value using the i-th model, R i k represents the reconstructed sample value of the kth sample using the i-th model, and P i k represents the predicted value of the k-th sample point using the i-th model. The method of claim 15 , wherein using the i-th model, the offset for the i-th model is determined as the average of the incremental values.

17. The method of claim 15, wherein the offset for the i-th model is determined as: D i =sign(sum)×((|sum|+off)>>W), and in D i represents the offset, and M represents the number of samples using the i-th model.

18. The method of claim 11, wherein a division operation is not used to calculate the offset for the template or the offset for the i-th model, or A lookup table is used to calculate the offset for the template or the offset for the i-th model. The method of claim 1 , wherein the modification of the predicted value is applied across component predicted target types.

20. The method of claim 1, wherein candidates of a type having non-adjacent cross-component prediction information are placed in a cross-component prediction candidate list. The method of claim 20 , wherein the cross-component prediction information comprises a position.

22. The method of claim 21, wherein if the candidate is used to predict the current block, a cross-component prediction model is derived using samples related to the position.

23. The method of claim 20, wherein positions stored in a backup position list are checked to place valid candidates into the cross-component prediction candidate list. 24 . The method of claim 1 , wherein if the number of candidates in the cross-component prediction candidate list is M, construction of the cross-component prediction candidate list is terminated, where M=D+1 and D represents an index indicating a selection candidate.

25. The method of claim 1, wherein if all possible potential candidates are checked and the size of the cross-component prediction candidate list is less than a threshold, a default candidate is placed into the cross-component prediction candidate list to enrich the cross-component prediction candidate list, wherein the threshold is equal to the maximum number of candidates.

26. The method of claim 1, wherein the cross-component prediction candidate list comprises at least one candidate extracted from a history-based table.

27. The method of claim 26, wherein the history-based table is an online table, or The history-based table is a storage table.

28. The method of claim 26, wherein to construct the cross-component prediction candidate list, potential candidates are examined sequentially.

29. The method of claim 28, wherein the order is as follows: (1) cross-component prediction information stored in spatially adjacent or non-adjacent blocks; (2) cross-component prediction candidates of non-adjacent types; (3) history-based candidates from an online table; (4) history-based candidates from a stored table; (5) default candidates.

30. The method of claim 28, wherein the order is as follows: (1) cross-component prediction information stored in spatially adjacent blocks; (2) cross-component prediction information stored in spatially non-adjacent blocks; (3) cross-component prediction candidates of non-adjacent types; (4) history-based candidates from an online table; (5) history-based candidates from a stored table; (6) default candidates.

31. The method of claim 28, wherein the order is as follows: (1) cross-component prediction information stored in spatially adjacent blocks; (2) cross-component prediction information stored in spatially non-adjacent blocks; (3) history-based candidates from an online table; (4) history-based candidates from a stored table; (5) cross-component prediction candidates of non-adjacent types; and (6) default candidates.

32. The method of claim 28, wherein the order is as follows: (1) cross-component prediction information stored in spatially adjacent blocks; (2) history-based candidates from an online table; (3) cross-component prediction information stored in spatially non-adjacent blocks; (4) cross-component prediction candidates of non-adjacent types; (5) history-based candidates from a stored table; and (6) default candidates.

33. The method of claim 28, wherein candidate types are removed from the order. 34 . The method of claim 1 , wherein if a chroma block is encoded by using at least one cross-component prediction candidate, cross-component prediction information of the cross-component prediction candidate is stored. 35 . The method of claim 1 , wherein if a chroma block is encoded by using at least one cross-component prediction candidate, cross-component prediction information of the cross-component prediction candidate is placed in a history-based table.

36. The method of any one of claims 1 to 35, wherein an indication of whether and / or how to generate the cross-component prediction candidate list for the chroma block associated with the video unit is indicated at one of: Sequence level, Picture group level, Picture level, stripe level, or Film group level.

37. The method of any one of claims 1 to 35, wherein the indication of whether and / or how to generate the cross-component prediction candidate list for the chroma block associated with the video unit is indicated in one of the following: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependent Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.

38. The method of any one of claims 1 to 35, wherein an indication of whether and / or how to generate the cross-component prediction candidate list for the chroma block associated with the video unit is included in one of: Prediction Block (PB), Transform Block (TB), Codec Block (CB), Prediction Unit (PU), Transformation Unit (TU), Codec Unit (CU), Codec Tree Block (CTB), or Codec Tree Unit (CTU).

39. The method according to any one of claims 1 to 35, further comprising: Determine whether and / or how to generate the cross-component prediction candidate list for the chroma block associated with the video unit based on codec information of the video unit, the codec information including at least one of the following: Block size, Color format, Single-tree and / or dual-tree partitioning, Color component, Strip type, or Image type.

40. The method according to any one of claims 1 to 35, wherein the SE is binarized as one of the following: a flag, a fixed length code, an EG(x) code, a unary code, a truncated unary code, or a truncated binary code. The method of claim 40 , wherein the SE is signed or unsigned.

42. The method according to any one of claims 1 to 41, wherein the SE is encoded or decoded using at least one context model, or The SE is bypassed.

43. The method according to any one of claims 1 to 41, wherein the SE is signaled in a conditional manner.

44. A method according to claim 43, wherein the SE is signalled only if a corresponding function is applicable, or The SE is transmitted via a signal only when the dimension of the video unit meets a condition.

45. The method of any one of claims 1 to 44, wherein the SE is indicated at one of: Sequence level, Picture group level, Picture level, stripe level, or Film group level.

46. The method of any one of claims 1 to 44, wherein the SE is indicated at one of: Prediction Block (PB), Transform Block (TB), Codec Block (CB), Prediction Unit (PU), Transformation Unit (TU), Codec Unit (CU), Codec Tree Block (CTB), or Codec Tree Unit (CTU).

47. The method of any one of claims 1 to 46, wherein the converting comprises encoding the video unit into the bitstream.

48. The method of any one of claims 1 to 46, wherein the converting comprises decoding the video unit from the bitstream.

49. An apparatus for video processing, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 48.

50. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to execute the method according to any one of claims 1 to 48.

51. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: generating a prediction value for a video unit of the video based on the cross-component prediction candidates; modifying the predicted value of the video unit; Obtaining a reconstructed sample value based on the modified predicted value; as well as The bitstream is generated based on the reconstructed sample values.

52. A method for storing a bitstream of a video, comprising: generating a prediction value for a video unit of the video based on the cross-component prediction candidates; modifying the predicted value of the video unit; Obtaining a reconstructed sample value based on the modified predicted value; generating the bitstream based on the reconstructed sample values; as well as The bitstream is stored in a non-transitory computer-readable medium.