Method and device for video processing and medium

By using cross-component residual models and cross-component prediction techniques, the problems of insufficient encoding and decoding efficiency and performance in existing video encoding and decoding technologies are solved, and more efficient video processing is achieved.

CN121587022APending Publication Date: 2026-02-27DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480049132.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-24
Filing Date
2024-07-23
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have shortcomings in terms of encoding and decoding efficiency and performance, and need to be further improved.

Method used

Cross-component residual model (CCRM) and cross-component prediction (CCP) are used to improve the encoding and decoding efficiency and performance of video processing. By determining the conversion relationship between video units and bitstreams, conversion is performed using multi-mode CCRM, multi-filter CCRM, or CCRM Merge mode.

Benefits of technology

It improves the efficiency and performance of video encoding and decoding, and enhances encoding and decoding gain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121587022A_ABST
    Figure CN121587022A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises: for a conversion between a video unit of the video and a bitstream of the video, determining that one or more parameters of a cross-component residual model (CCRM) of the video unit are inherited from a previous filter-based codec block, where the CCRM is a filter model comprising a cross-component model or the same component model; and performing the conversion based on the CCRM.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to cross-component models for residual encoding and decoding. Background Technology

[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, conventional video codecs have some undesirable issues. Therefore, there is an overall expectation to further improve the encoding and decoding gain of conventional video codec technologies. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. This method includes: for the conversion between video units and the video bitstream, determining one or more parameters of the cross-component residual model (CCRM) of the video unit inherited from a previous filter-based codec block, where the CCRM is a filter model including a cross-component model or a same-component model; and performing the conversion based on the CCRM. In this manner, it can improve encoding / decoding efficiency and performance.

[0005] In a second aspect, another method for video processing is proposed. This method includes: a conversion between video units and the video bitstream for a given video; determining the final prediction of the video units based on a weighted sum of fusion assumptions, where at least one assumption is based on cross-component prediction (CCP); and performing the conversion based on the final prediction. In this manner, encoding / decoding efficiency and performance can be improved.

[0006] In the third aspect, another method for video processing is proposed. This method includes: for the conversion between video units and the video bitstream, determining one or more cross-component residual models (CCRMs) for the video units, wherein the one or more CCRMs include at least one of the following: multi-mode CCRM (MM-CCRM) mode, multi-filter (MF) CCRM (MF-CCRM) mode, or CCRM Merge mode; and performing the conversion based on the one or more CCRMs. In this way, encoding / decoding efficiency and performance can be improved.

[0007] In a fourth aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform a method according to the first, second, or third aspect of this disclosure.

[0008] In a fifth aspect, a non-transitory computer-readable storage medium is provided. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform a method according to the first, second, or third aspect of this disclosure.

[0009] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining one or more parameters of a cross-component residual model (CCRM) of video units inherited from a previous filter-based codec block, wherein the CCRM is a filter model including a cross-component model or a same-component model; and generating a bitstream of video units based on the CCRM.

[0010] In a seventh aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a final prediction of video units of the video based on a weighted sum of fusion hypotheses, wherein at least one hypothesis is based on a prediction of cross-component prediction (CCP); and generating a bitstream of video units based on the final prediction.

[0011] In an eighth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining one or more cross-component residual models (CCRMs) for video units, wherein the one or more CCRMs include at least one of the following: a multi-mode CCRM (MM-CCRM) model, a multi-filter (MF) CCRM (MF-CCRM) model, or a CCRMMerge model; and generating a bitstream of video units based on the one or more CCRMs.

[0012] In the ninth aspect, a method for storing a bitstream of video is proposed. The method includes: determining one or more parameters of a cross-component residual model (CCRM) for video units inherited from a previous filter-based codec block, wherein the CCRM is a filter model that includes a cross-component model or a same-component model; generating a bitstream of video units based on the CCRM; and storing the bitstream in a non-transitory computer-readable recording medium.

[0013] In a tenth aspect, a method for storing a bitstream of video is proposed. The method includes: determining a final prediction of video units based on a weighted sum of fusion hypotheses, wherein at least one hypothesis is based on a prediction of cross-component prediction (CCP); generating a bitstream of video units based on the final prediction; and storing the bitstream in a non-transitory computer-readable recording medium.

[0014] In the eleventh aspect, a method for storing a bitstream of video is proposed. The method includes: determining one or more cross-component residual models (CCRMs) for video units, wherein the one or more CCRMs include at least one of the following: a multi-mode CCRM (MM-CCRM) model, a multi-filter (MF) CCRM (MF-CCRM) model, or a CCRM Merge model; generating a bitstream of the video units based on the one or more CCRMs; and storing the bitstream in a non-transitory computer-readable recording medium.

[0015] The present invention is provided to present, in a simplified form, the selection of concepts further described below in the detailed description. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0016] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0017] Figure 1 A block diagram illustrating an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram illustrating a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram illustrating an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 The diagram illustrates the effect of the slope adjustment parameter "u", with the model created using the current CCLM shown on the left and the updated model as proposed shown on the right. Figure 5 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown. Figure 6 The neighbor reconstructed sample points used for the DIMD chromaticity mode are shown; Figure 7 The intra-frame template matching search area used is shown; Figure 8 The use of the IntraTMP block vector for IBC blocks is shown; Figure 9 The method for dividing angle patterns is shown; Figure 10 An expanded list of MRL candidates is shown; Figure 11 An illustration of the template area is shown; Figure 12 The spatial portion of the convolution filter is shown; Figure 13 The reference region (with its padding) used to derive the filter coefficients is shown. Figure 14 Four Sobel-based gradient styles for GLM are shown; Figure 15 The spatial GPM candidates are shown; Figure 16 The GPM template is shown; Figure 17 GPM mixing is shown; Figure 18 The possible locations of the candidate regions are shown; Figure 19 The locations of adjacent airspace candidates are shown; Figure 20 The transformation selection process for directional planar modes is illustrated; Figure 21 A luminance block is shown for deriving the direct block vector; Figure 22 The diagram shows three types of reconstructed regions defined, comprising thirteen columns or rows of reconstructed pixels; Figure 23 The diagram illustrates three types of filter shapes defined with fifteen inputs and producing one output. Figure 24 Examples of predictions for different locations within the current block are shown; Figure 25 The proposed method is shown on the decoder; Figure 26 The luminance samples L0, ..., L5 are shown relative to the chrominance sample C (shown in a half-pixel luminance grid); Figure 27 The luminance samples L0, ..., L5 are shown relative to the chromaticity sample C; Figure 28 An example of the current template and reference template involved in CCRM encoding and decoding for the current inter-frame block is shown; Figure 29Examples of the current template and reference template involved in CCRM encoding and decoding for the current IBC block are shown; Figure 30 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 31 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 32 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and Figure 33 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0018] Throughout all the accompanying figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Implementation

[0019] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0020] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0021] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0022] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0024] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0025] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.

[0026] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.

[0027] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0028] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.

[0029] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.

[0030] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0031] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0032] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.

[0033] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2The examples are shown separately.

[0034] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0035] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).

[0036] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0037] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.

[0038] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0039] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images containing multiple reference video blocks in lists 0 and 1, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0040] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0041] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0042] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0043] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0044] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0045] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0046] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.

[0047] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0048] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0049] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0050] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0051] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0052] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.

[0053] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0054] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.

[0055] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.

[0056] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.

[0057] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolated values ​​of sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.

[0058] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some respects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.

[0059] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0060] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0061] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, although some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.

[0062] 1. Brief Overview This disclosure relates to video codec technology. Specifically, it concerns chroma prediction in image / video codecs. It can be applied to existing video codec standards such as HEVC, VVC, etc. It can also be applied to future video codec standards or video codecs.

[0063] 2 Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established in 2015 by VCEG and MPEG. JVET meetings are held quarterly, and the new video codec standard was officially named Multifunctional Video Codec (VVC) at the April 2018 JVET meeting, where the first version of the VVC Test Model (VTM) was also released. The VVC working draft and the VTM test model are updated after each meeting. The VVC project achieved technical completion (FDIS) at a meeting in July 2020.

[0064] 2.1 Intra-frame prediction In intra-frame prediction, the minimum chroma intra-frame prediction unit (SCIPU) constraint in the VVC is removed. Additionally, the VPDU constraint used to reduce CCLM prediction latency is also removed.

[0065] 2.1.1 Multi-Model Learning (MMLM) The VVC-included CCLM is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, reconstructed neighboring samples are classified into two classes using a threshold that serves as the average of the luminance reconstructed neighboring samples. A linear model for each class is derived using the least mean square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model. Slope adjustment is applied to both the cross-component linear model (CCLM) and multi-model LM predictions. This adjustment is a linear function that maps luminance values ​​to chrominance values, tilted relative to a center point determined by the average luminance values ​​of the reference samples.

[0066] 2.1.1.1 Slope Adjustment of CCLM CCLM uses a two-parameter model to map luminance values ​​to chrominance values. The slope parameter "a" and the bias parameter "b" define the mapping as follows:

[0067] The slope parameter "u" is adjusted via signal transmission to update the model to the following form:

[0068] in

[0069] This selection tilts or rotates the mapping function around a point with a brightness value yr. The average value of the reference brightness samples used in model creation is taken as yr to provide meaningful modifications to the model. The image below illustrates this process. Figure 4 The diagram illustrates the effect of the slope adjustment parameter "u". Left: Model created using the current CCLM. Right: Updated model as proposed.

[0070] Figure 4 The effect of the slope adjustment parameter "u" is shown. Left side: Model created using the current CCLM. Right side: Updated model as proposed.

[0071] Implementation The slope adjustment parameter is provided as an integer between -4 and 4 (inclusive) and is transmitted via signal in the bitstream. The unit of the slope adjustment parameter is 1 / 8 of the chroma sample value per luminance sample value (for 10-bit content).

[0072] The CCLM model (“LM_CHROMA_IDX” and “MMLM_CHROMA_IDX”) is adjusted to apply to reference samples used on both the top and left sides of the block, but not to “one-sided” mode. This choice is based on a trade-off between encoding / decoding efficiency and complexity.

[0073] When slope adjustment is applied to a multi-mode CCLM model, both models can be adjusted, so at most two slope updates are transmitted via signal transmission for a single chroma block.

[0074] Encoder solution The proposed encoder method performs a SATD-based search for the optimal slope update for Cr and a similar SATD-based search for Cb. If either parameter results in a non-zero slope adjustment, the combined slope adjustment pair (SATD-based update for Cr, SATD-based update for Cb) is included in the list of RD checks for TU.

[0075] 2.1.2 Gradient PDPC In VVC, PDPC may not be applied in some scenarios due to the unavailability of secondary reference samples. In these cases, gradient-based PDPC extended from the horizontal / vertical mode is applied. The PDPC weights (wT / wL) and nScale parameter, used to determine the decay of PDPC weights relative to the distance from the left / top boundary, are set to the corresponding parameters in the horizontal / vertical mode, respectively. Bilinear interpolation is applied when the secondary reference sample is located at the fractional sample position.

[0076] 2.1.3 Secondary MPM A secondary MPM list was introduced. The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list includes 16 entries. First, a general MPM list with 22 entries is constructed. Then, the first 6 entries from this general MPM list are included in the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the planar mode. The remaining entries consist of intra-modes from the left (L), top (A), bottom left (BL), top right (AR), and top left (AL) neighboring blocks, a band-oriented mode with an offset added from the first two available band-oriented modes from the neighboring blocks, and the default mode.

[0077] If the CU block is vertically oriented, the order of the neighboring blocks is A, L, BL, AR, AL; otherwise, it is L, A, BL, AR, AL. Figure 5 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown.

[0078] First, the PMPM flag is parsed. If it is equal to 1, the PMPM index is parsed to determine which entry in the PMPM list is selected. Otherwise, the SPMPM flag is parsed to determine whether to parse the SMPM index or the remaining patterns.

[0079] 2.1.4 Reference Sample Interpolation and Smoothing for Intra-Frame Prediction The 4-tap cubic interpolation is replaced by a 6-tap cubic interpolation filter, which is used to derive the predicted samples from the reference samples.

[0080] For reference sample filtering, a 6-tap Gaussian filter is applied to larger blocks (W>= 32 and H>= 32), otherwise the existing VVC 4-tap Gaussian interpolation filter is applied. Extended intra-frame reference samples are derived using a 4-tap interpolation filter instead of nearest-neighbor rounding.

[0081] 2.1.5 Decoder-side Intra-Frame Mode Derivation (DIMD) When DIMD is applied, two intra-frame modes are derived from the reconstructed neighboring samples, and these two predictions are combined with the planar mode predictions, where the weights are derived from the gradient. The division operation in the weight derivation is performed using an integerization scheme based on the same lookup table (LUT) used by CCLM. For example, division in orientation calculation.

[0082] The following LUT-based method is used for calculation:

[0083] in DivSigTable

[16] = {0, 7, 6, 5,5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}.

[0084] The derived intra-frame modes are included in the main list of most probable intra-frame modes (MPMs), so the DIMD process is performed before the MPM list is built. The main derived intra-frame modes of the DIMD block are stored with the block and used for building the MPM lists of neighboring blocks.

[0085] 2.1.5.1 DIMD Chroma Mode The DIMD chroma mode uses the DIMD derivation method to derive the chroma intra-prediction mode for the current block based on the reconstructed Y, Cb, and Cr samples from neighboring rows and columns in the second nearest neighbor sequence. Specifically, the horizontal and vertical gradients are calculated for each co-located reconstructed luma sample and the reconstructed Cb and Cr samples of the current chroma block to construct the HoG. The intra-prediction mode with the largest histogram amplitude value is then used to perform chroma intra-prediction for the current chroma block. Figure 6 The neighbor reconstructed samples used for the DIMD chromaticity mode are shown.

[0086] When the intra-prediction mode derived from the DIMD chroma mode is the same as the intra-prediction mode derived from the DM mode, the intra-prediction mode with the second largest histogram amplitude value is used as the DIMD chroma mode. A CU level flag is transmitted via signal transmission to indicate whether the proposed DIMD chroma mode is applied.

[0087] 2.1.6 Fusion of Chroma Intra-Frame Prediction Modes The DM mode and the four default modes can be merged with the MMLM_LT mode, as shown below:

[0088] in These are predicted values ​​obtained by applying a non-LM model. These are predicted values ​​obtained by applying the MMLM_LT mode, and This is the final predicted value for the current chroma block. Two weights. and Determined by the intra-prediction mode of adjacent chroma blocks, and It is set to equal to 2. Specifically, when the upper and left adjacent blocks are both encoded and decoded using LM mode, { }={1, 3}; When the blocks above and to the left are both encoded and decoded using non-LM mode, { }={3, 1}; otherwise, { }={2, 2}.

[0089] For syntax design, if a non-LM mode is selected, a flag is transmitted via signaling to indicate whether fusion is applied. This method is only applicable to I-stripes.

[0090] 2.1.7 Intra-frame template matching Intra-Template Matching Prediction (IntraTMP) is a special intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. The encoder then transmits the use of this mode via signal transmission, and the same prediction operation is performed on the decoder side.

[0091] The prediction signal is obtained by comparing the L-shaped causal nearest neighbors of the current block with... Figure 7 Another block in a predefined search region is generated by matching the following: R1: Current CTU R2: Top left CTU R3: Above CTU R4: Left CTU The sum of absolute differences (SAD) is used as the cost function.

[0092] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.

[0093] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is:

[0094] in" "" is a constant that controls the trade-off between gain and complexity. In practice, " "Equals 5." Figure 7 The intra-frame template matching search area used is shown.

[0095] To accelerate the template matching process, the search range of all search regions is downsampled by a factor of 2. This results in a reduction of 4 in the template matching search. After finding the best match, a refinement process is performed. Refinement is accomplished by a second template matching search around the best match with the reduced range. The reduced range is defined as min(BlkW, BlkH) / 2.

[0096] Intra-template matching is enabled for CUs with a width and height of 64 or less. This maximum CU size for intra-template matching is configurable. When DIMD is not used for the current CU, the intra-template matching prediction mode is signaled at the CU level via a dedicated flag.

[0097] 2.1.7.1 Block Vector Candidates Derived for IntraTMP in IBC In this method, the block vector (BV) derived from IntraTMP (Intra-Temporal Matching Prediction) is used for IntraBlock Copy (IBC). The stored IntraTMP BV and IBC BV of neighboring blocks are used as spatial BV candidates in the construction of the IBC candidate list.

[0098] The IntraTMP block vector is stored in the IBC block vector cache, and the current IBC block can use both the IBC BV and the IntraTMP BV of neighboring blocks as BV candidates for the IBC BV candidate list, such as... Figure 8 As shown.

[0099] IntraTMP block vectors are added to the IBC block vector candidate list as spatial domain candidates.

[0100] 2.1.8 Fusion for Template-Based Intra-Frame Mode Derivation (TIMD) For each intra-prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The two intra-prediction modes with the smallest SATD are selected as TIMD modes. These two TIMD modes are weighted and fused after applying the PDPC process, and this weighted intra-prediction is used for encoding and decoding the current CU. Position-dependent intra-prediction combination (PDPC) is included in the derivation of the TIMD modes.

[0101] The costs of the two selected modes are compared with a threshold, and a cost factor of 2 is applied during the test, as shown below:

[0102] If the condition is true, then fusion is applied; otherwise, only mode1 is used.

[0103] The weights of the patterns are calculated from their SATD costs as follows:

[0104] Division operations are performed using an integerization scheme based on the same lookup table (LUT) used by CCLM.

[0105] 2.1.9 Intra-frame prediction fusion This intra-frame prediction method derives predicted samples as a weighted combination of multiple predicted values ​​generated from different reference rows. In this process, multiple intra-frame predicted values ​​are generated and then fused by a weighted average. The process of deriving the predicted values ​​to be used in the fusion process is described below: For the intra-angle prediction mode in the single-mode case including TIMD and DIMD, the proposed method improves upon the representation of... Intra-prediction is derived by weighting the intra-prediction obtained from multiple reference rows, where It is an intra-frame prediction from the default reference line, and This is a prediction from the row above the default reference row. The weights are set to... and .

[0106] For TIMD modes with hybridity, Used in the first mode ( ),and Used in the second mode ( ).

[0107] For DIMD patterns with a mix, the number of predicted values ​​selected for the weighted average is increased from 3 to 6.

[0108] When the intra-frame mode has a non-integer slope (required reference sample interpolation) and the block size is greater than 16, the intra-frame prediction fusion method is applied to the luma block, used together with MRL, and not applied to the ISP-encoded block. In the method studied in subtest a, PDPC is applied to the intra-frame prediction mode using the reference line closest to the current block.

[0109] 2.1.10 Combination of CIIP with TIMD and TM Merge In CIIP mode, prediction samples are generated by weighting the inter-prediction signal using CIIP-TM Merge candidate prediction and the intra-prediction signal using the intra-prediction mode derived using TIMD. This method is only applied to codec blocks with an area of ​​1024 or less.

[0110] The TIMD derivation method was used to derive intra-prediction modes in CIIP. Specifically, the intra-prediction mode with the smallest SATD value in the TIMD mode list was selected and mapped to one of 67 regular intra-prediction modes.

[0111] Furthermore, it is proposed that if the derived intra-prediction mode is an angle mode, the weights (wIntra, wInter) for the two tests should be modified. For near-horizontal mode (2 <= angle mode index < 34), the current block is vertically partitioned; for near-vertical mode (34 <= angle mode index <= 66), the current block is horizontally partitioned.

[0112] For different sub-blocks (wIntra, wInter) such as Figure 9 As shown.

[0113] Table 1. Weights used for modifications to the angle mode

[0114] Using CIIP-TM, a CIIP-TM Merge candidate list is constructed for the CIIP-TM pattern. Merge candidates are refined through template matching. CIIP-TM Merge candidates are also reordered as regular Merge candidates using the ARMC method. The maximum number of CIIP-TM Merge candidates is 2.

[0115] 2.1.11 Extended Multi-Reference Row (MRL) List The MRL list in VVC is expanded to include more reference lines for intra-frame prediction. The expanded reference line list consists of line indices {1, 3, 5, 7, 12}. For Template-Based Intra-Frame Mode Derivation (TIMD), only the first two reference line candidates (i.e., {1, 3}) are used, instead of the complete MRL candidate list. Figure 10 An expanded list of MRL candidates is shown.

[0116] 2.1.12 Template-based multi-reference row intra-frame prediction Template-based Multi-Reference Line Intra-Prediction (TMRL) mode combines reference lines and prediction modes, using template matching to construct a list of candidate combinations. The index of the codec candidate combination list indicates which reference line and prediction mode to use when encoding and decoding the current block. Regular Multi-Reference Line (MRL) for non-TIMD portions is replaced by TMRL mode.

[0117] The TMRL mode expands the reference line candidate list and the intra-prediction mode candidate list. The expanded reference line candidate list is {1, 3, 5, 7, 12}. The restriction on the top CTU line remains unchanged. The size of the intra-prediction mode candidate list is 10. The construction of the intra-prediction mode candidate list is similar to that of MPM, except that planar modes are excluded from the intra-prediction mode candidate list, DC modes are added after the modes of the 5 neighboring PUs and DIMD modes (if they are not included), and has a range from... arrive An incremental angle mode (compared to existing angle modes in the intra-prediction mode candidate list) has been added.

[0118] The TMRL candidate is constructed as follows. There are 5 x 10 = 50 combinations of extended reference lines and allowed intra-prediction modes for the block. Since the extended reference lines start from reference line 1, the region covered by reference line 0 is used for template matching. The template region is calculated between prediction (generated from the 50 combinations) and reconstruction (see...). Figure 11 The SAD cost on the TMRL is calculated. The 20 combinations with the lowest SAD cost are selected in ascending order to form the TMRL candidate list.

[0119] For TMR signaling, instead of directly encoding and decoding the reference line and intra-frame mode, the index of the TMRL candidate list is encoded and decoded to indicate which combination of reference line and prediction mode is used to encode and decode the current block.

[0120] 2.1.13 Convolutional Cross-Component Intra-Frame Prediction Model In this method, a convolutional cross-component model (CCCM) is applied to predict chroma samples from reconstructed luminance samples, similar in spirit to what is done by the current CCLM model. As with CCLM, when chroma downsampling is used, the reconstructed luminance samples are downsampled to match a lower-resolution chroma grid. Similar to CCLM, top, left, or top and left reference samples are used as templates for model derivation.

[0121] In addition, similar to CCLM, there are options for using a single model or a multi-model variant of CCCM. The multi-model variant uses two models: one model is derived for samples above the average luminance reference value, and the other model is derived for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples.

[0122] 2.1.13.1 Convolution Filter The convolutional 7-tap filter consists of a 5-tap spatial component with a sign shape, a nonlinear term, and a bias term. The input of the 5-tap spatial component of the filter consists of the center (C) luminance sample that is co-located with the chrominance sample to be predicted, and its upper / north (N), lower / south (S), left / west (W), and right / east (E) neighbors, as shown below. Figure 12 The spatial portion of the convolution filter is shown.

[0123] The nonlinear term P is expressed as the square of the center luminance sample C and scaled to the range of sample values ​​for the content:

[0124] That is, for 10 bits of content, it is calculated as:

[0125] The bias term B represents the scalar offset between the input and output (similar to the offset term in CCLM) and is set to an intermediate chroma value (512 for 10-bit content).

[0126] The filter output is calculated as the convolution between the filter coefficients ci and the input values, and then clipped to the range of effective chromaticity samples. predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B 2.1.13.2 Calculation of Filter Coefficients The filter coefficients ci are calculated by minimizing the MSE between the predicted chromaticity samples and the reconstructed chromaticity samples in the reference region. Figure 13 The reference region is shown, consisting of six rows of chroma samples above and to the left of the PU. The reference region extends to the right by one PU width and below the PU boundary by one PU height. The region is adjusted to include only available samples. The expansion of the region shown in blue is necessary to support the "edge samples" of the shape spatial filter and is filled in unavailable areas.

[0127] MSE minimization is performed by computing the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chromaticity output. The autocorrelation matrix is ​​decomposed using LDL, and the final filter coefficients are computed using inverse substitution. This process largely follows the computation of ALF filter coefficients in ECM; however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations.

[0128] The autocorrelation matrix is ​​calculated using reconstructed values ​​from luma and chromaticity samples. These samples are full-range (e.g., between 0 and 1023 for 10-bit content), resulting in relatively large values ​​in the autocorrelation matrix. This requires high-bit-depth operations during model parameter calculation. A proposed solution is to remove a fixed offset from the luma and chromaticity samples in each PU for each model. This reduces the magnitude of the values ​​used in model creation and allows for a reduction in the precision required for fixed-point arithmetic. As a result, a 16-bit decimal precision is proposed instead of the 22-bit precision of the original CCCM implementation.

[0129] For simplicity, the reference sample values ​​immediately outside the top-left corner of the PU are used as offsets (offsetLuma, offsetCb, and offsetCr). The sample values ​​used in both model creation and final prediction (i.e., the luminance and chromaticity in the reference region and the luminance in the current PU) are reduced by these fixed values, as follows: C' = C – offsetLuma N' = N – offsetLuma S' = S – offsetLuma E' = E – offsetLuma W' = W – offsetLuma P' = nonLinear(C') B = midValue = 1<<(bitDepth - 1) Furthermore, the chromaticity values ​​are predicted using the following equation, where offsetChroma equals offsetCr for the Cr component and offsetCb for the Cb component: predChromaVal = c0C' + c1N' + c2S' + c3E' + c4W' + c5P' + c6B + offsetChroma.

[0130] To avoid any additional sample-level operations, the luminance offset is removed during luminance reference sample interpolation. For example, this can be done by replacing the rounding term used in luminance reference sample interpolation with an updated offset that includes both the rounding term and offsetLuma. Chroma offset can be removed by subtracting it directly from the reference chroma samples. Alternatively, the effect of chroma offset can be removed from the cross-component vectors, yielding the same result. To add the chroma offset back to the output of the convolution prediction operation, it is added to the bias term of the convolution model.

[0131] The calculation of CCCM model parameters requires division operations. Division operations are not always considered implementation-friendly. Division operations are replaced by multiplication (using scaling factors) and shift operations, where the scaling factor and the number of shifts are calculated based on the denominator, similar to the method used in the calculation of CCLM parameters.

[0132] 2.1.13.3 Gradient Linear Model For the YUV 4:2:0 color format, the Gradient Linear Model (GLM) method can be used to predict chromaticity samples from the luminance sample gradient. Two modes are supported: two-parameter GLM mode and three-parameter GLM mode.

[0133] Compared to CCLM, two-parameter GLM uses the gradient of luminance samples to derive a linear model, rather than downsampled luminance values. Specifically, when applying two-parameter GLM, the input to the CCLM process (i.e., downsampled luminance samples) is... Gradient of brightness sample points Replacement. Other parts of CCLM (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged.

[0134]

[0135] In a three-parameter GLM, chromaticity samples can be predicted based on both the gradient of luminance samples with different parameters and the downsampled luminance values. The model parameters of the three-parameter GLM are derived from neighboring samples in 6 rows and columns using an MSE minimization method based on LDL decomposition, as used in CCCM.

[0136]

[0137] For signaling, when CCLM mode is enabled for the current CU, a flag is transmitted via signaling to indicate whether GLM is enabled for both Cb and Cr components; if GLM is enabled, another flag is transmitted via signaling to indicate which of the two GLM modes is selected, and further, a syntax element is transmitted via signaling to select one of the four gradient filters used for gradient calculation. Four gradient filters are enabled for GLM, such as... Figure 14 As shown.

[0138] 2.1.13.4 Bitstream Signaling The use of this mode is signaled via a PU-level flag encoded and decoded by CABAC. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered a sub-mode of CCLM. That is, the CCCM flag is signaled only when the intra-frame prediction mode is LM_CHROMA.

[0139] 2.1.14 Spatial Geometric Partitioning Model (SGPM) SGPM is an intra-frame mode of inter-frame coding / decoding tools similar to GPM, where two prediction parts are generated from the intra-frame prediction process. In this mode, a candidate list is constructed, where each entry contains a segmentation partition and two intra-frame prediction modes, such as... Figure 15 As shown, 26 segmentation modes and 3 intra-frame prediction modes were used to form a combination. The length of the candidate list was set to 16. The selected candidate indices were transmitted via signaling.

[0140] Reorder the list using a template ( Figure 16 The SAD between the template's prediction and reconstruction was used for sorting. The template size was fixed at 1.

[0141] For each segmentation pattern, the same intra-to-inter-frame GPM list is used to derive the IPM list for each segment. The IPM list size is set to 3. In the list, the TIMD-derived pattern is replaced by two derived patterns with horizontal and vertical orientations.

[0142] The SGPM pattern is applied with a restricted block size, which is: 4 <= width <= 64, 4 <= height <= 64, and width < height. 8. Height < Width 8. Width Height >= 32.

[0143] Adaptive blending has also been used in spatial GPM, where Figure 17 The mixing depth τ shown is derived as follows:

[0144] 2.1.15 Nonlocal cross-component prediction Cross-component prediction (CCP), including CCLM, CCCM, and their variants, is employed by ECM to leverage cross-component correlations. With CCLM or CCCM, training samples are always adjacent to the current block. However, the cross-component relationships of the current block can be more relevant to cross-component relationships in non-local regions.

[0145] A nonlocal cross-component prediction method is proposed to enhance CCP by gaining more advantages from nonlocal regions.

[0146] Method #1: A Non-Adjacent Cross-Component Prediction (NA-CCP) model is proposed. Using the NA-CCP model, samples from regions that are not adjacent to the current block can be used to derive the CCCM model for the current block. A candidate region list with six candidates is constructed by sequentially examining potential 8×8 regions. If an examined region is available, it is added to the candidate region list. The top-left corner position of the potential 8×8 region is pre-determined as {(-xStep, 0), (0, -yStep), (xStep, -yStep), (-xStep, yStep), (-xStep, -yStep), (-2...} xStep, 0), (0, -2) yStep), (-2 xStep, 2 yStep), (2) xStep, -2 yStep), (-2 xStep, yStep), (xStep, -2 yStep), (-2 xStep, -yStep), (-xStep, -2 yStep), (-2 xStep, -2 yStep), (-xStep / 2, 0), (0, -yStep / 2), (xStep / 2, -yStep / 2), (-xStep / 2, yStep / 2), (-xStep / 2, -yStep / 2)}, where xStep = Max(width, 16), yStep = Max(height, 16). Figure 18 Some possible locations of the candidate regions are shown.

[0147] A flag is transmitted via signaling to indicate whether NA-CCP is applied to the chroma block. If NA-CCP is applied, an index is transmitted via signaling to indicate which candidate in the candidate region list was used to derive the CCCM model.

[0148] Method #2: A history-based cross-component prediction (H-CCP) model is proposed. H-CCP is used to maintain H-CCLM and H-CCCM tables, similar to HMVP tables. After decoding a block encoded / decoded using CCLM or CCCM, the corresponding table is updated. In the H-CCP implementation, the size of the H-CCLM or H-CCCM table is 6. If the current block is encoded / decoded in CCLM or CCCM mode, a flag is signaled to indicate whether H-CCP is applied. If H-CCP is used, an index is further signaled to indicate which candidate model in the H-CCLM or H-CCCM table is selected.

[0149] 2.1.16 Cross-component Merge Mode for Chroma Intra-Frame Coding / Decoding Cross-component prediction (CCP) using methods including Cross-component Linear Model (CCLM), Convolutional Cross-component Model (CCCM), and Gradient Linear Model (GLM) is employed by ECM to leverage cross-component correlations. The Cross-component Merge (CCMerge) mode is proposed as a new CCP mode. The cross-component model parameters of the current chroma block encoded using CCMerge can be inherited from neighboring blocks encoded using CCP. Through CCMerge, CCP can be more efficient and has less signaling overhead.

[0150] In CCMerge, the final cross-component model parameters for the current chroma block can be inherited from its spatially adjacent and non-adjacent neighbors or the default model. A list is created that includes CCP models from spatially adjacent and non-adjacent neighbors encoded and decoded in CCLM, MMLM, CCCM, GLM, chroma blending, and CCMerge modes. After including neighboring CCP models, the default model is further included to fill any remaining empty positions in the list. To avoid including redundant CCP models in the list, a deduplication operation is applied. More details are described below. Figure 19 The locations of adjacent airspace candidates are shown.

[0151] Airspace adjacent candidate The positions of adjacent candidates in the airspace are as follows Figure 19 As shown, the airspace candidates are included in the following order: B1 -> A1 -> B0 -> A0 -> B2.

[0152] Airspace not adjacent to neighboring candidates After examining all spatially adjacent neighbors, spatially non-adjacent neighbor candidates are considered. In the current ECM design, two sets of spatially non-adjacent neighbor candidates are obtained in inter-frame merge mode. In the proposed method, the positions and inclusion order of the spatially non-adjacent neighbor candidates from the first set are used.

[0153] CCLM candidates with default scaling parameters After including both spatially adjacent and non-adjacent candidates, if the list is not full, CCLM candidates with default scaling parameters are considered. The default scaling parameters are {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8}, and the offset parameters are derived based on the selected default scaling parameters, the average neighbor reconstructed luminance sample value (Yavg), and the average neighbor reconstructed Cb / Cr sample value (Cavg).

[0154] 2.1.16.1 Candidate Merging Models When merging CCLM candidates, only the scaling parameter is inherited. The offset parameter is derived using the inherited scaling parameter, Yavg, and Cavg.

[0155] When merging MMLM candidates, the scaling parameter and classification threshold are inherited. The offset parameter in each class is derived based on the inherited classification threshold and the Yavg and Cavg in each class. If no neighboring reconstructed samples are available in a class, the offset parameter is directly inherited from the candidate.

[0156] When merging CCCM candidates, all convolution parameters, offsets (i.e., offsetLuma, offsetCb, and offsetCr), and classification thresholds are inherited.

[0157] When merging GLM candidates, if the GLM candidate is a 3-parameter GLM mode, all gradient style indices and model parameters are inherited; otherwise, if the GLM candidate is a 2-parameter GLM mode, the offset parameters are derived by using the inherited scaling parameters, Yavg, and Cavg.

[0158] When merging chroma blending candidates, the derived MMLM parameters are inherited and used as the merging MMLM candidates.

[0159] For a CCMerge block, if its merging candidate mode is CCLM, MMLM, CCCM, or GLM, the merging candidate mode is stored as the propagation mode of the current chroma block; otherwise, if its merging candidate mode is chroma blending, the propagation mode is set to MMLM. How CCP parameters are inherited or derived when merging CCMerge candidates depends on the propagation mode of the CCMerge candidate, as described in the five paragraphs above.

[0160] 2.1.16.2 Signaling Following the `cclm_mode_flag` syntax element, additional flags are transmitted via semaphores to indicate whether CCMerge is used. If CCMerge is used, candidate indices are additionally transmitted via semaphores. Candidate indices transmitted via semaphores are shared for the Cb / Cr color components. Currently, the maximum allowed number of candidates is set to the default value of 6. If the maximum allowed number of candidates is modified to 1, candidate indices do not need to be transmitted via semaphores. Each bit of the candidate index is context-encoded using a separate context.

[0161] 2.1.17 Plane Mode with Direction Two additional planar modes are used, where either horizontal interpolation only or vertical interpolation only is used to obtain the predicted samples.

[0162] For the planar horizontal mode, only horizontal linear interpolation is performed based on the left reference sample and the upper right reference sample to predict the current sample as:

[0163] For the planar vertical mode, only vertical linear interpolation is performed based on the upper reference point and the lower left reference point to predict the current sample point as:

[0164] Transform kernel selection for horizontal and vertical planar modes, such as Figure 20As shown. If the intra-prediction mode of the current block is planar vertical mode, then the horizontal intra-prediction mode is used to derive the transform kernels in the MTS set and LFNST set. Furthermore, if the intra-prediction mode of the current block is planar horizontal mode, then the vertical intra-prediction mode is used to derive the transform kernels in the MTS set and LFNST set.

[0165] 2.1.18 Direct block vectors for chroma blocks Direct block vectors are used for chroma blocks in a dual-tree stripe. When the chroma dual-tree is activated, a flag is transmitted via signaling to indicate whether the chroma blocks are being encoded / decoded using IBC mode. Figure 21 If one of the luma blocks in the five locations shown is encoded or decoded in IBC or intraTMP mode, its block vector is scaled and used as the block vector for the chroma block. Template matching is used to perform the block vector scaling.

[0166] 2.1.19 Intra-frame prediction mode based on extrapolation filter (EFI mode) The proposed intra-frame prediction based on extrapolation filters is processed in two steps. First, extrapolation filter coefficients are obtained from the reconstructed pixels of the neighboring pixels of the current block using a predetermined template. Second, extrapolation generates predicted values ​​position by position within the current block from the top left to the bottom right.

[0167] 2.1.19.1 Searching for the mean, minimum, and maximum values Similar to CCCM mode, the mean should be removed when the input is fed to the EIP filter. The DC mode value of the current block is used as the mean for the EIP prediction. The minimum and maximum values ​​are searched from the reconstructed pixels in the reconstructed region with thirteen columns and thirteen rows.

[0168] 2.1.19.2 Calculation of Filter Coefficients Three types of reconstruction regions and three filter shapes are proposed, such as Figure 22 As shown, the three types of reconstructed regions defined consist of thirteen columns or rows of reconstructed pixels. When the current block is predicted using the proposed EIP pattern, the decoder decodes the relevant syntax elements to determine the selected reconstructed region type and filter shape for the current block.

[0169] Figure 22 The diagram shows three types of reconstructed regions, defined as comprising thirteen columns or rows of reconstructed pixels.

[0170] Figure 23 The diagram illustrates three types of filter shapes with fifteen inputs and one output.

[0171] The selected filter slides across the selected reconstruction region in a one-pixel step to collect input and output samples for EIP. The autocorrelation matrix and cross-correlation vector are constructed while removing the mean from the input and output samples. The EIP coefficients are then obtained using the same method as in CCCM.

[0172] 2.1.19.3 Prediction of the current block EIP mode makes predictions position by position for the current block, such as Figure 24 As shown.

[0173] For a position located at the top left of the current block, the input to the EIP filter is the reconstructed sample.

[0174] For a location located on the boundary of the current block, part of the input to the EIP filter is the reference sample, and part of the input to the EIP filter is the previously predicted sample.

[0175] For other locations in the current block, the input to the EIP filter is the previously predicted samples.

[0176] To reduce prediction error, the searched minimum and maximum values ​​are used to limit the output range of each predicted value.

[0177] It is the predicted value at (x, y) in the current block. These are the minimum and maximum values ​​searched from the thirteen reconstructed columns and rows. It is the derived EIP filter. coefficient, These are the reconstructed or predicted values ​​used for the current location. It is a value calculated using the DC prediction model.

[0178] 2.2 Cross-component residual model (CCRM) for inter-frame prediction 2.2.1 Introduction It is proposed that when blocks use inter-frame prediction or intra-block copy (IBC), a cross-component residual model (CCRM) is applied to predict chrominance samples from reconstructed luminance samples. Figure 25 The decoder side of the method is shown. A cross-component filter is derived using the predicted signals for both luminance and chrominance. The derived filter is applied to the reconstructed luminance signal to produce the final chrominance prediction.

[0179] 2.2.2 Calculation of Convolution Filter and Filter Coefficients The proposed 8-tap filter consists of 6 spatial brightness samples, a nonlinear term, and a bias term. For example... Figure 26 As shown, the spatial luminance samples (L0, ..., L5) are obtained from the luminance grid. The six luminance samples closest to the chromaticity position C are selected without downsampling. The predicted chromaticity values ​​are obtained as follows:

[0180] Where nonlinear is the nonlinear operator of CCCM, and B is the bias.

[0181] The filter coefficients were derived using the division-free Gaussian elimination method of ECM, and the necessary offset was applied to the samples before the filter derivation.

[0182] When a block has fewer than 64 chroma samples, intra-frame reference samples are used as additional input samples in filter derivation. A CCCM design with up to 6 rows and columns of intra-frame reference samples is used.

[0183] A block with 256 or more chromaticity samples is divided into sub-blocks with a maximum of 256 chromaticity samples each. Sub-blocks containing zero luminance residuals are skipped.

[0184] 2.2.3 Bitstream Signaling The use of this mode is transmitted via signaling through the TU level flag, which is encoded and decoded using CABAC. A new CABAC context is included to support this. The CCRM flag is transmitted via signaling only when the TU's luminance Cbf is not zero and the CU's predMode is MODE_INTER or MODE_IBC.

[0185] 3 questions There are several problems with existing video encoding and decoding technologies, and they will be further improved in order to achieve higher encoding and decoding gains.

[0186] 1. Several aspects of the video unit of CCRM encoding and decoding (such as filter terms, model type, and applied block type) can be further improved.

[0187] 2. The CCRM-estimated predictions will compete with the original inter-frame predictions. Ultimately, the residual with the lower cost is chosen. However, the concept of fusion can be incorporated to achieve better results.

[0188] 3. CCRM is applied to inter-blocks as long as the luminance component has a non-zero CBF. The CCRM on / off decision can be further designed.

[0189] 4. Currently, the CCRM model is applied, taking luminance reconstruction as input and estimated chromaticity prediction as output. However, estimated chromaticity predictions can be generated by adding residual blocks estimated by CCRM to chromaticity predictions without CCRM.

[0190] 4 Detailed Solutions The detailed solutions below should be considered as examples for explaining general concepts. These solutions should not be interpreted in a narrow sense. Furthermore, these solutions can be combined in any way.

[0191] The term "video unit" or "code-decoder unit" can refer to a picture, strip, slice, code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, ​​TU, PB, TB.

[0192] The term "block" can refer to code-decode tree block (CTB), code-decode tree unit (CTU), code-decode block (CB), CU, PU, ​​TU, PB, and TB.

[0193] The terms "motion vector" or "block vector" can refer to the vector of horizontal and vertical displacement between the position of a reference block and the position of the current block. The reference block can be a video unit in a reference image within the RPL list. Alternatively, the reference block can be a video unit in the current image.

[0194] The term "LM" can refer to any linear regression-based method, such as CCLM, MMLM, CCCM, GL-CCCM, CCCM without downsampling, GLM, GLM with luminance values, etc. It can also be referred to as "Cross-Component Prediction (CCP)". CCP models can be used for intra-frame prediction, IBC prediction, or inter-frame prediction.

[0195] The term "CCLM" can refer to a single-model LM mode, which can be a single-model CCLM, a single-model CCCM, a single-model GL-CCCM, a single-model CCCM without subsampling, a single-model GLM, a single-model GLM with luminance values, a multi-model CCLM, a MMLM, a multi-model CCCM, a multi-model GL-CCCM, a multi-model CCCM without subsampling, a multi-model GLM, a multi-model GLM with luminance values, etc.

[0196] The term "MMLM" can refer to a multi-model LM mode, which can be multi-model CCLM, MMLM, multi-model CCCM, multi-model GL-CCCM, multi-model CCCM without downsampling, multi-model GLM, multi-model GLM with luminance values, etc.

[0197] The term "MFLM" can refer to a multi-filter LM mode, which can be MF-CCLM, MF-CCCM, MF-GLM, MF-CCRM, MF-CCCM for inter-frame, multi-filter IBC filter, multi-filter intraTMP filter, and / or variants of the mentioned mode, etc.

[0198] The term "CCCM" can refer to regular CCCM mode, GL-CCCM mode, CCCM without downsampling, CCRM, etc.

[0199] The term "GL-CCCM" can refer to a CCCM mode that takes into account the gradient and location of the samples involved.

[0200] The term "CCCM without downsampling" can refer to a CCCM mode that takes into account unsampled luminance samples.

[0201] The term "CCRM" can refer to residual encoding / decoding or derivation based on cross-component models. It can also refer to inter-frame / IBC prediction based on CCCM models (such as inter-frame / IBC CCCM). It can also refer to intra-frame prediction based on CCCM models (such as intra-frame CCCM). It can refer to the generation and application of cross-component models (such as luma-to-chroma prediction). It can also refer to the generation and application of same-component models (such as luma-to-luma prediction).

[0202] In this document, Cross Component Prediction (CCP) can refer to any cross component prediction method, such as any kind of CCLM / CCCM / GLM / GL-CCCM.

[0203] It should be noted that the following terms are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.

[0204] 1) The residuals (and / or predictions) of chroma blocks can be derived based on cross-component models.

[0205] a. For example, the cross-component model can be a specific extrapolation filter (e.g., EIP, etc.).

[0206] b. For example, the cross-component model can be a specific interpolation filter (e.g., GLM, etc.).

[0207] c. For example, the cross-component model can be a specific convolutional filter (CCCM, GL-CCCM, CCCM without downsampling, CCRM, inter-frame CCCM, intra-frame CCCM, etc.).

[0208] d. For example, the cross-component model can be a specific linear filter (e.g., CCLM, MMLM, etc.).

[0209] 2) Cross-component models used for residual coding and decoding (e.g., CCRM) may not contain nonlinear terms.

[0210] a. For example, a cross-component model used for residual encoding and decoding may contain linear terms and / or bias terms, but not nonlinear terms.

[0211] 3) CCRM can be used for intra-frame or IBC blocks.

[0212] a. For example, it can be used for intra-frame or IBC blocks in intra-frame (such as I) stripes.

[0213] b. For example, it can be used in intra-frame or IBC blocks within inter-frame (such as B or P) stripes.

[0214] c. For example, it can also be used for single trees.

[0215] d. For example, it can also be used for two trees.

[0216] e. For example, in a single-tree I-strip, both luminance and chrominance are encoded and decoded by IBC (or IntraTMP). CCRM can be generated based on the reconstructed luminance and chrominance samples within a reference block retrieved / guided by block vectors, and a residual model is applied to estimate the reconstructed values ​​of the chrominance samples in the current block.

[0217] f. For example, in a dual-tree system, luminance is encoded and decoded by IBC (or IntraTMP), while chrominance is encoded and decoded by intra-frame. CCRM can generate reconstructed luminance samples within a reference luminance block retrieved / guided by block vectors, as well as reconstructed chrominance samples that are co-located (e.g., at the same position) within that luminance block, and a residual model is applied to estimate the reconstructed values ​​of the chrominance samples in the current block.

[0218] 4) CCRM can be used for chroma blocks encoded and decoded by DBV.

[0219] a. For example, a reference chroma block and its corresponding luma block can be identified based on the block vector of the chroma block encoded and decoded by DBV. These samples can be used as training samples for computation of the CCRM model.

[0220] b. For example, the derived CCRM model is applied to the reconstructed luminance signal of the DBV chromaticity block to produce the final chromaticity prediction.

[0221] 5) The CCRM model can be generated based on the correlation between the luminance reconstruction values ​​and chrominance reconstruction values ​​from neighboring / non-adjacent samples of the current block.

[0222] a. For example, alternatively, the CCRM model can be generated based on the correlation between the luminance reconstruction values ​​and chrominance reconstruction values ​​in a reference block of a reference image.

[0223] b. For example, alternatively, the CCRM model can be generated based on the correlation between the luminance reconstruction values ​​and chrominance reconstruction values ​​in a reference block in the current image.

[0224] 6) For example, the CCCM used for intra-frame prediction and the CCCM used for inter-frame prediction (e.g., CCRM) can share the same logic.

[0225] a. For example, both can follow the same logic to obtain training samples.

[0226] b. For example, both can follow the same logic to determine the training area.

[0227] 7) CCRM models can be generated based on unsampled luminance samples.

[0228] a. For example, CCRM model coefficients can be solved based on unsampled luminance samples from a reference region used as training samples.

[0229] b. For example, the CCRM model can be applied to chroma blocks, where the chroma prediction of the current chroma block is generated based on the non-downsampled luminance samples of the co-located luminance block.

[0230] 8) More than one CCRM model (e.g., MM-CCRM model, MF-CCRM model, CCRM Merge model) can be generated for blocks.

[0231] a. For example, the training samples of CCRM can be divided into more than one class (e.g., two classes), and each set of samples can contribute to a unique model. In this way, multiple models can be generated, each with its own filter coefficients. Each derived filter is applied to its corresponding set of luminance reconstructed signals to produce a final prediction value for the current chrominance sample belonging to the corresponding class.

[0232] i. For example, according to the multi-model CCRM (e.g., MM-CCRM) mode, training sample pairs of the luminance and chrominance sample pairs of the reference block (e.g., these training samples in the reference frame) can be divided into more than one category.

[0233] ii. For example, alternatively, training sample pairs (e.g., these training samples in the reference frame) of the brightness and chromaticity sample pairs of neighboring samples adjacent / non-adjacent to the reference block can be classified into more than one category, according to a multi-model CCRM (e.g., MM-CCRM) mode.

[0234] iii. For example, alternatively, training sample pairs (e.g., those training samples in the current frame) of the luminance and chrominance sample pairs of neighboring samples that are adjacent / non-adjacent to the current video cell can be classified into more than one class, according to a multi-model CCRM (e.g., MM-CCRM) mode.

[0235] iv. For example, in addition, by following the same criteria (e.g., by a threshold), the luminance samples in the current video unit are divided into more than one group, and for each luminance sample belonging to a category, the corresponding model can be applied to generate the model-estimated chrominance samples belonging to that group.

[0236] b. For example, multiple sets of training samples can be used to derive multiple models.

[0237] i. In one example, there are two groups where the distance between the training sample and the current sample is different.

[0238] c. For example, the threshold used to separate samples into different categories (e.g., category threshold) may depend on the values ​​of samples within or near the training area.

[0239] i. For example, the training region can be a reference block of the current video unit (e.g., these training samples are in the reference frame).

[0240] 1. For example, a reference block can be derived based on a block vector.

[0241] 2. For example, a reference block can be derived based on motion vectors.

[0242] ii. For example, the threshold can be derived based on samples that are adjacent to or not adjacent to the reference block of the current video unit (e.g., these training samples are in the reference frame).

[0243] iii. For example, the threshold can be derived based on samples that are adjacent to or not adjacent to the current video unit (e.g., these training samples are in the current frame).

[0244] iv. For example, the threshold can be derived based on the average / median / intermediate operation of more than one sample point within or adjacent to the training region.

[0245] v. For example, category thresholds can be derived based on unsampled luminance sample values.

[0246] 1. Alternatively, the category threshold can be derived based on the downsampled brightness sample values.

[0247] a. For example, a K-tap (such as K=6) downsampling filter can be used to reduce K surrounding luminance samples to a single downsampled luminance sample value.

[0248] vi. For example, the category threshold can be derived based on the offset removal scheme.

[0249] 1. For example, the offset can be derived based on luminance samples located at fixed positions (such as the upper left or center) within the reference video unit.

[0250] 2. For example, the offset values ​​calculated for category threshold derivation and CCRM model can be the same.

[0251] vii. For example, category thresholds can be derived based on the sub-block level.

[0252] viii. For example, category thresholds can be derived based on CU / PU / TU levels.

[0253] ix. For example, the category threshold can be calculated based on (downsampled or non-downsampled) brightness prediction samples.

[0254] x. For example, category thresholds can be derived based on luminance residual sample values.

[0255] 1. For example, for a second video unit (e.g., a sub-block) that does not have a non-zero residual, the predicted samples of such a video unit may not be included in the calculation of the category threshold for the first video unit.

[0256] a. For example, the second video unit may be a subset of the first video unit.

[0257] b. For example, the second video unit can be equal to the first video unit.

[0258] d. For example, MM-CCRM can be applied at the sub-block level.

[0259] i. For example, the size of the sub-block can be predefined.

[0260] 1. For example, the predefined sub-block size can be 16x16, or 32x32, etc.

[0261] 2. For example, predefined rules can be used to determine the sub-block size of the MM-CCRM for a specific video unit.

[0262] a. For example, the size of a sub-block can be adapted to the block dimensions (width and / or height) of the current video block.

[0263] b. For example, for a sub-block of a video unit encoded and decoded by MM-CCRM, a minimum number of chroma samples can be guaranteed.

[0264] ii. For example, if a video unit is larger than a predefined sub-block size, the video unit can be divided into more than one sub-block and MM-CCRM can be performed.

[0265] iii. For example, at least one sub-block of a video unit may have more than one CCRM model.

[0266] iv. For example, each sub-block (and its associated training region) can have its own category threshold.

[0267] 1. For example, the category threshold for a specific sub-block can be calculated based on the training sample values ​​belonging to that sub-block.

[0268] a. For example, brightness training samples in a reference block can be used to calculate a category threshold.

[0269] v. For example, all sub-blocks (and their associated training regions) can share the same class threshold.

[0270] 1. For example, a category threshold can be calculated and used for all sub-blocks.

[0271] 2. For example, the category threshold for all sub-blocks in the current video unit can be calculated based on the training sample values ​​of the current video unit.

[0272] 3. For example, the category threshold for all applicable sub-blocks in the current video unit can be calculated based on the training sample values ​​of the current video unit.

[0273] a. For example, sub-blocks that do not contain non-zero residuals may not be counted.

[0274] vi. For example, each sub-block of a video unit can have its own training samples, and the training samples of a particular sub-block can be divided into more than one category.

[0275] 1. For example, training samples in the reference video unit of a reference image can be classified based on sub-blocks.

[0276] vii. For example, training samples from the current image can be classified into more than one group, but may not be divided into sub-blocks.

[0277] e. For example, MM-CCRM / MF-CCRM / CCRM Merge can be applied at the TU level (or PU / CU level).

[0278] i. For example, MM-CCRM / MF-CCRM / CCRM Merge can be applied based on TU / CU / PU (e.g., for MM-CCRM applications, TU / CU / PU may not be divided into sub-blocks).

[0279] ii. For example, whether to use a multi-model CCRM / MF-CCRM / CCRM Merge based on TU / PU / CU can be determined at the TU / PU / CU level.

[0280] 1. For example, a video unit (e.g., TU / PU / CU) may choose to use a sub-block-based CCRM (e.g., CCRM / MM-CCRM / MF-CCRM / CCRM Merge) or a TU / PU / CU-based CCRM (e.g., CCRM / MM-CCRM / MF-CCRM / CCRM Merge).

[0281] a. For example, decisions can be made at the TU / PU / CU level.

[0282] f. For example, whether and / or how to apply MM-CCRM (and / or CCRM / MF-CCRM / CCRM Merge) can be deduced based on encoding and decoding information at both the encoder and decoder sides (e.g., not through signal transmission).

[0283] i. In one example, it can be derived on the fly, for example, using information from previously encoded / reconstructed samples.

[0284] ii. For example, the determination of whether to use a sub-block-based CCRM / MM-CCRM / MF-CCRM / CCRM Merge or a TU / CU / PU level CCRM can be based on implicit deduction of codec information (e.g., not through signal transmission).

[0285] iii. For example, the determination of whether to use CCRM based on M1xM2 subblocks or CCRM based on N1xN2 subblocks can be based on the implicit derivation of encoding and decoding information (e.g., without signal transmission).

[0286] 1. For example, M1 = 16 or 8 or 32 or TU / CU / PU.

[0287] 2. For example, M2 = 16 or 8 or 32 or TU / CU / PU.

[0288] 3. For example, N1 = 16 or 8 or 32 or TU / CU / PU.

[0289] 4. For example, N2 = 16 or 8 or 32 or TU / CU / PU.

[0290] 5. For example, M1 != N1 and / or M2 != N2.

[0291] iv. For example, the determination of whether to use a sub-block-based MM-CCRM / MF-CCRM / CCRM Merge or a TU / CU / PU level MM-CCRM / MF-CCRM / CCRM Merge can be based on implicit deduction of codec information (e.g., not through signal transmission).

[0292] v. For example, the determination of whether to use the MM-CCRM / MF-CCRM / CCRM Merge based on M1xM2 sub-blocks or the MM-CCRM / MF-CCRM / CCRM Merge based on N1xN2 sub-blocks can be based on the implicit derivation of the encoding / decoding information (e.g., not through signal transmission).

[0293] 1. For example, M1 = 16 or 8 or 32 or TU / CU / PU.

[0294] 2. For example, M2 = 16 or 8 or 32 or TU / CU / PU.

[0295] 3. For example, N1 = 16 or 8 or 32 or TU / CU / PU.

[0296] 4. For example, N2 = 16 or 8 or 32 or TU / CU / PU.

[0297] 5. For example, M1 != N1 and / or M2 != N2.

[0298] vi. For example, the determination of whether to use SM-CCRM / MF-CCRM / CCRM Merge or MM-CCRM / MF-CCRM / CCRMMerge can be based on the implicit derivation of codec information (e.g., not through signal transmission).

[0299] vii. For example, determination can be based on a cost derived from the decoder.

[0300] 1. For example, the cost of decoder derivation can be calculated based on minimizing the SAD / SATD / SSE / MSE between the model estimated sample values ​​and the true reconstructed sample values, where a sample can refer to at least one training sample among the training samples.

[0301] 2. For example, a method with lower cost can be chosen as the final method to be applied to the current video unit.

[0302] viii. For example, it can be determined that the information is based on a reference image.

[0303] 1. For example, the determination can be based on the POC distance between the current image and its reference image.

[0304] 2. For example, the determination can be based on a reference index.

[0305] g. Alternatively, whether and / or how to apply MM-CCRM (and / or CCRM / MF-CCRM / CCRM Merge) can be signaled in the bitstream.

[0306] i. For example, syntax elements (e.g., flags, indices, etc.) can be signaled conditionally based on whether the current block is coded / decoded by CCRM.

[0307] 1. For example, if a video unit is coded / decoded by CCRM, syntax elements (e.g., flags, indices, etc.) can be further signaled to indicate whether it is MM-CCRM / MF-CCRM / CCRM Merge.

[0308] ii. For example, syntax elements (e.g., flags, indices, etc.) can be signaled to indicate whether it is MM-CCRM based on sub-blocks or MM-CCRM / MF-CCRM / CCRM Merge based on TU / CP / PU.

[0309] iii. For example, syntax elements (e.g., flags, indices, etc.) can be signaled to indicate whether it is CCRM based on sub-blocks or CCRM / MF-CCRM / CCRM Merge based on TU / CP / PU.

[0310] iv. For example, syntax elements can be signaled conditionally based on the block dimension (width and / or height).

[0311] 1. For example, if W*H < T (such as T = 16 or 32), it may not be signaled.

[0312] v. For example, syntax elements can be signaled depending on the residual / coefficients of the current luma block.

[0313] 1. For example, whether to signal a syntax element can be conditional based on whether there is a residual (or non-zero coefficient) in the current luma block.

[0314] 2. For example, whether to signal a syntax element can be conditional based on the distribution / number / value of the residual (or non-zero coefficient) in the current luma block.

[0315] vi. For example, syntax elements can be signaled conditionally based on the prediction method of neighboring blocks.

[0316] 1. For example, it can be based on whether neighboring blocks (such as the left neighbor and / or the upper neighbor) use the CCRM / MM-CCRM / MF-CCRM / CCRM Merge mode.

[0317] vii. For example, the context model of a syntax element can depend on the coding and decoding information of neighboring blocks or the current block.

[0318] 1. For example, the context model can be derived based on whether neighboring blocks (such as the left neighbor and / or the upper neighbor) use the CCRM / MM-CCRM / MF-CCRM / CCRM Merge mode.

[0319] 2. For example, the context model can be derived based on whether the block dimension of the current block meets specific conditions.

[0320] a. For example, if the current block (e.g., TU / PU / CU) is long or wide (e.g., W > a*H, and / or H > b*W, where W and H are the width and height of the current block, and a and b are predefined constants, e.g., a = b = 2), then the specified context model can be used.

[0321] h. For example, block restrictions can be applied to indicate the allowance of the MM-CCRM mode.

[0322] i. In one example, assuming the width and height of the chrominance CU / PU / TU are represented as W and H, then MM-CCRM can be allowed when at least one of the following conditions is satisfied: 1. W [[ID=2​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​ 10.W H < T9, or W H <= T9 ii. In one example, for a block enabled for a specific tool (e.g., enabling affine motion compensation), MM-CCRM can be prohibited.

[0323] i. For example, a video unit encoded / decoded by CCRM can always use multi-model CCRM.

[0324] i. Alternatively, a video unit encoded / decoded by CCRM can use single-model CCRM or multi-model CCRM.

[0325] 9) Chrominance Cb and Cr can share one CCRM.

[0326] a. Alternatively, chrominance Cb and Cr can construct their own CCRM.

[0327] 10) For the filter design of the CCRM model, sample values and / or gradients and / or position information can be considered.

[0328] a. For example, at least one K-tap filter can be used for the CCRM model, which consists of K1 sample terms, K2 gradient terms, K3 positioning / position terms, K4 non-linear terms, K5 bias terms, etc.

[0329] i. For example, K1 = 0 or 1 or 2 or 5 or 6 ii. For example, K2 = 0 or 1 or 2 or 4 iii. For example, K3 = 0 or 1 or 2 or 4 iv. For example, K4 = 0 or 1 or 2 or 4 v. For example, K5 = 0 or 1 vi. For example, K = K1 + K2 + K3 + K4 + K5 vii. For example, the sample terms can be calculated based on the luminance sample values.

[0330] viii. For example, the gradient terms can be calculated based on more than one sample adjacent to a specific luminance sample.

[0331] ix. For example, the positioning / position terms can be calculated based on the horizontal coordinate and / or vertical coordinate of a specific luminance sample, where the coordinates can be relative to the upper left position of a specific reference region.

[0332] x. For example, the non-linear terms can be the square of a specific value (e.g., an intermediate value related to bit depth, such as 512 or 256, or a specific luminance value).

[0333] xi. For example, the non-linear terms can be the square of the gradient value based on a specific gradient term.

[0334] xii. For example, the offset can be subtracted from the terms of the K-tap filter.

[0335] 1. For example, the offset can be derived based on predefined rules such as the value of the training sample in the upper left of the training area, or the average / median value of more than one sample in the training area.

[0336] xiii. For example, the coefficients of a K-tap filter can be solved using a Gaussian elimination solver.

[0337] xiv. For example, the coefficients of a K-tap filter can be solved using the LDL decomposition method.

[0338] i. For example, the coefficients of a K-tap filter can be solved by linear regression.

[0339] ii. For example, the coefficients of a K-tap filter can be solved using linear equations.

[0340] b. For example, more than one filter can be used, and the final prediction can be derived by fusing the filtered outputs of multiple filters together.

[0341] i. For example, the weights that fuse multiple filter values ​​can be solved using a Gaussian elimination solver.

[0342] ii. For example, the weights for fusing multiple filter values ​​can be solved using the LDL decomposition method.

[0343] 11) For example, for a video unit encoded or decoded by CCRM, more than one filter may be allowed, and which filter is ultimately selected may be transmitted through the signal or divided.

[0344] a. For example, syntax elements can be transmitted via signals to indicate which filter (e.g., CCLM or CCCM) is used in CCRM mode.

[0345] b. For example, indicating which filter (e.g., CCLM or CCCM) is used in CCRM mode can be determined based on template costs from both the encoder and decoder.

[0346] c. For example, indicating which filter (e.g., CCLM or CCCM) is used in CCRM mode can be determined based on the cost derived from both the encoder and decoder.

[0347] 12) The filter output can be limited to a certain value.

[0348] a. For example, it can be limited based on the reconstructed values ​​in the training region.

[0349] i. For example, the training region can be derived based on block vectors (or motion vectors).

[0350] ii. For example, the training region can be adjacent to the current block.

[0351] iii. For example, the training region can be the reference region of the current block.

[0352] iv. For example, the filter output can be limited to the minimum and maximum values ​​of the reconstructed (or predicted) luminance sample values ​​in the training region.

[0353] b. For example, it can be limited based on the reconstructed (or predicted) value in the co-located luminance block of the current chroma block.

[0354] i. For example, it can be limited to the minimum and maximum values ​​of the current block brightness reconstruction (or prediction) value.

[0355] c. For example, if the value is outside the valid range, it can be ignored / discarded / not used.

[0356] 13) CCRM parameters can be stored in the cache and used for encoding and decoding of future blocks.

[0357] a. For example, CCRM parameters for video units (e.g., CU, PU, ​​color components, Cb, Cr, etc.) may include model type, model coefficients, whether it is a single model or multiple models, threshold for separating samples into multiple models, etc.

[0358] b. For example, it can be stored in a local cache for encoding and decoding future blocks in the current image.

[0359] c. For example, it can be stored in the temporal domain / picture / frame buffer for encoding and decoding future blocks in a future decoded picture.

[0360] i. For example, the CCRM parameters of the current frame / image can be stored and referenced for the CCP process of future frames / images.

[0361] ii. For example, it can be stored in association with motion and pattern information of the video unit.

[0362] 14) Video blocks can inherit model parameters from previous filter-based codec blocks. In the sub-items below, CCRM can refer to any filter model that includes cross-component models or same-component models.

[0363] a. For example, a cross-component model can refer to a cross-component residual model or cross-component prediction model used within or between frames, where model generation and application are based on the relationship between different color components (such as luminance to chrominance, chrominance Cb to chrominance Cr).

[0364] b. For example, a common component model can refer to an inter-frame / IBC / IBC-LIC / inter-frame-LIC / intraTMP / EIF filter, where model generation and model application are based on the relationship in the common component (such as luminance to luminance S).

[0365] c. For example, video blocks can be encoded and decoded using a type of CCP inheritance pattern.

[0366] d. For example, video blocks can be encoded and decoded using a type of CCP Merge (e.g., CCMerge) mode.

[0367] e. For example, video blocks can be encoded and decoded using a type of filter model inheritance / merge mode. For example, model parameters of previously encoded and decoded blocks using CCRM can be stored in a cache (e.g., local cache, image cache, temporal cache, history-based LUT, etc.).

[0368] f. In one example, parameters could refer to filter information, linear or nonlinear parameters of the model, model index, etc.

[0369] g. For example, at least one syntax element may be signaled at the video unit level (e.g., block level, tu / pu / cu level, etc.) to specify whether and / or how to use CCRM model inheritance mode (e.g., CCRM Merge mode, or CCP Merge mode, or IBC / intraTMP filter Merge mode).

[0370] i. For example, an indicator can be transmitted via signaling at the video unit level to specify whether the current video unit uses the regular CCRM mode or the CCRM model inheritance mode.

[0371] 1. For example, alternatively, based on at least one type of CCRM (e.g., conventional CCRM mode) used for the current video unit, the indicator is conditionally transmitted via signaling.

[0372] 2. For example, a first syntax is transmitted via signal to indicate that the current video unit uses a certain type of CCRM mode, and then a second syntax is further transmitted via signal to indicate which type of CCRM mode is used.

[0373] a. In addition, alternatively, the second syntax may be transmitted via signaling only if at least one available CCRM candidate exists (e.g., at least one valid CCRM Merge candidate exists).

[0374] ii. For example, alternatively, if the CCRM model inheritance pattern is used, another syntax (e.g., index) can be further specified via signaling to indicate which CCRM model candidate is selected to be inherited.

[0375] 1. For example, candidate indexes can be encoded or decoded to indicate CCRM candidates from a candidate list.

[0376] 2. For example, candidate indices can be encoded or decoded using the Rounding Rice (TR) binarization process, or the Rounding Binary (TB) binarization process, or the k-th order Exp-Golomb (EGk) binarization process, or the Fixed Length (FL) binarization process.

[0377] 3. For example, the maximum allowed CCRM candidates (e.g., the maximum length of the candidate list) can be specified in the codec (such as 12 or 8 or 4 or 2, etc.).

[0378] iii. For example, alternatively, an indicator may be transmitted via signaling at the video unit level to specify whether the current video unit uses the CCRM model inheritance mode, and / or which candidate is used for the CCRM model inheritance mode.

[0379] iv. For example, syntax elements can be transmitted via signaling depending on the residual / coefficient of the current luma block.

[0380] 1. For example, whether to transmit a syntax element via signal can be conditional based on the presence of a residual (or non-zero coefficient) in the current luma block.

[0381] 2. For example, whether a signal transmission syntax element can be conditional on the distribution / number / value of residuals (or non-zero coefficients) in the current luma block.

[0382] v. For example, syntax elements can be conditionally transmitted via signaling based on a prediction method of neighboring blocks.

[0383] 1. For example, it can be based on whether the CCRM / CCRMMerge mode is used based on neighboring blocks (such as left nearest and / or top nearest).

[0384] vi. For example, the context model of a syntax element can depend on the encoding / decoding information of neighboring blocks or the current block.

[0385] 1. For example, the context model can be derived based on whether neighboring blocks (such as left nearest and / or top nearest) use the CCRM / MM-CCRM / MF-CCRM / CCRM Merge pattern.

[0386] 2. For example, the context model can be derived based on whether the block dimension of the current block satisfies specific conditions.

[0387] a. For example, if the current block (e.g., TU / PU / CU) is long or wide (e.g., W>a*H, and / or H>b*W, where W and H are the width and height of the current block, and a and b are predefined constants, e.g., a=b=2), then the specified context model can be used.

[0388] h. For example, if the CCRM model inheritance pattern is used, a list of CCRM model candidates can be generated.

[0389] i. For example, the maximum length of the list size can be predefined in the bitstream (e.g., the size is equal to 6, 10, or 12 candidate models).

[0390] 1. In addition, the size of the history table can be predefined (e.g., size equal to 5 or 6).

[0391] ii. For example, CCRM model candidates can be obtained based on previously encoded and decoded CCRM blocks that are spatially adjacent, and / or temporally adjacent, and / or spatially non-adjacent, and / or based on historical CCRM candidates, and / or shifted candidates, and / or default CCRM candidates.

[0392] 1. For example, the candidate insertion order can follow predefined rules, such as spatially adjacent -> temporally adjacent -> spatially non-adjacent -> history -> shift -> default.

[0393] a. Alternatively, the candidate insertion order can follow predefined rules, such as spatially adjacent -> spatially non-adjacent (if applicable) -> history (if applicable) -> shift (if applicable) -> default (if applicable).

[0394] 2. For example, CCRM candidates can be inspected at a sub-block (e.g., 4x4) granularity.

[0395] a. For example, each consecutive sub-block within a predefined area can be inspected; for instance, all 4x4 sub-blocks above and to the left of the current video unit can be inspected.

[0396] b. Alternatively, a predefined scatter check order can be used.

[0397] 3. For example, the position of a non-adjacent neighboring block can be based on the block dimension of the current video unit, such as a certain distance from the current video unit, where the distance is proportional to the width and / or height of the current video unit.

[0398] 4. For example, motion shifts (e.g., zero vectors or non-zero vectors) can be used to locate temporal candidates.

[0399] a. For example, motion shift can be based on the motion vector of neighboring blocks.

[0400] b. For example, temporal candidates can come from co-located images.

[0401] c. Alternative locations: Temporal candidates can be derived from reference images, which do not necessarily have to be in the same position as the image.

[0402] 5. For example, history-based CCRM candidates can come from a first-in-first-out history table.

[0403] a. For example, history tables can be initialized at the slice / ctu row / strip / image level.

[0404] 6. For example, time-domain candidates can be derived based on motion displacement.

[0405] a. For example, the temporal candidate of a block encoded by IBC / IntraTMP / inter-frame coding can be derived based on the MV / BV of neighboring blocks.

[0406] iii. For example, it can be based on the MV / BV of the current block. For example, deduplication / redundancy / similarity checks can be applied to CCRM candidate list construction.

[0407] 1. For example, if the candidate to be inserted is different from the specified candidate already in the list, the candidate to be inserted is inserted into the list.

[0408] a. For example, specifying a candidate could refer to all available CCRM candidates in the list.

[0409] b. Alternatively, a specified candidate may refer to one or more specified CCRM candidates in a list (e.g., the last one, and / or the last one in the list – X, where X is a predefined constant).

[0410] 2. For example, the same deduplication / redundancy / similarity check rules can be applied to all types of CCRM candidates.

[0411] a. Alternatively, different deduplication / redundancy / similarity check rules can be applied to different types of CCRM candidates. iv. For example, CCRM candidate reordering can be applied.

[0412] 1. For example, CCRM candidates in the list can be sorted based on the cost inferred by the decoder (e.g., template cost).

[0413] 2. For example, the cost can be derived based on applying CCRM candidates to the reference region / block of the current video unit.

[0414] a. For example, a reference region / block can be identified by the motion vector of the current block.

[0415] b. For example, for each CCRM candidate, the model is first applied to the reference luminance to obtain the predicted reference chromaticity, and then the cost is calculated as the absolute difference between the true reference chromaticity and the predicted reference chromaticity.

[0416] 3. For example, the cost can be derived based on applying CCRM candidates to the neighboring regions / blocks of the current video unit.

[0417] a. For example, a neighboring region / block can be the current block's upper / left nearest neighbor.

[0418] b. For example, for each CCRM candidate, the model is first applied to neighboring luminance to obtain the predicted neighboring chromaticity, and then the cost is calculated as the absolute difference between the true neighboring chromaticity and the predicted neighboring chromaticity.

[0419] 4. For example, based on the cost derived from the decoder above, CCRM candidates can be sorted from lowest cost to highest cost, and the one with the lowest cost is sorted first in the list.

[0420] i. For example, if the CCRM model inheritance mode is used, the specified CCRM candidate model is directly applied to the current video unit without model estimation.

[0421] i. For example, the first candidate in the CCRM list can always be used in the CCRM model inheritance pattern.

[0422] ii. Alternatively, in the bitstream, it can be signaled which candidate from the CCRM list was used for the block encoded / decoded via the CCRM model inheritance mode.

[0423] iii. For example, for RRIBC blocks encoded and decoded using the CCRM model inheritance pattern, the inherited CCRM model can be applied based on the inherited RRIBC flip type.

[0424] 1. For example, if the inherited RRIBC flip type indicates that the inherited CCRM model comes from an RRIBC codec block (e.g., the inherited RRIBC flip type is non-zero), then the CCRM filter taps for the current block can be flipped according to the inherited RRIBC flip type.

[0425] a. For example, when generating CCRM filter taps from unsampled luminance samples, the filter taps of the unsampled luminance samples can be swapped / flipped.

[0426] 2. For example, suppose an 8-tap CCRM model consists of 6 spatial luminance samples, a nonlinear term, and a bias term. Without downsampling, the 6 luminance samples closest to the chromaticity position C are selected from the luminance grid to obtain the spatial luminance samples (L0, ..., L5). The predicted chromaticity value is obtained as: predChromaVal=c0L0+c1L1+c2L2+c3L3+c4L4+c5L5+c6nonlinear((L0+L3+1)>>1)+c7B, where nonlinear is the nonlinear operator of CCRM, and B is the bias. Figure 27 The luminance samples L0, ..., L5 are shown relative to the chromaticity sample C.

[0427] a. For example, if the inherited CCRM model comes from a horizontally flipped RRIBC codec block, then when the inherited CCRM model is applied to the current block (e.g., regardless of whether the current block is RRIBC codec), luma samples L1 and L2 can be swapped. And luma samples L4 and L5 can be swapped.

[0428] b. For example, if the inherited CCRM model comes from a vertically flipped RRIBC codec block, then when the inherited CCRM model is applied to the current block (e.g., regardless of whether the current block is RRIBC codec), luma samples L1 and L4 can be swapped. Furthermore, luma samples L0 and L3 can be swapped. Additionally, luma samples L2 and L5 can be swapped.

[0429] j. For example, the CCRM model of a CCRM codec block can be stored in a cache.

[0430] i. For example, the stored CCRM model information may include the following information.

[0431] 1. CCRM model coefficients / tap / parameters for Cb and Cr components.

[0432] 2. Intermediate values ​​of CCRM codec blocks.

[0433] 3. Bit depth of CCRM codec blocks.

[0434] 4. The offset value of the Y component of the CCRM codec block.

[0435] 5. Offset values ​​of the U and / or V components of the CCRM codec block.

[0436] 6. RRIBC flip type of CCRM codec block.

[0437] k. For example, the CCRM model of a block encoded and decoded by CCRM can be stored in a cache.

[0438] i. For example, the stored CCRM model information may include the following information.

[0439] 1. CCRM model coefficients / tap / parameters for Cb and Cr components.

[0440] 2. Intermediate values ​​of blocks encoded / decoded by CCRM 3. Bit depth of the block encoded / decoded by CCRM 4. Offset of the Y component of the block encoded / decoded by CCRM 5. Offset values ​​of the U and / or V components of the block encoded / decoded by CCRM. 6. RRIBC flip type of blocks encoded / decoded by CCRM ii. Alternatively, model offsets may not be stored in the cache.

[0441] 1. For example, the model offset of a neighboring block may not be reused in the current block / inherited for use in the current block.

[0442] 2. For example, the model offset of the current block is recalculated based on specific available samples.

[0443] l. Alternatively, filter model information (e.g., model coefficients / tap / parameters / offsets for the Y component) can be stored in a cache, where the filter coefficients are calculated based on the relationship between samples in the same component (e.g., all training samples are in the luminance component domain).

[0444] i. Alternatively, model offsets may not be stored in the cache.

[0445] 1. For example, the model offset of a neighboring block may not be reused in the current block / inherited for use in the current block.

[0446] 2. For example, the model offset of the current block is recalculated based on specific available samples.

[0447] 15) The final prediction can be generated from a weighted sum of multiple fusion / hybrid hypotheses, where at least one hypothesis is based on predictions of CCPs (e.g., CCRM, CCRM Merge, CCP Merge, CCLM, LM, CCCM, GLM, etc.).

[0448] a. In one example, the final prediction of a block can be generated based on multiple prediction candidates from different CCPs (e.g., CCRM, CCRMMerge, CCP Merge, CCLM, LM, CCCM, GLM, etc.).

[0449] i. For example, more than one CCP prediction can be merged together.

[0450] ii. For example, the weights / coefficients of different fusion terms can be solved based on the Gaussian elimination method.

[0451] iii. For example, the weights / coefficients of different fusion terms can be solved based on the LDL decomposition method.

[0452] iv. For example, bias terms can be involved for fusion.

[0453] v. For example, nonlinear terms can be involved in fusion.

[0454] b. In one example, multiple CCP models can be derived to obtain a fused prediction.

[0455] i. Fusion predictions can refer to predictions generated through weighted summation.

[0456] ii. In one example, P0 in luminance and chrominance can be used to derive CCP model M0, P1 in luminance and chrominance can be used to derive CCP model M1, and the final chrominance prediction can be derived as wc0×Pc0+wc1×Pc1, where Pc0 and Pc1 are chrominance predictions obtained using M0 and M1, and wc0 and wc1 are weighting factors.

[0457] iii. P0 and P1 can be predictions from two different directions in a bidirectional prediction.

[0458] iv. For example, predictions for two hypotheses can be generated based on the top two candidates in the CCP Merge pattern list, and the final prediction can be derived based on a weighted sum of the predictions for the two hypotheses.

[0459] c. In one example, the chromaticity prediction obtained through CCP can be fused with other predictions.

[0460] i. For example, chromaticity predictions obtained through CCP can be fused with intra-frame angular predictions.

[0461] ii. For example, chromaticity predictions obtained through CCP can be fused with CCLM predictions.

[0462] iii. For example, chromaticity predictions obtained through CCP can be fused with CCCM predictions.

[0463] iv. For example, chroma predictions obtained through CCP can be fused with original predictions (e.g., intra-frame, inter-frame, or IBC predictions) without CCP.

[0464] 1. For example, suppose the final chromaticity prediction can be derived as (w0×P0 + w1×P1 + offset) >> shift, where P0 represents the chromaticity prediction obtained through CCRM and P1 represents the chromaticity prediction without CCP.

[0465] a. The fusion weights w0 and w1 can be fixed and / or predefined, for example, w0=3 and w1=1, or w0=2 and w1=2.

[0466] b. The shift value can be a constant value that can be derived from w0 and w1, for example, shift = log2(w0 + w1).

[0467] c. The offset can be a constant value that can be derived based on the shift and / or fusion weights, for example, offset = shift >> 1, or offset = log2(w0+w1) >> 1.

[0468] d. Alternatively, the fusion weights w0 and w1 can be adaptively determined based on encoding and decoding information (e.g., block dimension, prediction patterns of nearest neighbors, etc.).

[0469] d. In one example, the weighting factors for different hypotheses in the fusion / mixing process can be derived based on predefined rules. The final prediction P is assumed to be derived as P = w0 × P0 + w1 × P1 + w2 × P2 + ..., where w0, w1, and w2 are weighting factors.

[0470] i. For example, fixed values ​​can be assigned to w0, w1, w2… ii. For example, block-based w0, w1, w2... can be assigned.

[0471] iii. For example, for each prediction, a sample-based weighting factor can be assigned (e.g., for different samples in a hypothetical prediction block, at least two different weights can be applied).

[0472] iv. For example, w0, w1, w2 can be derived on the fly (e.g., based on decoded neighboring samples, nearest neighbor prediction patterns, and / or template costs).

[0473] e. In one example, indicators of the weighting factors for different assumptions in the fusion / mixing process can be transmitted via signals in the bitstream.

[0474] i. For example, a lookup table containing multiple sets of weighting factors can be defined, and the index can be signaled to look up the corresponding weight.

[0475] f. For example, the sample values ​​of the final fused / mixed prediction can be clipped to a predefined range, e.g., it can be required to be no less than T1 and no greater than T2, where T2 can depend on the bit depth. For example, Clip1(x) = Clip3(0, (1< <BitDepth) 1,x).

[0476] g. For example, the proposed method can be applied to fuse more than one hypothesis, where at least one hypothesis is based on predictions from a filtering model.

[0477] i. For example, filter coefficients can be calculated based on the relationship between two sets of samples in the same component domain (e.g., one set consists of luminance samples adjacent to the current block, and the other set consists of luminance samples adjacent to the reference block).

[0478] ii. For example, prediction based on a filter model can refer to predictions generated based on applying filters to MC-compensated video units.

[0479] 1. For example, filters can be based on IBC filters, intra-frame TMP filters, LICs for inter-frame, LICs for IBC, EIFs, etc.

[0480] iii. For example, the prediction based on the filter model can be generated based on filter-based merge / inheritance patterns (e.g., IBC filter-based merge / inheritance pattern, intra-frame TMP filter-based merge / inheritance pattern, inter-frame LIC-based merge / inheritance pattern, IBC LIC-based merge / inheritance pattern, EIF-based merge / inheritance pattern).

[0481] 1. For example, a filter-based Merge pattern can refer to a pattern in which the filter model is inherited / derived from the candidate model (e.g., derived from a list of candidate models).

[0482] iv. For example, a prediction based on a filtering model (e.g., encoded or decoded by prediction type A) can be fused with another prediction (e.g., encoded or decoded by prediction type B).

[0483] 1. For example, prediction type B may not be based on a filtering model.

[0484] 2. For example, prediction type B can be a prediction based on a filter model, while the filter types in A and B are different.

[0485] 3. Alternatively, prediction type B can be a prediction based on a filter model, and the filter types in A and B are the same.

[0486] a. For example, predictions for two hypotheses can be generated based on the top two candidates in a filter-based Merge pattern list, and the final prediction can be derived from a weighted sum of the predictions for the two hypotheses.

[0487] 16) Whether CCRM predictions are fused with another prediction can be transmitted via signaling in the bitstream.

[0488] a. For example, a flag can be transmitted via signaling at the video unit level (e.g., TU / PU / CU / strip header / picture header / SPS / PPS level) to indicate such a CCRM fusion mode.

[0489] b. Alternatively, CCRM predictions can always be merged with another prediction without signal transmission.

[0490] 17) For CCRM, bidirectional forecasts can be managed in a different way than unidirectional forecasts. In the following discussion, assume that the forecasts from the two directions are P0 and P1, and that bidirectional forecasts are expressed as Pb = w0 × P0 + w1 × P1, where w0 and w1 are weighting factors.

[0491] a. In one example, Pb in luminance and chrominance can be used to derive the CCRM model.

[0492] b. In one example, P0 or P1 in luminance and chrominance can be used to derive the CCRM model.

[0493] c. In one example, which prediction was used to deduce that the CCRM model could be transmitted via signaling?

[0494] 18) The permission for the CCRM model may depend on at least one of the following: a. Prediction mode of video unit (e.g., MODE_INTRA, MODE_INTER, MODE_IBC, MODE_PLT, etc.); b. Transformation type of video unit (e.g., ACT, color transformation, transformation skip, etc.). c. SBT (e.g., whether SBT is applied to the current video unit); d. The number of non-zero coefficients in a video unit; e. Luminance coefficients (e.g., luminance coefficient values, sum of absolute values ​​of all luminance coefficients, last scan position of non-zero luminance coefficients, AC value, DC value, etc.) and segmentation tree type (e.g., single tree, dual tree). f. Strip type (e.g., I, B, P stripes); g. Color format (e.g., whether it is 4:0:0); h. Availability of chromaticity components; i. For example, CCRM may not be allowed for ACT and / or 4:0:0 color formats.

[0495] j. For example, CCRM on / off can be determined based on the last scan position of a non-zero luminance coefficient.

[0496] i. For example, if the last scan position is less than a threshold, CCRM can be presumed to be disabled for the current chroma unit, and therefore no syntax element is signaled for CCRM use.

[0497] 1. For example, the threshold can be a fixed constant (such as 1).

[0498] 2. For example, the threshold can be a variable based on encoding / decoding information, such as block dimensions.

[0499] ii. For example, if the brightness is not transformed and the encoding / decoding is skipped, such a condition can be checked.

[0500] iii. For example, conditions such as skipping encoding / decoding regardless of whether brightness is transformed can be checked.

[0501] k. For example, CCRM on / off can be determined based on the absolute value of the non-zero luminance coefficient.

[0502] i. For example, it can be determined based on the absolute values ​​of all luminance coefficients (e.g., both AC and DC).

[0503] ii. For example, it can be determined based on the absolute values ​​of all luminance AC coefficients.

[0504] iii. For example, it can be determined based on the luminance DC coefficient value.

[0505] iv. For example, it can be determined based on at least one luminance coefficient value (e.g., DC and / or AC).

[0506] v. For example, if the absolute values ​​are less than a threshold, CCRM can be presumed to be disabled for the current chroma unit, and therefore no syntax element is signaled for CCRM use.

[0507] 1. For example, the threshold can be a fixed constant value.

[0508] 2. For example, the threshold can be a variable based on encoding / decoding information, such as block dimensions.

[0509] vi. For example, if the brightness is not transformed and encoding / decoding is skipped, such a condition can be checked.

[0510] vii. For example, conditions such as skipping encoding / decoding regardless of whether brightness is transformed can be checked.

[0511] viii. For example, such conditions can be checked together with conditions based on block size (e.g., TU width and / or width).

[0512] For example, if a transform skip is used on the luminance component, CCRM may not be applied to the chrominance component.

[0513] i. For example, if a transform skip is used on the luminance component, CCRM can be presumed to be disabled for the current chromaticity unit (e.g., Cb and / or Cr).

[0514] 1. Furthermore, in this case, no syntax elements are used by CCRM on this video unit via signal transmission.

[0515] ii. Alternatively, if transform skip is used on the luminance component, CCRM can be presumed to be always enabled for the current chromaticity unit (e.g., Cb and / or Cr).

[0516] 1. Furthermore, in this case, no syntax elements are used by CCRM on this video unit via signal transmission.

[0517] 19) The application of CCRM can depend on template information.

[0518] a. For example, whether CCRM is allowed for use in video units may depend on the template cost.

[0519] i. For example, if it is determined by a template cost-based approach that CCRM is disabled for the current video unit (i.e., CCRM on / off is presumed rather than signaled), then no syntax element is signaled for CCRM use on that video unit.

[0520] b. For example, such as Figure 28 As shown, assuming the current block is inter-frame encoded / decoded, two costs (e.g., SAD) can be calculated: the first cost is calculated based on the absolute difference between the current template predicted by the CCRM model and the actual reconstruction of the current template, and the second cost is calculated based on the difference between the reference template and the actual reconstruction of the current template. If the first cost is lower than the second cost, CCRM is presumed to be used for the current chroma unit; otherwise, the current chroma unit is encoded / decoded without CCRM. Figure 28 Examples of the current template (2760) and reference template (2750) involved in CCRM encoding and decoding for the current inter-frame block (2720, 2740) are shown.

[0521] i. For example, the CCRM model can be calculated based on the relationship between the reference luminance block (2710) and the reference chrominance block (2730).

[0522] ii. For example, if it is determined that CCRM is to be used, the CCRM model can be applied to the current luminance reconstruction block (i.e., the input of the CCRM model) and generate the current chromaticity prediction predicted by the CCRM model (i.e., the output of the CCRM model).

[0523] iii. For example, in this case (i.e., CCRM on / off is presumed rather than transmitted via signaling), no syntax element is transmitted via signaling for CCRM use on that video unit.

[0524] iv. For example, such a template cost method can be applied to blocks that have undergone inter-frame encoding and decoding.

[0525] c. For example, such as Figure 29 As shown, assuming the current block is encoded / decoded using IBC, two costs (e.g., SAD) can be calculated: the first cost is calculated based on the absolute difference between the current template predicted by the CCRM model and the actual reconstruction of the current template, and the second cost is calculated based on the difference between the reference template and the actual reconstruction of the current template. If the first cost is lower than the second cost, CCRM is presumed to be used for the current chroma unit; otherwise, the current chroma unit is encoded / decoded without CCRM. Figure 29 Examples of the current template (2860) and reference template (2850) involved in CCRM encoding and decoding for the current IBC blocks (2820, 2840) are shown.

[0526] i. For example, the CCRM model can be calculated based on the relationship between the reference luminance block (2810) and the reference chrominance block (2830).

[0527] ii. For example, if it is determined that CCRM is to be used, the CCRM model can be applied to the current luminance reconstruction block (i.e., the input of the CCRM model) and generate the current chromaticity prediction predicted by the CCRM model (i.e., the output of the CCRM model).

[0528] iii. For example, in this case (i.e., CCRM on / off is presumed rather than transmitted via signaling), no syntax element is transmitted via signaling for CCRM use on that video unit.

[0529] iv. For example, such a template cost method can be applied to blocks encoded and decoded by IBC.

[0530] d. For example, whether a template cost-based approach is used to determine CCRM on / off may depend on whether the current video unit (e.g., TU) has residual / non-zero coefficients and / or SBT usage.

[0531] i. For example, for the zero-residual portion of the current CU encoded and decoded by SBT, the CCRM decision based on template cost may not be applied.

[0532] ii. For example, for the portion of the current CU with residuals after SBT encoding and decoding, the CCRM decision based on template cost may not be applied.

[0533] iii. For example, if the CBF flag of the current luminance TU is false, then the CCRM decision based on template cost may not be applied.

[0534] e. For example, for a TU generated from a CU encoded and decoded by SBT (e.g., the TU size is smaller than the CU size), the template can be constructed from neighboring samples outside the entire CU.

[0535] f. For example, in sub-block / sub-segmentation-based inter-frame / IBC modes, since each sub-block can have its own motion vector, the motion vectors of predefined sub-blocks can be used to locate the reference template.

[0536] i. For example, for a TU encoded with affine / sbTMVP, the MV of a specific sub-block (e.g., the top left corner or the center) can be used.

[0537] ii. For example, for a TU that has been encoded and decoded by GPM inter-frame-to-inter-frame, a specific segment of the MV (e.g., part 0 or part 1) can be used.

[0538] iii. For example, for a TU that has been encoded and decoded via GPM inter-intra-frame, the MV of the inter-frame portion can be used.

[0539] iv. For example, for a TU encoded and decoded by GPM, the MV after TM / MMVD can be used.

[0540] 1. As an alternative, MV prior to TM / MMVD can be used.

[0541] v. Alternatively, if the current TU is encoded and decoded in a sub-block / sub-segmentation-based inter-frame / IBC mode, then the template-based approach may not be applied to the CCRM on / off decision.

[0542] g. For example, if sub-block-based CCRM is applied, the samples of the current template predicted by the CCRM model can be constructed based on sub-blocks.

[0543] i. For example, the sample points of the current template predicted by the CCRM model can be constructed by applying multiple CCRM models with boundary sub-blocks (e.g., upper and / or left boundary sub-blocks).

[0544] ii. For example, if the boundary sub-block does not have a valid CCRM model, the corresponding template samples may not be calculated for cost calculation.

[0545] 1. For example, alternatively, it can be filled with real reconstructed template points.

[0546] h. For example, the samples of the current template predicted by the CCRM model can be constructed from the same CCRM model.

[0547] 20) The CCRM model can be applied to the luminance residual block and output CCRM estimated chrominance residual block.

[0548] a. For example, the final chromaticity prediction block can be generated by adding the first candidate to the second candidate.

[0549] i. For example, the first candidate can be based on the CCRM to estimate the chromaticity residual block, and the second candidate can be based on the chromaticity prediction block before / before the CCRM.

[0550] ii. Alternatively, the two candidates can be mixed / fused based on a weighted sum method.

[0551] 1. For example, the weights of the two candidates can be fixed and / or based on predefined rules.

[0552] b. For example, a model can be derived from reference reconstructed samples and applied to the current residual samples.

[0553] i. For example, CCRM model coefficients can be solved / derived based on a set of training samples, where the training samples can refer to luminance (unsampled or downsampled) and chrominance samples in a reference block.

[0554] ii. For example, the derived model coefficients can be applied to the luminance residual block (unsampled or downsampled) and output the chrominance residual block estimated by CCRM.

[0555] c. For example, different offset values ​​can be used during the CCRM model coefficient derivation and CCRM model application processes. Assume an 8-tap CCRM model consists of 6 spatial luminance samples, a nonlinear term, and an offset term; for model coefficient derivation, the estimated reference chromaticity value is obtained as estChromaVal. ref =c0(L0 ref –offset ref )+c1(L1 ref –offset ref )+c2(L2 ref –offset ref )+c3(L3 ref –offsetref )+c4(L4 ref –offset ref )+c5(L5 ref –offset ref )+c6nonlinear((L0 ref +L3 ref +1)>>1)+c7B ref , of which (L0 ref ,…,L5 ref ) represents the six brightness reconstruction samples in the reference block, nonlinear is the nonlinear operator of CCCM, and B ref It is a bias, and offset ref These are block-based variables; for model applications, the estimated current chromaticity residual value is obtained as estChromaResiVal. cur =c0(L0 cur –offset cur )+c1(L1 cur –offset cur )+c2(L2 cur –offset cur )+c3(L3 cur –offset cur )+c4(L4 cur –offset cur )+c5(L5 cur –offset cur )+c6nonlinear((L0 cur +L3 cur +1)>>1)+c7B cur , of which (L0 cur ,…,L5 cur ) represents the six brightness residual samples in the current block, nonlinear is the nonlinear operator of CCCM, and B cur It is a bias, and offset cur These are block-based variables.

[0556] i. For example, the offset value can be derived based on at least one training sample.

[0557] 1. For example, the offset can be derived based on specific brightness training samples (e.g., samples at fixed positions in the brightness reference block, such as the upper left sample or the center sample value).

[0558] 2. For example, the offset can be derived based on the average / median of at least two training samples (e.g., all brightness training samples).

[0559] ii. For example, the offset can be defined as a fixed constant (e.g., 0).

[0560] iii. For example, the first offset (e.g., offset) ref ) can be used in the derivation of model coefficients (e.g., c0…c7).

[0561] 1. For example, a Gaussian elimination solver can be used to minimize the difference between the reference chromaticity block estimated by the CCRM model (e.g., the model input could be a real reference luminance reconstruction block) and the real reference chromaticity reconstruction block.

[0562] 2. For example, the first offset (e.g., offset) ref The average value of all sample points in the reconstructed block can be derived based on the real reference brightness.

[0563] iv. For example, the second offset (e.g., offset) cur () can be used in the application process of the model.

[0564] 1. For example, the CCRM model associated with the derived model coefficients can be applied to the current luminance residual block and output an estimated chrominance residual block for the current block.

[0565] 2. For example, the second offset (e.g., offset) cur ) can be fixed to be equal to 0.

[0566] d. For example, different bias values ​​can be used during the CCRM model coefficient derivation process and the CCRM model application process.

[0567] i. For example, the bias value can be derived based on the bit depth of the luminance and chrominance prediction / reconstruction samples in the bitstream (e.g., it can be equal to 1 << (bit depth - 1)).

[0568] 1. Alternatively, it can be equal to a fixed constant (e.g., 0).

[0569] ii. For example, the first bias (e.g., B) ref () can be used in the derivation of model coefficients.

[0570] 1. For example, B ref It can be equal to 1 << (bit depth - 1).

[0571] iii. For example, the second offset (e.g., B) cur () can be used in the model application process.

[0572] 1. For example, B cur It can be fixed to be equal to 0.

[0573] e. For example, whether the CCRM model is used to predict the current chrominance prediction or the current chrominance residual can be signaled in the bitstream.

[0574] i. Alternatively, it can be implicitly derived based on decoder information.

[0575] ii. Alternatively, the CCRM model can always be applied to predict the current chrominance residual.

[0576] 21) The disclosed CCRM mode can be based on one of the following filters: a. CCLM and / or its variants; b. MMLM and / or its variants; c. CCCM and / or its variants (e.g., GL-CCCM, non-downsampled CCCM, BVG-CCCM, inter-frame CCCM, intra-frame CCCM, etc.); d. GLM and / or its variants; e. Any cross-component prediction that uses information in one channel / component to predict information in another channel / component; f. Any filter-based prediction where the filter coefficients are solved based on the correlation between prediction and / or reconstruction information.

[0577] 22) Block restrictions can be applied to limit the application of specific types of CCP modes.

[0578] a. For example, the CCP mode can only be allowed to be used in cases where the block size meets a predefined rule.

[0579] b. For example, a syntax element can only be signaled when the CCP mode is applicable.

[0580] c. For example, if the CCP mode is not allowed to be used, the syntax element can be presumed to indicate a specific value that no such CCP mode is used for such a block.

[0581] d. For example, at least one of the following block restrictions can be applied to the CCRM mode (assuming W represents the block width and H represents the block height): i. W < T1, or, W <= T1, ii. H < T2, or, H <= T2, iii. Min(W, H) > T3, or, Min(W, H) >= T3, iv. Max(W, H) < T4, or, Max(W, H) <= T4, v. W < T5 H, or, W <= T5 H, vi. W > T6 H, or, W >= T6 H, vii. H < T7 W, or, H <= T7 W, viii. H > T8 W, or, H >= T8 W, ix. W H < T9, or W H <= T9, x. For example, T1, T2, … T9 can be predefined integer constants.

[0582] e. For example, the CCRM mode can be allowed only for small blocks.

[0583] i. For example, it can be allowed for blocks smaller than 4x4, or 8x8, or 16x16, or 32x32.

[0584] ii. For example, it can be allowed for blocks with the number of samples less than 32, or 64, or 128.

[0585] iii. For example, it can be allowed for blocks with the number of samples less than 32, or 64, or 128.

[0586] iv. For example, it can be not allowed for 2xN blocks, where N can be greater than 4 or 8 or 16.

[0587] v. For example, it can be not allowed for Nx2 blocks, where N can be greater than 4 or 8 or 16.

[0588] 23) The disclosed method can be used in a single tree.

[0589] 24) The disclosed method can be used in a dual tree.

[0590] 25) The disclosed method can be used in inter-frame (such as B or P) stripes.

[0591] 26) The disclosed method can be used in intra-frame (such as I) stripes.

[0592] 27) The “block vector” in the disclosed method can be a “motion vector”.

[0593] 28) The training / reference samples in the disclosed method can refer to the predicted samples and / or reconstructed samples in the training / reference regions.

[0594] 29) Whether and / or how the methods disclosed above can be applied can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0595] 30) Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.

[0596] 31) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / double tree segmentation, color components, and stripe / picture type.

[0597] As used herein, the term “video unit” or “video block” can be a sequence, picture, strip, slice, brick, subpicture, codec tree unit (CTU) / codec tree block (CTB), CTU / CTB row, one or more codec units (CU) / codec blocks (CB), one or more CTU / CTB, one or more virtual pipeline data units (VPDU), or a sub-region within a picture / strip / slice / brick.

[0598] Figure 30 A flowchart of a method 3000 for video processing according to an embodiment of the present disclosure is shown. Method 3000 is implemented during the conversion between video units of a video and a bitstream of a video.

[0599] At box 3010, for the conversion between video units and video bitstreams, one or more parameters of the cross-component residual model (CCRM) of the video unit are determined to be inherited from the previous filter-based codec block. CCRM is a filter model that includes a cross-component model or a same-component model.

[0600] At box 3020, the conversion is performed based on CCRM. In some embodiments, the conversion includes encoding video units into a bitstream. In some other embodiments, the conversion includes decoding video units from the bitstream. In this way, estimated chroma predictions can be generated by adding CCRM-estimated residual blocks to chroma predictions without CCRM, thereby improving encoding / decoding efficiency and performance.

[0601] In some embodiments, the cross-component model includes a cross-component residual model or a cross-component prediction model for intra-frame or inter-frame use. In some embodiments, model generation and model application are based on the relationship between different color components.

[0602] In some embodiments, the same component model includes one of the following: inter-frame filter, intra-block copy (IBC) filter, IBC local illumination compensation (IBC-LIC) filter, intra-template matching prediction (IntraTMP) filter, and extended information filter (EIF) filter. In some embodiments, model generation and model application are based on relationships in the same component. In some embodiments, video units are encoded and decoded using either a type of filter model inheritance mode or a type of filter model merge mode.

[0603] In some embodiments, at least one syntax element is transmitted via signaling at the video unit level to specify whether and / or how to use CCRM model inheritance modes. In some embodiments, CCRM model inheritance modes include at least one of the following: CCRM Merge mode, Cross Component Prediction (CCP) Merge mode, IBC Filter Merge mode, or IntraTMP Filter Merge mode.

[0604] In some embodiments, an indicator specifying whether a video unit uses a conventional CCRM mode or a CCRM model inheritance mode is transmitted via signaling based on conditions associated with at least one type of CCRM used for the video unit. In some embodiments, at least one type of CCRM includes a conventional CCRM mode.

[0605] In some embodiments, a first syntax is transmitted via signaling to indicate that the video unit uses a type of CCRM mode, and a second syntax is further transmitted via signaling to indicate which type of CCRM mode is used. In some embodiments, the second syntax is transmitted via signaling if at least one available CCRM candidate exists. In some embodiments, the second syntax is transmitted via signaling if at least one valid CCRM Merge candidate exists.

[0606] In some embodiments, if a CCRM model inheritance mode is used, another syntax is further signaled to specify which CCRM model candidate is selected for inheritance. In some embodiments, a candidate index is encoded to indicate a CCRM candidate from a candidate list. In some embodiments, the candidate index is encoded using a rounded Rice (TR) binarization process, or a rounded binary (TB) binarization process, or a k-th order Exp-Golomb (EGk) binarization process, or a fixed-length (FL) binarization process.

[0607] In some embodiments, the maximum allowed CCRM candidates are specified in the codec. In some embodiments, the maximum allowed CCRM candidates are the maximum length of the candidate list.

[0608] In some embodiments, at least one syntax element depends on the residual or coefficient of the current luma block being transmitted via signaling. In some embodiments, whether at least one syntax element is transmitted via signaling is based on the presence of a residual or non-zero coefficient in the current luma block.

[0609] In some embodiments, whether to transmit the signal syntax element can be based on at least one of the distribution, number, or value of the residuals in the current luma block. Alternatively, whether to transmit the signal syntax element can be based on at least one of the distribution, number, or value of the non-zero coefficients in the current luma block.

[0610] In some embodiments, at least one syntax element is signaled based on a neighbor block prediction method. In some embodiments, at least one syntax element is signaled based on whether the neighbor block prediction method uses CCRM mode or CCRM Merge mode.

[0611] In some embodiments, the context model of at least one syntax element depends on the encoding / decoding information of a neighboring block or the current block. In some embodiments, a neighboring block includes at least one of a left neighboring block or an upper neighboring block.

[0612] In some embodiments, the context model is derived based on whether the neighboring block uses at least one of the following: CCRM mode, multi-mode CCRM (MM-CCRM) mode, multi-filter CCRM (MF-CCRM) mode, or CCRM Merge mode.

[0613] In some embodiments, the context model is derived based on whether the block dimensions of the current block satisfy specific conditions. In some embodiments, the specified context model is used if the dimensions of the current block are W > a*H and / or H > b*W, where W and H represent the width and height of the current block, respectively, and a and b are predefined constants. In some embodiments, a = b = 2.

[0614] In some embodiments, if the CCRM model inheritance pattern is used, a CCRM model candidate list is generated. In some embodiments, CCRM model candidates are obtained based on previously encoded and decoded CCRM blocks.

[0615] In some embodiments, the candidate insertion order follows predefined rules. In some embodiments, the candidate insertion order is the following sequence: spatially adjacent, spatially non-adjacent, history, shift, default.

[0616] In some embodiments, temporal candidates are derived based on motion displacement. In some embodiments, temporal candidates for one of the IBC-encoded, IntraTMP-encoded, or inter-frame-encoded blocks are derived based on the motion vector (MV) or block vector (BV) of neighboring blocks. In some embodiments, temporal candidates for one of the IBC-encoded, IntraTMP-encoded, or inter-frame-encoded blocks are derived based on the MV or BV of the current block.

[0617] In some embodiments, model offsets are not stored in the cache. In some embodiments, model offsets of neighboring blocks are not reused in the current block or inherited for the current block. In some embodiments, the model offset of the current block is recalculated based on available samples.

[0618] In some embodiments, the filtering model information is stored in a cache. Filter coefficients can be calculated based on the relationship between samples in the same component. In some embodiments, the filtering model information includes at least one of the following: model coefficients, model taps, model parameters, and model offset for the Y component. Alternatively or additionally, all training samples are in the luminance component domain.

[0619] In some embodiments, the CCRM model inheritance mode is used for at least one of the following: single-tree or dual-tree. In some embodiments, the CCRM model inheritance mode is used in inter-frame stripes. In some embodiments, the inter-frame stripe is a B-strip or a P-strip.

[0620] In some embodiments, the CCRM model inheritance mode is used in intra-slices. In some embodiments, the intra-slice is an I-slice.

[0621] In some embodiments, training samples or reference samples are predicted samples in the training region or reference region. In some embodiments, training samples or reference samples are reconstructed samples in the training region or reference region.

[0622] In some embodiments, an indication of whether one or more parameters of a video unit's CCRM are inherited from a previous filter-based codec block and / or how to determine that one or more parameters of a video unit's CCRM are inherited from a previous filter-based codec block is indicated at one of the following: sequence level, picture group level, picture level, stripe level, or slice group level. In some embodiments, an indication of whether one or more parameters of a video unit's CCRM are inherited from a previous filter-based codec block and / or how to determine that one or more parameters of a video unit's CCRM are inherited from a previous filter-based codec block is indicated at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header. In some embodiments, an indication of whether one or more parameters of the CCRM of a video unit are inherited from a previous filter-based codec block and / or how one or more parameters of the CCRM of a video unit are inherited from a previous filter-based codec block is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample or pixel.

[0623] In some embodiments, method 3000 further includes: determining, based on the encoded and decoded information of the target block, whether one or more parameters of the CCRM of the video unit are inherited from a previous filter-based codec block and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from a previous filter-based codec block. The encoded and decoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.

[0624] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining one or more parameters of a cross-component residual model (CCRM) of video units inherited from a previous filter-based codec block, wherein the CCRM is a filter model including a cross-component model or a same-component model; and generating a bitstream of video units based on the CCRM.

[0625] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining one or more parameters of a cross-component residual model (CCRM) of video units inherited from a prior filter-based codec block, wherein the CCRM is a filter model including a cross-component model or a same-component model; generating a bitstream of video units based on the CCRM; and storing the bitstream in a non-transitory computer-readable recording medium.

[0626] Figure 31 A flowchart of a method 3100 for video processing according to an embodiment of the present disclosure is shown. Method 3100 is implemented during the conversion between video units of a video and a bitstream of a video.

[0627] At box 3110, for the conversion between video units and the video bitstream, the final prediction of the video unit is determined based on a weighted sum of fusion assumptions. At least one assumption is based on a prediction of cross-component prediction (CCP). In some embodiments, CCP includes at least one of the following: cross-component residual model (CCRM), CCRM Merge, CCP Merge, cross-component linear model (CCLM), convolutional cross-component model (CCCM), gradient linear model (GLM), or linear model (LM).

[0628] At box 3120, the conversion is performed based on the final prediction. In some embodiments, the conversion includes encoding video units into a bitstream. In some other embodiments, the conversion includes decoding video units from the bitstream. In this way, encoding / decoding efficiency and performance can be improved.

[0629] In some embodiments, the final prediction of a video unit is generated based on multiple prediction candidates from different CCPs. In some embodiments, multiple CCP-based predictions are fused together. In some embodiments, multiple CCP models are derived to obtain the fused prediction.

[0630] In some embodiments, the final prediction is a weighted sum of a first prediction and a second prediction, whereby the first prediction is used to derive a first CCP model and the second prediction is used to derive a second CCP model. In some embodiments, P0 in luminance and chrominance is used to derive a CCP model M0, P1 in luminance and chrominance is used to derive a CCP model M1, and the final chrominance prediction is derived as (w0×P0 + w1×P1 + offset) >> shift. In this case, Pc0 and Pc1 represent the chrominance predictions obtained using M0 and M1, respectively, and wc0 and wc1 represent weighting factors.

[0631] In some embodiments, the predictions of the two hypotheses are generated based on the top two candidates in the CCP Merge pattern list. The final prediction can be derived based on a weighted sum of the predictions of the two hypotheses.

[0632] In some embodiments, the chroma prediction obtained by CCP is fused with other predictions. In some embodiments, the chroma prediction obtained by CCP is fused with at least one of the following: angular intra-frame prediction, cross-component linear model (CCLM) prediction, or convolutional cross-component model (CCCM) prediction.

[0633] In some embodiments, the chroma prediction obtained by CCP is fused with the original prediction without CCP. In some embodiments, the original prediction includes at least one of the following: intra-frame prediction, inter-frame prediction, or intra-block copy (IBC) prediction.

[0634] In some embodiments, the final chroma prediction is derived as (w0×P0 + w1×P1 + offset) >> shift. In this case, P0 represents the chroma prediction obtained through CCP, and P1 represents the chroma prediction without CCP, w0 and w1 are the fusion weights, and shift represents the parameter.

[0635] In some embodiments, the determination of the final prediction is applied to the fusion of more than one hypothesis. In this case, at least one hypothesis may be a prediction based on a filtering model.

[0636] In some embodiments, the filter coefficients are calculated based on the relationship between two sets of samples in the same component domain. In some embodiments, one set includes luminance samples adjacent to the current block, and the other set includes luminance samples adjacent to a reference block.

[0637] In some embodiments, the prediction based on the filtering model is a prediction generated by applying the filter to the MC-compensated video unit. In some embodiments, the filter is based on at least one of the following: IBC filter, intra-frame TMP filter, local illumination compensation (LIC) for inter-frame, LIC for IBC, or extended information filter (EIF).

[0638] In some embodiments, the prediction based on the filter model is generated based on either a filter-based merge mode or a filter-based inheritance mode. In some embodiments, the filter-based merge mode includes at least one of the following: an IBC filter-based merge mode, an intra-frame TMP filter-based merge mode, an inter-frame LIC-based merge mode, an IBC LIC-based merge mode, or an EIF-based merge mode; or wherein the filter-based inheritance mode includes at least one of the following: an IBC filter-based inheritance mode, an intra-frame TMP filter-based inheritance mode, an inter-frame LIC-based inheritance mode, an IBC LIC-based inheritance mode, or an EIF-based inheritance mode.

[0639] In some embodiments, the filter-based merge pattern includes a pattern in which the filter model is inherited / derived from the candidate model. In some embodiments, the filter-based prediction is fused with another prediction. In some embodiments, the filter-based prediction is encoded / decoded using a first prediction type, and the other prediction is encoded / decoded using a second prediction type.

[0640] In some embodiments, the second prediction type is not a prediction based on the filter model. In some embodiments, the second prediction type is a prediction based on the filter model, and the first prediction type is different from the second prediction type. In some embodiments, the second prediction type is a prediction based on the filter model, and the first prediction type is the same as the second prediction type. In some embodiments, the predictions of the two hypotheses are generated based on the top two candidates in the filter-based Merge pattern list, and the final prediction is derived based on the weighted sum of the predictions of the two hypotheses.

[0641] In some embodiments, the CCRM model inheritance mode is used for at least one of the following: single-tree or dual-tree. In some embodiments, the CCRM model inheritance mode is used in inter-frame stripes. Alternatively, the CCRM model inheritance mode is used in intra-frame stripes. In some embodiments, the inter-frame stripe is a B-strip or a P-strip. In some embodiments, the intra-frame stripe is an I-strip.

[0642] In some embodiments, training samples or reference samples are predicted samples in the training region or reference region. In some embodiments, training samples or reference samples are reconstructed samples in the training region or reference region.

[0643] In some embodiments, an indication of whether and / or how the final prediction of a video unit is determined based on a weighted sum of fusion assumptions is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, an indication of whether and / or how the final prediction of a video unit is determined based on a weighted sum of fusion assumptions is indicated at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header. In some embodiments, an indication of whether the final prediction of a video unit is determined based on a weighted sum of fusion assumptions and / or how the final prediction of a video unit is determined based on a weighted sum of fusion assumptions is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample point or pixel.

[0644] In some embodiments, method 3100 further includes: determining, based on the encoded and decoded information of the target block, whether to determine the final prediction of the video unit based on a weighted sum of fusion assumptions and / or how to determine the final prediction of the video unit based on a weighted sum of fusion assumptions. The encoded and decoded information may include at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.

[0645] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a final prediction of video units of the video based on a weighted sum of fusion hypotheses, wherein at least one hypothesis is based on a prediction of cross-component prediction (CCP); and generating a bitstream of video units based on the final prediction.

[0646] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining a final prediction of video units of the video based on a weighted sum of fusion hypotheses, wherein at least one hypothesis is based on a prediction of cross-component prediction (CCP); generating a bitstream of video units based on the final prediction; and storing the bitstream in a non-transitory computer-readable recording medium.

[0647] Figure 32A flowchart of a method 3200 for video processing according to an embodiment of the present disclosure is shown. Method 3200 is implemented during the conversion between video units of a video and a bitstream of a video.

[0648] At box 3210, for the conversion between video units and video bitstreams, one or more cross-component residual models (CCRMs) for the video units are determined. The one or more CCRMs include at least one of the following: multi-mode CCRM (MM-CCRM) mode, multi-filter (MF) CCRM (MF-CCRM) mode, or CCRM Merge mode.

[0649] At box 3220, the conversion is performed based on one or more CCRMs. In some embodiments, the conversion includes encoding video units into a bitstream. In some other embodiments, the conversion includes decoding video units from the bitstream. In this way, the estimated chroma prediction can be generated by adding residual blocks estimated by CCRM to chroma predictions without CCRM, thereby improving encoding / decoding efficiency and performance.

[0650] In some embodiments, MM-CCRM, MF-CCRM, or CCRM Merge are applied based on one of the TU, PU, ​​or CU levels.

[0651] In some embodiments, for MM-CCRM applications, TUs are not divided into sub-blocks. In some embodiments, for MM-CCRM applications, CUs are not divided into sub-blocks. In some embodiments, for MM-CCRM applications, PUs are not divided into sub-blocks. In some embodiments, whether to use one of the following is determined at the TU level, PU level, or CU level: TU-based multi-model CCRM, PU-based multi-model CCRM, CU-based multi-model CCRM, TU-based MF-CCRM, PU-based MF-CCRM, CU-based MF-CCRM, TU-based CCRM Merge, PU-based CCRM Merge, or CU-based CCRM Merge.

[0652] In some embodiments, the video unit decides to use a sub-block-based CCRM or one of the following: TU-based CCRM, PU-based CCRM, or CU-based CCRM. In some embodiments, a sub-block-based CCRM includes one of the following: a sub-block-based MM-CCRM, a sub-block-based MF-CCRM, or a sub-block-based CCRM Merge. In some embodiments, a TU-based CCRM includes one of the following: a TU-based MF-CCRM, a TU-based MM-CCRM, or a TU-based CCRM Merge. In some embodiments, a PU-based CCRM includes one of the following: a PU-based MF-CCRM, a PU-based MM-CCRM, or a PU-based CCRM Merge. In some embodiments, a CU-based CCRM includes one of the following: a CU-based MF-CCRM, a CU-based MM-CCRM, or a CU-based CCRM Merge.

[0653] In some embodiments, whether and / or how to apply at least one of the following is derived based on encoding / decoding information at both the encoder and decoder sides: MM-CCRM, CCRM, MF-CCRM, or CCRM Merge. In some embodiments, the determination of whether to use one of the following is based on encoding / decoding information implicitly derived: TU-based multi-model CCRM, PU-based multi-model CCRM, CU-based multi-model CCRM, TU-based MF-CCRM, PU-based MF-CCRM, CU-based MF-CCRM, TU-based CCRM Merge, PU-based CCRM Merge, or CU-based CCRM Merge.

[0654] In some embodiments, sub-block-based CCRM includes one of the following: sub-block-based MM-CCRM, sub-block-based MF-CCRM, or sub-block-based CCRM Merge. In some embodiments, TU-level CCRM includes one of the following: TU-level MF-CCRM, TU-level MM-CCRM, or TU-level CCRM Merge. In some embodiments, PU-level CCRM includes one of the following: PU-level MF-CCRM, PU-level MM-CCRM, or PU-level CCRM Merge, or wherein CU-level CCRM includes one of the following: CU-based MF-CCRM, CU-level MM-CCRM, or CU-level CCRM Merge.

[0655] In some embodiments, the determination of whether to use MM-CCRM, MF-CCRM, CCRMMerge based on M1xM2 sub-blocks, or MM-CCRM, MF-CCRM, or CCRMMerge based on N1xN2 sub-blocks, is implicitly derived based on codec information. In some embodiments, M1 = 16 or 8 or 32 or TU or CU or PU. In some embodiments, M2 = 16 or 8 or 32 or TU or CU or PU. In some embodiments, N1 = 16 or 8 or 32 or TU or CU or PU. In some embodiments, N2 = 16 or 8 or 32 or TU or CU or PU. In some embodiments, M1 is not equal to N1 and / or M2 is not equal to N2.

[0656] In some embodiments, the determination of whether to use single-mode CCRM, MF-CCRM, CCRM Merge, multi-mode CCRM, MF-CCRM, or CCRM Merge is derived based on codec information. In some embodiments, whether and / or how to apply at least one of the following is signaled in the bitstream: MM-CCRM, CCRM, MF-CCRM, or CCRM Merge. In some embodiments, if the video unit is CCRM encoded, the syntax element is further signaled to indicate whether it is MM-CCRM, MF-CCRM, or CCRM Merge.

[0657] In some embodiments, syntax elements are signaled to indicate whether they are sub-block-based MM-CCRMs or one of the following: TU-based MM-CCRM, CP-based MM-CCRM, PU-based MM-CCRM, TU-based MF-CCRM, CP-based MF-CCRM, PU-based MF-CCRM, TU-based CCRM Merge, CP-based CCRM Merge, or PU-based CCRM Merge.

[0658] In some embodiments, a syntax element is signaled based on a condition regarding a block dimension. In some embodiments, the block dimension includes at least one of width or height. In some embodiments, if W*H < T, the syntax element is not signaled, where W represents the width, H represents the height, and T represents a threshold. In some embodiments, T = 16 or 32.

[0659] In some embodiments, a syntax element is signaled depending on the residual or coefficients of a current luminance block. In some embodiments, whether to signal a syntax element is based on whether there is a residual or non-zero coefficients in the current luminance block.

[0660] In some embodiments, whether to signal a syntax element is based on the distribution or number or value of the residuals in the current luminance block. In some embodiments, whether to signal a syntax element is based on the distribution or number or value of the non-zero coefficients in the current luminance block.

[0661] In some embodiments, a syntax element is signaled based on a prediction method of neighboring blocks. In some embodiments, a syntax element is signaled based on whether neighboring blocks use at least one of the following: CCRM, MM-CCRM, MF-CCRM, or CCRM Merge mode.

[0662] In some embodiments, the context model of a syntax element depends on the coding and decoding information of neighboring blocks or the current block. In some embodiments, the context model is derived based on whether neighboring blocks use CCRM or MM-CCRM or MF-CCRM or CCRM Merge mode.

[0663] In some embodiments, the context model is derived based on whether the block dimension of the current block satisfies a condition. In some embodiments, if the dimension of the current block is W > a*H and / or H > b*W, a specified context model is used, where W and H represent the width and height of the current block respectively, and a and b are predefined constants. For example, a = b = 2.

[0664] In some embodiments, the CCRM model inheritance mode is used for at least one of the following: single tree or double tree. In some embodiments, the CCRM model inheritance mode is used in inter-frame stripes. Alternatively, the CCRM model inheritance mode is used in intra-frame stripes.

[0665] In some embodiments, the inter-frame stripe is a B stripe or a P stripe. In some embodiments, the intra-frame stripe is an I stripe.

[0666] In some embodiments, a training sample or a reference sample is a predicted sample in a training region or a reference region. In some embodiments, a training sample or a reference sample is a reconstructed sample in a training region or a reference region.

[0667] In some embodiments, the indication of whether one or more parameters of a video unit's CCRM are inherited from a previously CCRM-encoded block and / or how to determine that one or more parameters of a video unit's CCRM are inherited from a previously CCRM-encoded block is indicated at one of the following: sequence level, picture group level, picture level, stripe level, or slice group level. In some embodiments, the indication of whether one or more parameters of a video unit's CCRM are inherited from a previously CCRM-encoded block and / or how to determine that one or more parameters of a video unit's CCRM are inherited from a previously CCRM-encoded block is indicated at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header. In some embodiments, an indication of whether or not one or more parameters of the CCRM of a video unit are inherited from a previously CCRM-encoded block and / or how one or more parameters of the CCRM of a video unit are inherited from a previously CCRM-encoded block is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample or pixel.

[0668] In some embodiments, method 3200 further includes: determining, based on the encoded / decoded information of the target block, whether one or more parameters of the CCRM of the video unit are inherited from a previously CCRM-encoded block and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from a previously CCRM-encoded block. The encoded / decoded information may include at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.

[0669] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining one or more cross-component residual models (CCRMs) for video units, wherein the one or more CCRMs include at least one of the following: a multi-mode CCRM (MM-CCRM) model, a multi-filter (MF) CCRM (MF-CCRM) model, or a CCRM Merge model; and generating a bitstream of video units based on the one or more CCRMs.

[0670] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining one or more cross-component residual models (CCRMs) for video units of the video, wherein the one or more CCRMs include at least one of the following: a multi-mode CCRM (MM-CCRM) model, a multi-filter (MF) CCRM (MF-CCRM) model, or a CCRM Merge model; generating a bitstream of the video units based on the one or more CCRMs; and storing the bitstream in a non-transitory computer-readable recording medium.

[0671] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.

[0672] Item 1. A method for video processing, comprising: a conversion between a video unit and a bitstream of the video; determining that one or more parameters of a cross-component residual model (CCRM) of the video unit are inherited from a previous filter-based codec block, wherein the CCRM is a filter model including a cross-component model or a same-component model; and performing the conversion based on the CCRM.

[0673] Item 2. The method according to Item 1, wherein the cross-component model includes a cross-component residual model or a cross-component prediction model for intra-frame or inter-frame use, wherein model generation and model application are based on the relationship between different color components.

[0674] Item 3. The method according to Item 1, wherein the same component model includes one of the following: inter-frame filter, intra-block copy (IBC) filter, IBC local illumination compensation (IBC-LIC) filter, intra-template matching prediction (IntraTMP) filter, extended information filter (EIF) filter, wherein the model generation and model application are based on the relationship in the same component.

[0675] Item 4. According to the method described in Item 1, the video unit is encoded or decoded using either a type of filter model inheritance mode or a type of filter model merge mode.

[0676] Item 5. According to the method described in Item 1, at least one syntax element is transmitted via signal at the video unit level to specify whether and / or how to use the CCRM model inheritance mode.

[0677] Item 6. The method according to Item 5, wherein the CCRM model inheritance mode includes at least one of the following: CCRM Merge mode, cross-component prediction (CCP) Merge mode, IBC filter Merge mode, or IntraTMP filter Merge mode.

[0678] Item 7. The method according to Item 5, wherein an indicator specifying whether the video unit uses a conventional CCRM mode or a CCRM model inheritance mode is transmitted via signaling based on conditions associated with at least one type of CCRM used for the video unit.

[0679] Item 8. The method according to Item 7, wherein the at least one type of CCRM includes a conventional CCRM pattern.

[0680] Item 9. The method according to Item 8, wherein a first syntax is transmitted by signal to indicate that the video unit uses a type of CCRM mode, and a second syntax is further transmitted by signal to indicate which type of CCRM mode is used.

[0681] Item 10. The method according to Item 9, wherein if at least one available CCRM candidate exists, the second syntax is transmitted via signaling.

[0682] Item 11. The method according to Item 9, wherein the second syntax is transmitted by signaling if at least one valid CCRM Merge candidate exists.

[0683] Item 12. The method according to Item 5, wherein if the CCRM model inheritance mode is used, another syntax is further transmitted via signaling to specify which CCRM model candidate is selected to be inherited.

[0684] Item 13. The method according to Item 12, wherein the candidate index is encoded or decoded to indicate CCRM candidates from the candidate list.

[0685] Item 14. The method according to Item 12, wherein the candidate index is encoded or decoded using a rounded Rice (TR) binarization process, or a rounded binary (TB) binarization process, or a k-th order Exp-Golomb (EGk) binarization process, or a fixed length (FL) binarization process.

[0686] Item 15. The method according to Item 12, wherein the maximum allowed CCRM candidate is specified in the codec.

[0687] Item 16. The method according to Item 15, wherein the maximum allowed CCRM candidate is the maximum length of the candidate list.

[0688] Item 17. The method according to Item 5, wherein the at least one syntax element depends on the residual or coefficient of the current luma block being transmitted via signal.

[0689] Item 18. The method according to Item 17, wherein whether the at least one syntax element is transmitted via signal is based on the presence of a residual or non-zero coefficient in the current luminance block.

[0690] Item 19. The method according to Item 17, wherein whether the syntax element is transmitted by signaling is based on at least one of the distribution, number, or value of residuals in the current luminance block, or wherein whether the syntax element is transmitted by signaling is based on at least one of the distribution, number, or value of non-zero coefficients in the current luminance block.

[0691] Item 20. The method according to Item 5, wherein the at least one syntax element is transmitted via signaling based on a neighbor block prediction method.

[0692] Item 21. The method according to Item 20, wherein the at least one syntax element is transmitted via signaling based on whether the neighboring block uses CCRM mode or CCRM Merge mode.

[0693] Item 22. According to the method described in Item 5, the context model of the at least one syntax element depends on the encoding / decoding information of the neighboring block or the current block.

[0694] Item 23. The method according to Item 21 or 22, wherein the neighboring block includes at least one of the left neighboring block or the upper neighboring block.

[0695] Item 24. The method according to Item 22, wherein the context model is derived based on whether the neighboring block uses at least one of the following: CCRM mode, multi-mode CCRM (MM-CCRM) mode, multi-filter CCRM (MF-CCRM) mode, or CCRM Merge mode.

[0696] Item 25. The method according to Item 22, wherein the context model is derived based on whether the block dimension of the current block satisfies a specific condition.

[0697] Item 26. The method according to Item 25, wherein a specified context model is used if the dimension of the current block is W>a*H and / or H>b*W, wherein W and H represent the width and height of the current block, respectively, and a and b are predefined constants.

[0698] Item 27. The method described in Item 26, where a=b=2.

[0699] Item 28. The method described in Item 1, wherein if the CCRM model inheritance pattern is used, a CCRM model candidate list is generated.

[0700] Item 29. The method according to Item 28, wherein CCRM model candidates are obtained based on previously encoded and decoded CCRM blocks.

[0701] Item 30. The method according to Item 29, wherein the candidate insertion order follows a predefined rule.

[0702] Item 31. The method according to Item 30, wherein the candidate insertion order is the following sequence: spatially adjacent, spatially non-adjacent, history, shift, default.

[0703] Item 32. The method according to Item 29, wherein the time-domain candidate is derived based on motion displacement.

[0704] Item 33. The method according to Item 32, wherein the temporal candidate of one of the IBC-encoded blocks, IntraTMP-encoded blocks, or inter-frame-encoded blocks is derived based on the motion vector (MV) or block vector (BV) of neighboring blocks.

[0705] Item 34. The method according to Item 32, wherein the temporal candidate of one of the IBC-encoded block, the IntraTMP-encoded block, or the inter-frame-encoded block is derived based on the MV or BV of the current block.

[0706] Item 35. The method described in Item 1, wherein the model offset is not stored in the cache.

[0707] Item 36. The method according to Item 35, wherein the model offset of neighboring blocks is not reused in the current block or is not inherited for the current block.

[0708] Item 37. The method according to Item 35, wherein the model offset of the current block is recalculated based on available samples.

[0709] Item 38. The method according to Item 1, wherein the filtering model information is stored in a cache, and wherein the filter coefficients are calculated based on the relationship between samples in the same component.

[0710] Item 39. The method according to Item 38, wherein the filtering model information includes at least one of the following: model coefficients, model taps, model parameters, model offset of the Y component, and / or wherein all training samples are in the luminance component domain.

[0711] Item 40. The method described in Item 1, wherein the CCRM model inheritance pattern is used for at least one of the following: single tree or double tree.

[0712] Item 41. The method described in Item 1, wherein the CCRM model inheritance mode is used in inter-frame stripes.

[0713] Item 42. The method according to Item 41, wherein the inter-frame stripe is a B stripe or a P stripe.

[0714] Item 43. The method described in Item 1, wherein the CCRM model inheritance mode is used in intra-strip.

[0715] Item 44. The method according to Item 43, wherein the intra-frame stripe is an I-strip.

[0716] Item 45. The method according to Item 1, wherein the training sample or reference sample is a predicted sample in the training region or reference region.

[0717] Item 46. The method according to Item 1, wherein the training samples or reference samples are reconstructed samples in the training region or reference region.

[0718] Item 47. The method according to any one of Items 1 to 46, wherein an indication of whether one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.

[0719] Item 48. The method according to any one of Items 1 to 46, wherein an indication of whether or not it is determined that one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.

[0720] Item 49. The method according to any one of items 1 to 46, wherein an indication of whether one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample or pixel.

[0721] Item 50. The method according to any one of items 1 to 46 further comprises: determining, based on the encoded and decoded information of the target block, whether one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, the encoded and decoded information including at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.

[0722] Item 51. A method for video processing, comprising: a conversion between video units of a video and a bitstream of the video; determining a final prediction of the video units based on a weighted sum of fusion assumptions, wherein at least one assumption is based on a prediction of cross-component prediction (CCP); and performing the conversion based on the final prediction.

[0723] Item 52. The method according to Item 51, wherein the CCP includes at least one of the following: cross-component residual model (CCRM), CCRM Merge, CCP Merge, cross-component linear model (CCLM), convolution cross-component model (CCCM), gradient linear model (GLM), or linear model (LM).

[0724] Item 53. The method according to Item 51, wherein the final prediction of the video unit is generated based on multiple prediction candidates from different CCPs.

[0725] Item 54. The method described in Item 53, wherein multiple CCP-based predictions are fused together.

[0726] Item 55. The method according to Item 51, wherein multiple CCP models are derived to obtain fused predictions.

[0727] Item 56. The method according to Item 55, wherein the final prediction is a weighted sum of a first prediction and a second prediction, and the first prediction is used to derive a first CCP model, and the second prediction is used to derive a second CCP model.

[0728] Item 57. The method according to Item 56, wherein P0 in luminance and chrominance is used to derive CCP model M0, P1 in luminance and chrominance is used to derive CCP model M1, and the final chrominance prediction is derived as wc0×Pc0+wc1×Pc1, where Pc0 and Pc1 represent chrominance predictions obtained using M0 and M1, respectively, and wc0 and wc1 represent weighting factors.

[0729] Item 58. The method according to Item 55, wherein predictions of two hypotheses are generated based on the first two candidates in the CCP Merge pattern list, and a final prediction is derived based on a weighted sum of the predictions of the two hypotheses.

[0730] Item 59. The method according to Item 51, wherein the chromaticity prediction obtained by CCP is fused with other predictions.

[0731] Item 60. The method according to Item 59, wherein the chromaticity prediction obtained by CCP is fused with at least one of the following: angular intra-frame prediction, cross-component linear model (CCLM) prediction, or convolutional cross-component model (CCCM) prediction.

[0732] Item 61. The method according to Item 59, wherein the chromaticity prediction obtained by CCP is fused with the original prediction without CCP.

[0733] Item 62. The method according to Item 59, wherein the original prediction includes at least one of the following: intra-frame prediction, inter-frame prediction, or intra-block copy (IBC) prediction.

[0734] Item 63. The method according to Item 61, wherein the final chroma prediction is derived as (w0×P0+w1×P1+ offset)>>shift, where P0 represents the chroma prediction obtained through CCP and P1 represents the chroma prediction without CCP, w0 and w1 are fusion weights, and shift represents the parameter.

[0735] Item 64. The method according to Item 51, wherein the determination of the final prediction is applied to the fusion of more than one hypothesis, wherein at least one hypothesis is a prediction based on a filtering model.

[0736] Item 65. The method according to Item 64, wherein the filter coefficients are calculated based on the relationship between two sets of samples in the same component domain.

[0737] Item 66. The method according to Item 65, wherein one set includes luminance samples adjacent to the current block and the other set includes luminance samples adjacent to the reference block.

[0738] Item 67. The method according to Item 64, wherein the prediction based on the filter model is a prediction generated by applying a filter to a MC-compensated video unit.

[0739] Item 68. The method according to Item 67, wherein the filter is based on at least one of the following: an IBC filter, an intra-frame TMP filter, a local illumination compensation (LIC) for inter-frame, a LIC for IBC, or an extended information filter (EIF).

[0740] Item 69. The method according to Item 64, wherein the prediction based on the filter model is generated based on either a filter-based Merge pattern or a filter-based Inheritance pattern.

[0741] Item 70. The method according to Item 69, wherein the filter-based Merge mode includes at least one of the following: IBC filter-based Merge mode, intra-frame TMP filter-based Merge mode, inter-frame LIC-based Merge mode, IBC LIC-based Merge mode, EIF-based Merge mode, or wherein the filter-based inheritance mode includes at least one of the following: IBC filter-based inheritance mode, intra-frame TMP filter-based inheritance mode, inter-frame LIC-based inheritance mode, IBC LIC-based inheritance mode, EIF-based inheritance mode.

[0742] Item 71. The method according to Item 70, wherein the filter-based Merge pattern includes a pattern in which the filter model inherits / derives from the candidate model.

[0743] Item 72. The method according to Item 64, wherein a prediction based on a filtering model is fused with another prediction.

[0744] Item 73. The method according to Item 72, wherein the prediction based on the filter model is encoded or decoded using a first prediction type, and the other prediction is encoded or decoded using a second prediction type.

[0745] Item 74. The method according to Item 73, wherein the second prediction type is not a prediction based on a filtering model.

[0746] Item 75. The method according to Item 73, wherein the second prediction type is a prediction based on a filtering model, and the first prediction type is different from the second prediction type.

[0747] Item 76. The method according to Item 73, wherein the second prediction type is a prediction based on a filtering model, and the first prediction type is the same as the second prediction type.

[0748] Item 77. The method according to Item 76, wherein the predictions of the two hypotheses are generated based on the first two candidates in a filter-based Merge pattern list, and the final prediction is derived based on a weighted sum of the predictions of the two hypotheses.

[0749] Item 78. The method described in Item 51, wherein the CCRM model inheritance pattern is used for at least one of the following: single tree or double tree.

[0750] Item 79. The method according to Item 51, wherein the CCRM model inheritance mode is used in inter-frame stripes, or wherein the CCRM model inheritance mode is used in intra-frame stripes.

[0751] Item 80. The method according to Item 79, wherein the inter-frame stripe is a B stripe or a P stripe.

[0752] Item 81. The method according to Item 79, wherein the intra-frame stripe is an I-strip.

[0753] Item 82. The method according to Item 51, wherein the training sample or reference sample is a predicted sample in the training region or reference region.

[0754] Item 83. The method according to Item 51, wherein the training samples or reference samples are reconstructed samples in the training region or reference region.

[0755] Item 84. The method according to any one of items 51 to 83, wherein an indication of whether the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, and / or how the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.

[0756] Item 85. The method according to any one of items 51 to 83, wherein an indication of whether the final prediction of the video unit is determined based on the weighted sum of the fusion assumptions, and / or how the final prediction of the video unit is determined based on the weighted sum of the fusion assumptions, is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.

[0757] Item 86. The method according to any one of items 51 to 83, wherein an indication of whether the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, and / or how the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample or pixel.

[0758] Item 87. The method according to any one of items 51 to 83 further comprises: determining, based on the encoded and decoded information of the target block, whether to determine the final prediction of the video unit based on the weighted sum of the fusion hypothesis, and / or how to determine the final prediction of the video unit based on the weighted sum of the fusion hypothesis, wherein the encoded and decoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.

[0759] Item 88. A method for video processing, comprising: a conversion between video units of a video and a bitstream of the video; determining one or more cross-component residual models (CCRMs) for the video units, wherein the one or more CCRMs include at least one of: a multi-mode CCRM (MM-CCRM) model, a multi-filter (MF) CCRM (MF-CCRM) model, or a CCRM Merge model; and performing the conversion based on the one or more CCRMs.

[0760] Item 89. The method according to Item 88, wherein MM-CCRM or MF-CCRM or CCRM Merge is applied based on one of the TU level, PU level or CU level.

[0761] Item 90. The method according to Item 89, wherein the MM-CCRM or the MF-CCRM or the CCRMMerge is applied based on one of TU, ​​CU or PU.

[0762] Item 91. The method according to Item 90, wherein for the application of the MM-CCRM, TU is not divided into sub-blocks, and / or wherein for the application of the MM-CCRM, CU is not divided into sub-blocks, and / or wherein for the application of the MM-CCRM, PU is not divided into sub-blocks.

[0763] Item 92. The method according to Item 89, wherein whether or not to use one of the following is determined at the TU level, PU level or CU level: TU-based multi-model CCRM, PU-based multi-model CCRM, CU-based multi-model CCRM, TU-based MF-CCRM, PU-based MF-CCRM, CU-based MF-CCRM, TU-based CCRM Merge, PU-based CCRM Merge or CU-based CCRM Merge.

[0764] Item 93. The method according to Item 92, wherein the video unit decides to use a sub-block-based CCRM or one of the following: a TU-based CCRM, a PU-based CCRM, or a CU-based CCRM.

[0765] Item 94. The method according to Item 93, wherein the sub-block-based CCRM includes one of the following: sub-block-based MM-CCRM, sub-block-based MF-CCRM, or sub-block-based CCRM Merge; or wherein the TU-based CCRM includes one of the following: TU-based MF-CCRM, TU-based MM-CCRM, or TU-based CCRM Merge; or wherein the PU-based CCRM includes one of the following: PU-based MF-CCRM, PU-based MM-CCRM, or PU-based CCRM Merge; or wherein the CU-based CCRM includes one of the following: CU-based MF-CCRM, CU-based MM-CCRM, or CU-based CCRM Merge.

[0766] Item 95. The method according to Item 89, wherein whether and / or how to apply at least one of the following encoding / decoding information at both the encoder side and the decoder side is derived: MM-CCRM, CCRM, MF-CCRM, or CCRM Merge.

[0767] Item 96. The method according to Item 95, wherein the determination of whether to use one of the following is based on the implicit derivation of the encoding / decoding information: TU-based multi-model CCRM, PU-based multi-model CCRM, CU-based multi-model CCRM, TU-based MF-CCRM, PU-based MF-CCRM, CU-based MF-CCRM, TU-based CCRM Merge, PU-based CCRM Merge, or CU-based CCRM Merge.

[0768] Item 97. The method according to Item 96, wherein the sub-block-based CCRM includes one of the following: sub-block-based MM-CCRM, sub-block-based MF-CCRM, or sub-block-based CCRM Merge; or wherein the TU-level CCRM includes one of the following: TU-level MF-CCRM, TU-level MM-CCRM, or TU-level CCRM Merge; or wherein the PU-level CCRM includes one of the following: PU-level MF-CCRM, PU-level MM-CCRM, or PU-level CCRM Merge; or wherein the CU-level CCRM includes one of the following: CU-based MF-CCRM, CU-level MM-CCRM, or CU-level CCRM Merge.

[0769] Item 98. The method according to Item 95, wherein the determination of whether to use MM-CCRM, MF-CCRM, or CCRM Merge based on M1xM2 sub-blocks or MM-CCRM, MF-CCRM, or CCRM Merge based on N1xN2 sub-blocks is implicitly derived based on the encoding / decoding information.

[0770] Item 99. The method according to Item 98, wherein M1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein M2 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N1 = 16 or 8 or 32 or TU or CU or PU, and / or wherein N2 = 16 or 8 or 32 or TU or CU or PU, and / or wherein M1 is not equal to N1 and / or M2 is not equal to N2.

[0771] Item 100. The method described in Item 95, wherein the determination of whether to use a single-mode CCRM, MF-CCRM, CCRM Merge, multi-mode CCRM, MF-CCRM, or CCRM Merge is derived based on encoding / decoding information.

[0772] Item 101. The method according to Item 88, wherein at least one of the following is applied and / or at least one of the following is applied in a way that is transmitted by signal in the bitstream: MM-CCRM, CCRM, MF-CCRM, or CCRM Merge.

[0773] Item 102. The method according to Item 101, wherein if the video unit is CCRM encoded or decoded, the syntax element is further signaled to indicate whether it is MM-CCRM, or whether it is MF-CCRM, or whether it is CCRM Merge.

[0774] Item 103. The method according to Item 101, wherein syntax elements are signaled to indicate whether it is a sub-block-based MM-CCRM or one of the following: TU-based MM-CCRM, CP-based MM-CCRM, PU-based MM-CCRM, TU-based MF-CCRM, CP-based MF-CCRM, PU-based MF-CCRM, TU-based CCRM Merge, CP-based CCRM Merge, or PU-based CCRM Merge.

[0775] Item 104. The method according to Item 101, wherein the syntax element is signaled to indicate whether it is sub-block based CCRM or one of the following: TU based MM-CCRM, CP based MM-CCRM, PU based MM-CCRM, TU based MF-CCRM, CP based MF-CCRM, PU based MF-CCRM, TU based CCRM Merge, CP based CCRM Merge, or PU based CCRM Merge.

[0776] Item 105. The method according to Item 101, wherein the syntax element is signaled based on a condition regarding the block dimension.

[0777] Item 106. The method according to Item 105, wherein the block dimension includes at least one of width or height.

[0778] Item 107. The method according to Item 106, wherein if W*H < T, the syntax element is not signaled, where W represents the width, H represents the height, and T represents a threshold.

[0779] Item 108. The method according to Item 107, wherein T = 16 or 32.

[0780] Item 109. The method according to Item 101, wherein the syntax element is signaled depending on the residual or coefficients of the current luma block.

[0781] Item 110. The method according to Item 109, wherein whether the syntax element is signaled is based on the presence of residual or non-zero coefficients in the current luma block.

[0782] Item 111. The method according to Item 109, wherein whether the syntax element is signaled is based on the distribution or number or value of the residual in the current luma block, or wherein whether the syntax element is signaled is based on the distribution or number or value of the non-zero coefficients in the current luma block.

[0783] Item 112. The method according to Item 101, wherein the syntax element is signaled based on the prediction method of neighboring blocks.

[0784] Item 113. The method according to Item 112, wherein the syntax element is signaled based on whether the neighboring blocks use at least one of the following: CCRM, MM-CCRM, MF-CCRM or CCRM Merge mode.

[0785] Item 114. The method according to Item 101, wherein the context model of the syntax element depends on the encoding / decoding information of neighboring blocks or the current block.

[0786] Item 115. The method according to Item 114, wherein the context model is derived based on whether the neighboring block uses CCRM, MM-CCRM, MF-CCRM, or CCRM Merge mode.

[0787] Item 116. The method according to Item 114, wherein the context model is derived based on whether the block dimension of the current block satisfies the conditions.

[0788] Item 117. The method according to Item 116, wherein a specified context model is used if the dimension of the current block is W>a*H and / or H>b*W, where W and H represent the width and height of the current block, respectively, and a and b are predefined constants.

[0789] Item 118. The method described in Item 117, where a=b=2.

[0790] Item 119. The method described in Item 88, wherein the CCRM model inheritance pattern is used for at least one of the following: single tree or double tree.

[0791] Item 120. The method according to Item 88, wherein the CCRM model inheritance mode is used in inter-frame stripes, or wherein the CCRM model inheritance mode is used in intra-frame stripes.

[0792] Item 121. The method according to Item 120, wherein the inter-frame stripe is a B stripe or a P stripe.

[0793] Item 122. The method according to Item 120, wherein the intra-frame stripe is an I-strip.

[0794] Item 123. The method according to Item 88, wherein the training sample or reference sample is a predicted sample in the training region or reference region.

[0795] Item 124. The method according to Item 88, wherein the training samples or reference samples are reconstructed samples in the training region or reference region.

[0796] Item 125. The method according to any one of Items 88 to 124, wherein an indication of whether or not it is determined that one or more parameters of the CCRM of the video unit are inherited from a previously CCRM-encoded block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previously CCRM-encoded block, is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.

[0797] Item 126. The method according to any one of Items 88 to 124, wherein an indication of whether or not it is determined that one or more parameters of the CCRM of the video unit are inherited from a previously CCRM-encoded block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previously CCRM-encoded block, is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.

[0798] Item 127. The method according to any one of Items 88 to 124, wherein an indication of whether or not it is determined that one or more parameters of the CCRM of the video unit are inherited from a previously CCRM-encoded block, and / or how it is determined that one or more parameters of the CCRM of the video unit are inherited from the previously CCRM-encoded block, is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample or pixel.

[0799] Item 128. The method according to any one of items 88 to 124, further comprising: determining, based on the encoded information of the target block, whether one or more parameters of the CCRM of the video unit are inherited from the previously CCRM-encoded block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previously CCRM-encoded block, the encoded information including at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.

[0800] Item 129. The method according to any one of items 1 to 128, wherein the conversion includes encoding the video unit into the bitstream.

[0801] Item 130. The method according to any one of items 1 to 128, wherein the conversion includes decoding the video unit from the bitstream.

[0802] Item 131. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1 to 130.

[0803] Item 132. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1 to 130.

[0804] Item 133. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining one or more parameters of a cross-component residual model (CCRM) of video units of the video that are inherited from a previous filter-based codec block, wherein the CCRM is a filter model that includes a cross-component model or a same-component model; and generating the bitstream of the video units based on the CCRM.

[0805] Item 134. A method for storing a bitstream of video, comprising: determining one or more parameters of a cross-component residual model (CCRM) of video units of the video that are inherited from a previous filter-based codec block, wherein the CCRM is a filter model that includes a cross-component model or a same-component model; generating the bitstream of the video units based on the CCRM; and storing the bitstream in a non-transitory computer-readable recording medium.

[0806] Item 135. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining a final prediction of video units of the video based on a weighted sum of fusion hypotheses, wherein at least one hypothesis is based on a prediction of cross-component prediction (CCP); and generating the bitstream of the video units based on the final prediction.

[0807] Item 136. A method for storing a bitstream of video, comprising: determining a final prediction of video units of the video based on a weighted sum of fusion hypotheses, wherein at least one hypothesis is based on a prediction of cross-component prediction (CCP); generating the bitstream of the video units based on the final prediction; and storing the bitstream in a non-transitory computer-readable recording medium.

[0808] Item 137. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining one or more cross-component residual models (CCRMs) for video units of the video, wherein the one or more CCRMs comprise at least one of: a multi-mode CCRM (MM-CCRM) model, a multi-filter (MF) CCRM (MF-CCRM) model, or a CCRM Merge model; and generating the bitstream of the video units based on the one or more CCRMs.

[0809] Item 138. A method for storing a bitstream of video, comprising: determining one or more cross-component residual models (CCRMs) for video units of the video, wherein the one or more CCRMs include at least one of the following: a multi-mode CCRM (MM-CCRM) model, a multi-filter (MF) CCRM (MF-CCRM) model, or a CCRM Merge model; generating the bitstream of the video units based on the one or more CCRMs; and storing the bitstream in a non-transitory computer-readable recording medium.

[0810] Example device Figure 33 A block diagram of a computing device 3300 in which various embodiments of the present disclosure may be implemented is shown. The computing device 3300 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0811] It should be understood that, Figure 33 The computing device 3300 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0812] like Figure 33 As shown, computing device 3300 includes general-purpose computing device 3300. Computing device 3300 may include at least one or more processors or processing units 3310, memory 3320, storage unit 3330, one or more communication units 3340, one or more input devices 3350, and one or more output devices 3360.

[0813] In some embodiments, the computing device 3300 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 3300 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0814] Processing unit 3310 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 3320. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 3300. Processing unit 3310 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0815] Computing device 3300 typically includes various computer storage media. Such media can be any media accessible by computing device 3300, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 3320 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 3330 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 3300.

[0816] The computing device 3300 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 33 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0817] Communication unit 3340 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 3300 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 3300 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0818] Input device 3350 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 3360 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 3340, computing device 3300 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 3300 can also communicate with one or more devices that enable a user to interact with computing device 3300, or, if necessary, with any device (e.g., network card, modem, etc.) that enables computing device 3300 to communicate with one or more other computing devices. Such communication can be performed via an input / output (I / O) interface (not shown).

[0819] In some embodiments, some or all of the components of computing device 3300 may be deployed in a cloud computing architecture, rather than being integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although they appear as a single access point to users. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.

[0820] In embodiments of this disclosure, computing device 3300 may be used to implement video encoding / decoding. Memory 3320 may include one or more video codec modules 3325 having one or more program instructions. These modules are accessible and executable by processing unit 3310 to perform the functions of the various embodiments described herein.

[0821] In an example embodiment of performing video encoding, input device 3350 may receive video data as input 3370 to be encoded. The video data may be processed, for example, by video codec module 3325 to generate an encoded bitstream. The encoded bitstream may be provided as output 3380 via output device 3360.

[0822] In an example embodiment of performing video decoding, input device 3350 may receive an encoded bitstream as input 3370. The encoded bitstream may be processed, for example, by video codec module 3325 to generate decoded video data. The decoded video data may be provided as output 3380 via output device 3360.

[0823] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.

Claims

1. A method for video processing, comprising: For the conversion between video units and the bitstream of the video, one or more parameters of the cross-component residual model (CCRM) of the video unit are determined to be inherited from the previous filter-based codec block, wherein the CCRM is a filter model that includes a cross-component model or a same-component model; and The conversion is performed based on the CCRM.

2. The method of claim 1, wherein the cross-component model includes a cross-component residual model or a cross-component prediction model for intra-frame or inter-frame use, wherein model generation and model application are based on the relationship between different color components.

3. The method according to claim 1, wherein the same component model comprises one of the following: an inter-frame filter, an intra-block copy (IBC) filter, an IBC local illumination compensation (IBC-LIC) filter, an intra-template matching prediction (IntraTMP) filter, and an extended information filter (EIF) filter, wherein the model generation and model application are based on the relationship in the same component.

4. The method according to claim 1, wherein the video unit is encoded and decoded using a type of filter model inheritance mode or a type of filter model merge mode.

5. The method of claim 1, wherein at least one syntax element is transmitted via signal at the video unit level to specify whether and / or how to use the CCRM model inheritance mode.

6. The method of claim 5, wherein the CCRM model inheritance mode includes at least one of the following: CCRMMerge mode, cross-component prediction (CCP) Merge mode, IBC filter Merge mode, or IntraTMP filter Merge mode.

7. The method of claim 5, wherein an indicator specifying whether the video unit uses a conventional CCRM mode or a CCRM model inheritance mode is transmitted via signaling based on conditions associated with at least one type of CCRM used for the video unit.

8. The method of claim 7, wherein the at least one type of CCRM includes a conventional CCRM pattern.

9. The method of claim 8, wherein the first syntax is transmitted by signal to indicate that the video unit uses a type of CCRM mode, and the second syntax is further transmitted by signal to indicate which type of CCRM mode is used.

10. The method of claim 9, wherein if at least one available CCRM candidate exists, the second syntax is transmitted via signaling.

11. The method of claim 9, wherein the second syntax is transmitted via signaling if at least one valid CCRM Merge candidate exists.

12. The method of claim 5, wherein if the CCRM model inheritance mode is used, another syntax is further transmitted via signaling to specify which CCRM model candidate is selected to be inherited.

13. The method of claim 12, wherein the candidate index is encoded / decoded to indicate CCRM candidates from the candidate list.

14. The method of claim 12, wherein the candidate index is encoded or decoded using a rounding Rice (TR) binarization process, or a rounding binary (TB) binarization process, or a k-th order Exp-Golomb (EGk) binarization process, or a fixed length (FL) binarization process.

15. The method of claim 12, wherein the maximum allowed CCRM candidate is specified in the codec.

16. The method of claim 15, wherein the maximum allowed CCRM candidate is the maximum length of the candidate list.

17. The method of claim 5, wherein the at least one syntax element depends on the residual or coefficient of the current luma block being transmitted via signal.

18. The method of claim 17, wherein whether the at least one syntax element is transmitted via signal is based on the presence of a residual or non-zero coefficient in the current luma block.

19. The method of claim 17, wherein whether the syntax element is transmitted via signaling is based on at least one of the distribution, number, or value of residuals in the current luma block, or Whether the syntax element is transmitted via signal can be based on at least one of the distribution, number, or value of non-zero coefficients in the current luminance block.

20. The method of claim 5, wherein the at least one syntax element is transmitted via a neighbor-block-based prediction method.

21. The method of claim 20, wherein the at least one syntax element is transmitted via signaling based on whether the neighboring block uses CCRM mode or CCRM Merge mode.

22. The method of claim 5, wherein the context model of the at least one syntax element depends on the encoding / decoding information of the neighboring block or the current block.

23. The method of claim 21 or 22, wherein the neighboring block includes at least one of the left neighboring block or the upper neighboring block.

24. The method of claim 22, wherein the context model is derived based on whether the neighboring block uses at least one of the following: CCRM mode, multi-mode CCRM (MM-CCRM) mode, multi-filter CCRM (MF-CCRM) mode, or CCRMMerge mode.

25. The method of claim 22, wherein the context model is derived based on whether the block dimension of the current block satisfies a specific condition.

26. The method of claim 25, wherein a specified context model is used if the dimension of the current block is W>a*H and / or H>b*W, wherein W and H represent the width and height of the current block, respectively, and a and b are predefined constants.

27. The method according to claim 26, wherein a = b = 2.

28. The method of claim 1, wherein if the CCRM model inheritance pattern is used, a CCRM model candidate list is generated.

29. The method of claim 28, wherein the CCRM model candidate is obtained based on previously encoded and decoded CCRM blocks.

30. The method of claim 29, wherein the candidate insertion order follows a predefined rule.

31. The method according to claim 30, wherein the candidate insertion order is the following sequence: spatially adjacent, spatially non-adjacent, history, shift, default.

32. The method of claim 29, wherein the time-domain candidate is derived based on motion displacement.

33. The method of claim 32, wherein the temporal candidate of one of the IBC-encoded blocks, IntraTMP-encoded blocks, or inter-frame-encoded blocks is derived based on the motion vector (MV) or block vector (BV) of neighboring blocks.

34. The method of claim 32, wherein the temporal candidate of one of the IBC-encoded, IntraTMP-encoded, or inter-frame-encoded blocks is derived based on the MV or BV of the current block.

35. The method of claim 1, wherein the model offset is not stored in the cache.

36. The method of claim 35, wherein the model offset of neighboring blocks is not reused in the current block or inherited for the current block.

37. The method of claim 35, wherein the model offset of the current block is recalculated based on available samples.

38. The method of claim 1, wherein the filtering model information is stored in a cache, and wherein the filter coefficients are calculated based on the relationship between samples in the same component.

39. The method of claim 38, wherein the filtering model information includes at least one of the following: model coefficients, model taps, model parameters, model offset of the Y component, and / or All training samples are located in the brightness component domain.

40. The method of claim 1, wherein the CCRM model inheritance pattern is used for at least one of the following: single tree or double tree.

41. The method of claim 1, wherein the CCRM model inheritance mode is used in inter-frame stripes.

42. The method of claim 41, wherein the inter-frame stripe is a B-strip or a P-strip.

43. The method of claim 1, wherein the CCRM model inheritance mode is used in intra-frame stripes.

44. The method of claim 43, wherein the intra-frame stripe is an I-strip.

45. The method of claim 1, wherein the training samples or reference samples are predicted samples in the training region or reference region.

46. ​​The method of claim 1, wherein the training samples or reference samples are reconstructed samples in the training region or reference region.

47. The method of any one of claims 1 to 46, wherein an indication of whether one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, is indicated at one of the following: sequence level, Image group level, Image quality, strip level, or Film series level.

48. The method of any one of claims 1 to 46, wherein an indication of whether one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.

49. The method of any one of claims 1 to 46, wherein an indication of whether it is determined that one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, is included in one of the following: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Virtual Pipeline Data Unit (VPDU). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or A region containing more than one sample point or pixel.

50. The method according to any one of claims 1 to 46, further comprising: Based on the encoded and decoded information of the target block, determine whether one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block and / or how to determine whether one or more parameters of the CCRM of the video unit are inherited from the previous filter-based codec block, wherein the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.

51. A method for video processing, comprising: For the conversion between video units and the bitstream of the video, the final prediction of the video unit is determined based on a weighted sum of fusion assumptions, wherein at least one assumption is based on cross-component prediction (CCP) prediction. as well as The transformation is performed based on the final prediction.

52. The method of claim 51, wherein the CCP comprises at least one of the following: cross-component residual model (CCRM), CCRM Merge, CCP Merge, cross-component linear model (CCLM), convolution cross-component model (CCCM), gradient linear model (GLM), or linear model (LM).

53. The method of claim 51, wherein the final prediction of the video unit is generated based on a plurality of prediction candidates from different CCPs.

54. The method of claim 53, wherein multiple CCP-based predictions are fused together.

55. The method of claim 51, wherein multiple CCP models are derived to obtain fused predictions.

56. The method of claim 55, wherein the final prediction is a weighted sum of the first prediction and the second prediction, and the first prediction is used to derive the first CCP model, and the second prediction is used to derive the second CCP model.

57. The method of claim 56, wherein P0 in luminance and chrominance is used to derive CCP model M0, P1 in luminance and chrominance is used to derive CCP model M1, and the final chrominance prediction is derived as wc0×Pc0+wc1×Pc1, wherein Pc0 and Pc1 represent chrominance predictions obtained using M0 and M1, respectively, and wc0 and wc1 represent weighting factors.

58. The method of claim 55, wherein predictions based on the two hypotheses are generated based on the first two candidates in the CCP Merge pattern list, and a final prediction is derived based on a weighted sum of the predictions based on the two hypotheses.

59. The method of claim 51, wherein the chromaticity prediction obtained by CCP is fused with other predictions.

60. The method of claim 59, wherein the chromaticity prediction obtained by CCP is fused with at least one of the following: Intra-frame prediction of angles, Cross-component linear model (CCLM) prediction, or Convolutional Cross-Component Model (CCCM) Prediction.

61. The method of claim 59, wherein the chromaticity prediction obtained by CCP is fused with the original prediction without CCP.

62. The method of claim 59, wherein the original prediction comprises at least one of the following: intra-frame prediction, inter-frame prediction, or intra-block copy (IBC) prediction.

63. The method of claim 61, wherein the final chromaticity prediction is derived as (w0×P0+w1×P1+offset)>>shift, where P0 represents the chromaticity prediction obtained through CCP and P1 represents the chromaticity prediction without CCP, w0 and w1 are fusion weights, and shift represents the parameter.

64. The method of claim 51, wherein the determination of the final prediction is applied to fuse more than one hypothesis, wherein at least one hypothesis is a prediction based on a filtering model.

65. The method of claim 64, wherein the filter coefficients are calculated based on the relationship between two sets of samples in the same component domain.

66. The method of claim 65, wherein one set comprises luminance samples adjacent to the current block and the other set comprises luminance samples adjacent to the reference block.

67. The method of claim 64, wherein the prediction based on the filtering model is a prediction generated by applying the filter to the MC-compensated video unit.

68. The method of claim 67, wherein the filter is based on at least one of: an IBC filter, an intra-frame TMP filter, a local illumination compensation (LIC) for inter-frame, a LIC for IBC, or an extended information filter (EIF).

69. The method of claim 64, wherein the prediction based on the filter model is generated based on either a filter-based Merge pattern or a filter-based Inheritance pattern.

70. The method of claim 69, wherein the filter-based merge mode comprises at least one of the following: a merge mode based on an IBC filter, a merge mode based on an intra-frame TMP filter, a merge mode based on an inter-frame LIC, a merge mode based on an IBC LIC, a merge mode based on an EIF, or The filter-based inheritance mode includes at least one of the following: IBC filter-based inheritance mode, intra-frame TMP filter-based inheritance mode, inter-frame LIC-based inheritance mode, IBC LIC-based inheritance mode, and EIF-based inheritance mode.

71. The method of claim 70, wherein the filter-based Merge pattern includes a pattern in which the filter model inherits / derives from the candidate model.

72. The method of claim 64, wherein the prediction based on the filtering model is fused with another prediction.

73. The method of claim 72, wherein the prediction based on the filtering model is encoded or decoded using a first prediction type, and the other prediction is encoded or decoded using a second prediction type.

74. The method of claim 73, wherein the second prediction type is not a prediction based on a filtering model.

75. The method of claim 73, wherein the second prediction type is a prediction based on a filtering model, and the first prediction type is different from the second prediction type.

76. The method of claim 73, wherein the second prediction type is a prediction based on a filtering model, and the first prediction type is the same as the second prediction type.

77. The method of claim 76, wherein the predictions of the two hypotheses are generated based on the first two candidates in a filter-based Merge pattern list, and the final prediction is derived based on a weighted sum of the predictions of the two hypotheses.

78. The method of claim 51, wherein the CCRM model inheritance pattern is used for at least one of the following: single tree or double tree.

79. The method of claim 51, wherein the CCRM model inheritance mode is used in inter-frame stripes, or The CCRM model inheritance mode mentioned above is used in intra-frame stripes.

80. The method of claim 79, wherein the inter-frame stripe is a B-strip or a P-strip.

81. The method of claim 79, wherein the intra-frame stripe is an I-strip.

82. The method of claim 51, wherein the training samples or reference samples are predicted samples in the training region or reference region.

83. The method of claim 51, wherein the training samples or reference samples are reconstructed samples in the training region or reference region.

84. The method of any one of claims 51 to 83, wherein an indication of whether the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, and / or how the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, is indicated at one of the following: sequence level, Image group level, Image quality, strip level, or Film series level.

85. The method of any one of claims 51 to 83, wherein an indication of whether the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, and / or how the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.

86. The method of any one of claims 51 to 83, wherein an indication of whether the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, and / or how the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, is included in one of the following: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Virtual Pipeline Data Unit (VPDU). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or A region containing more than one sample point or pixel.

87. The method according to any one of claims 51 to 83, further comprising: Based on the encoded and decoded information of the target block, it is determined whether the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, and / or how the final prediction of the video unit is determined based on the weighted sum of the fusion assumption, wherein the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.

88. A method for video processing, comprising: For the conversion between video units and the bitstream of the video, one or more cross-component residual models (CCRMs) are determined for the video units, wherein the one or more CCRMs include at least one of the following: multi-mode CCRM (MM-CCRM) mode, multi-filter (MF) CCRM (MF-CCRM) mode, or CCRM Merge mode; and The conversion is performed based on one or more CCRMs.

89. The method of claim 88, wherein MM-CCRM or MF-CCRM or CCRM Merge is applied based on one of TU level, PU level or CU level.

90. The method of claim 89, wherein the MM-CCRM or the MF-CCRM or the CCRM Merge is applied based on one of TU, ​​CU or PU.

91. The method of claim 90, wherein for the application of the MM-CCRM, the TU is not divided into sub-blocks, and / or In the application of the MM-CCRM, the CU is not divided into sub-blocks, and / or In the application of the MM-CCRM, the PU is not divided into sub-blocks.

92. The method of claim 89, wherein whether to use one of the following is determined at one of the TU level, PU level, or CU level: TU-based multi-model CCRM, PU-based multi-model CCRM, CU-based multi-model CCRM, TU-based MF-CCRM, PU-based MF-CCRM, CU-based MF-CCRM, TU-based CCRM Merge, PU-based CCRM Merge, or CU-based CCRM Merge.

93. The method of claim 92, wherein the video unit decides to use a sub-block-based CCRM or one of the following: a TU-based CCRM, a PU-based CCRM, or a CU-based CCRM.

94. The method of claim 93, wherein sub-block-based CCRM comprises one of the following: sub-block-based MM-CCRM, sub-block-based MF-CCRM, or sub-block-based CCRM Merge, or CCRM based on TU includes one of the following: MF-CCRM based on TU, MM-CCRM based on TU, or CCRM Merge based on TU, or Among them, PU-based CCRM includes one of the following: PU-based MF-CCRM, PU-based MM-CCRM, or PU-based CCRM Merge, or The CU-based CCRM includes one of the following: CU-based MF-CCRM, CU-based MM-CCRM, or CU-based CCRM Merge.

95. The method of claim 89, wherein whether and / or how at least one of the following is applied is derived based on encoding / decoding information at both the encoder side and the decoder side: MM-CCRM, CCRM, MF-CCRM, or CCRM Merge.

96. The method of claim 95, wherein the determination of whether to use any of the following is based on implicit derivation of the encoding / decoding information: TU-based multi-model CCRM, PU-based multi-model CCRM, CU-based multi-model CCRM, TU-based MF-CCRM, PU-based MF-CCRM, CU-based MF-CCRM, TU-based CCRM Merge, PU-based CCRM Merge, or CU-based CCRM Merge.

97. The method of claim 96, wherein sub-block-based CCRM comprises one of the following: sub-block-based MM-CCRM, sub-block-based MF-CCRM, or sub-block-based CCRM Merge, or TU-level CCRM includes one of the following: TU-level MF-CCRM, TU-level MM-CCRM, or TU-level CCRM Merge, or The PU level CCRM includes one of the following: PU level MF-CCRM, PU level MM-CCRM, or PU level CCRM Merge, or The CU-level CCRM includes one of the following: CU-based MF-CCRM, CU-level MM-CCRM, or CU-level CCRMMerge.

98. The method of claim 95, wherein the determination of whether to use MM-CCRM, MF-CCRM, or CCRM Merge based on M1xM2 sub-blocks or MM-CCRM, MF-CCRM, or CCRM Merge based on N1xN2 sub-blocks is implicitly derived based on encoding / decoding information.

99. The method of claim 98, wherein M1 = 16 or 8 or 32 or TU or CU or PU, and / or Where M2 = 16 or 8 or 32 or TU or CU or PU, and / or Where N1 = 16 or 8 or 32 or TU or CU or PU, and / or Where N2 = 16 or 8 or 32 or TU or CU or PU, and / or Where M1 is not equal to N1 and / or M2 is not equal to N2.

100. The method of claim 95, wherein the determination of whether to use a single-mode CCRM, MF-CCRM, CCRMMerge, multi-mode CCRM, MF-CCRM, or CCRM Merge is derived based on encoding / decoding information.

101. The method of claim 88, wherein at least one of the following is applied and / or how at least one of the following is applied is transmitted by signal in the bitstream: MM-CCRM, CCRM, MF-CCRM, or CCRM Merge.

102. The method of claim 101, wherein if the video unit is CCRM encoded / decoded, the syntax element is further signaled to indicate: Is the CCRM described as MM-CCRM, or Is the CCRM described as MF-CCRM, or Is the CCRM a CCRM Merge? 103. The method of claim 101, wherein syntax elements are transmitted via signals to indicate whether the CCRM is based on sub-block MM-CCRM or one of the following: TU-based MM-CCRM CP-based MM-CCRM PU-based MM-CCRM MF-CCRM based on TU, MF-CCRM based on CP PU-based MF-CCRM CCRM Merge based on TU CCRM Merge based on CP, or PU-based CCRM Merge.

104. The method according to claim 101, wherein the syntax element is signaled to indicate whether it is sub-block-based CCRM or one of the following: TU-based MM-CCRM, CP-based MM-CCRM, PU-based MM-CCRM, TU-based MF-CCRM, CP-based MF-CCRM, PU-based MF-CCRM, TU-based CCRM Merge, CP-based CCRM Merge, or PU-based CCRM Merge.

105. The method according to claim 101, wherein the syntax element is signaled based on a condition regarding block dimensions.

106. The method according to claim 105, wherein the block dimension includes at least one of width or height.

107. The method according to claim 106, wherein if W*H < T, the syntax element is not signaled, where W represents the width, H represents the height, and T represents a threshold.

108. The method according to claim 107, wherein T = 16 or 32.

109. The method according to claim 101, wherein the syntax element is signaled depending on the residual or coefficients of the current luminance block.

110. The method according to claim 109, wherein whether the syntax element is signaled is based on the presence of a residual or non-zero coefficients in the current luminance block.

111. The method according to claim 109, wherein whether the syntax element is signaled is based on the distribution or number or value of the residuals in the current luminance block, or wherein whether the syntax element is signaled is based on the distribution or number or value of the non-zero coefficients in the current luminance block.

112. The method according to claim 101, wherein the syntax element is signaled based on the prediction method of neighboring blocks.

113. The method according to claim 112, wherein the syntax element is signaled based on whether the neighboring blocks use at least one of the following: CCRM, MM-CCRM, MF-CCRM, or CCRM Merge mode.

114. The method according to claim 101, wherein the context model of the syntax element depends on the coding and decoding information of neighboring blocks or the current block.

115. The method according to claim 114, wherein the context model is derived based on whether the neighboring blocks use CCRM or MM-CCRM or MF-CCRM or CCRM Merge mode.

116. The method according to claim 114, wherein the context model is derived based on whether the block dimension of the current block satisfies a condition.

117. The method according to claim 116, wherein if the dimension of the current block is W > a*H and / or H > b*W, a specified context model is used, where W and H respectively represent the width and height of the current block, and a and b are predefined constants.

118. The method according to claim 117, wherein a=b=2.

119. The method of claim 88, wherein the CCRM model inheritance pattern is used for at least one of the following: single tree or double tree.

120. The method of claim 88, wherein the CCRM model inheritance mode is used in inter-frame stripes, or The CCRM model inheritance mode mentioned above is used in intra-frame stripes.

121. The method of claim 120, wherein the inter-frame stripe is a B-strip or a P-strip.

122. The method of claim 120, wherein the intra-frame stripe is an I-strip.

123. The method of claim 88, wherein the training samples or reference samples are predicted samples in the training region or reference region.

124. The method of claim 88, wherein the training samples or reference samples are reconstructed samples in the training region or reference region.

125. The method of any one of claims 88 to 124, wherein an indication of whether it is determined that one or more parameters of the CCRM of the video unit are inherited from a previously CCRM-encoded block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previously CCRM-encoded block, is indicated at one of the following: sequence level, Image group level, Image quality, strip level, or Film series level.

126. The method of any one of claims 88 to 124, wherein an indication of whether it is determined that one or more parameters of the CCRM of the video unit are inherited from a previously CCRM-encoded block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previously CCRM-encoded block, is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.

127. The method of any one of claims 88 to 124, wherein an indication of whether it is determined that one or more parameters of the CCRM of the video unit are inherited from a previously CCRM-encoded block, and / or how to determine that one or more parameters of the CCRM of the video unit are inherited from the previously CCRM-encoded block, is included in one of the following: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Virtual Pipeline Data Unit (VPDU). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or A region containing more than one sample point or pixel.

128. The method according to any one of claims 88 to 124, further comprising: Based on the encoded / decoded information of the target block, determine whether one or more parameters of the CCRM of the video unit are inherited from the previously CCRM-encoded block, and / or how to determine whether one or more parameters of the CCRM of the video unit are inherited from the previously CCRM-encoded block, wherein the encoded / decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.

129. The method according to any one of claims 1 to 128, wherein the conversion comprises encoding the video unit into the bitstream.

130. The method according to any one of claims 1 to 128, wherein the conversion comprises decoding the video unit from the bitstream.

131. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 130.

132. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 130.

133. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: One or more parameters of the cross-component residual model (CCRM) of the video unit of the video are inherited from the previous filter-based codec block, wherein the CCRM is a filter model that includes a cross-component model or a same-component model; and The bitstream of the video unit is generated based on the CCRM.

134. A method for storing a bitstream of video, comprising: One or more parameters of the cross-component residual model (CCRM) of the video unit of the video are inherited from the previous filter-based codec block, wherein the CCRM is a filter model that includes a cross-component model or a same-component model. The bitstream of the video unit is generated based on the CCRM; as well as The bitstream is stored in a non-transitory computer-readable recording medium.

135. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: The final prediction of the video units of the video is determined by a weighted sum based on fusion assumptions, wherein at least one of the assumptions is based on cross-component prediction (CCP) prediction. as well as The bitstream of the video unit is generated based on the final prediction.

136. A method for storing a bitstream of video, comprising: The final prediction of the video units of the video is determined by a weighted sum based on fusion assumptions, wherein at least one of the assumptions is based on cross-component prediction (CCP) prediction. The bitstream of the video unit is generated based on the final prediction; as well as The bitstream is stored in a non-transitory computer-readable recording medium.

137. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Determine one or more cross-component residual models (CCRMs) for the video units of the video, wherein the one or more CCRMs include at least one of the following: multi-mode CCRM (MM-CCRM) mode, multi-filter (MF) CCRM (MF-CCRM) mode, or CCRM Merge mode; and The bitstream of the video unit is generated based on the one or more CCRMs.

138. A method for storing a bitstream of video, comprising: Determine one or more cross-component residual models (CCRMs) for the video unit of the video, wherein the one or more CCRMs include at least one of the following: multi-mode CCRM (MM-CCRM) mode, multi-filter (MF) CCRM (MF-CCRM) mode, or CCRM Merge mode; The bitstream of the video unit is generated based on the one or more CCRMs; and The bitstream is stored in a non-transitory computer-readable recording medium.