Methods, apparatuses, and media for video processing
By employing cross-component prediction model derivation technology in video encoding and decoding, and utilizing the CCP candidate list for inter-frame and intra-frame chroma encoding and decoding, the problem of insufficient encoding and decoding efficiency in existing technologies is solved, achieving more efficient video encoding and decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-10-10
- Publication Date
- 2026-05-29
AI Technical Summary
The efficiency of existing video encoding and decoding technologies needs to be further improved, especially in inter-frame and intra-frame chroma encoding and decoding.
The cross-component prediction (CCP) model derivation technique is adopted. By converting between the current block of the video and the bitstream, the inter-frame CCP mode or intra-frame CCP mode is applied, and the CCP candidate list is used for encoding and decoding.
It improves the efficiency and performance of video encoding and decoding, especially in inter-frame and intra-frame chroma encoding and decoding.
Smart Images

Figure CN122122900A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to the derivation of cross-component prediction (CCP) models for inter-frame and intra-frame chroma coding and decoding. Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-TH.265 High Efficiency Video Coding (HEVC) standard, and Versatile Video Coding (VVC) standard. However, the encoding and decoding efficiency of video encoding and decoding technologies is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for a conversion between a current block of video and a bitstream of video, deriving at least one of filters or models based on decoding information; applying at least one of the filters or models to the encoding / decoding of at least one of the current block or future blocks; and performing the conversion based on at least one of the filters or models. Compared to conventional solutions, the method according to the first aspect of this disclosure advantageously improves encoding / decoding efficiency and performance by deriving and applying filters and / or models.
[0005] In a second aspect, another method for video processing is proposed. This method includes: for a conversion between the current block of the video and the video bitstream, applying at least one of an inter-frame CCP mode or an intra-frame CCP mode based on a Cross-Component Prediction (CCP) candidate list, wherein the CCP candidate list is derived based on the decoder; and performing the conversion based on at least one of the inter-frame CCP mode or the intra-frame CCP mode. Compared to conventional solutions, the method according to the second aspect of this disclosure advantageously improves encoding / decoding efficiency and performance by applying an inter-frame CCP mode or an intra-frame CCP mode based on a CCP candidate list.
[0006] In a third aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform a method according to the first or second aspect of this disclosure.
[0007] In a fourth aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute a method according to the first or second aspect of this disclosure.
[0008] In a fifth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by a video processing apparatus. The method includes: deriving at least one of filters or models based on decoding information; applying at least one of the filters or models to at least one of the current or future blocks of the video for encoding and decoding; and generating a bitstream based on at least one of the filters or models.
[0009] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by a video processing apparatus. The method includes: applying at least one of an inter-frame CCP mode or an intra-frame CCP mode based on a Cross-Component Prediction (CCP) candidate list, wherein the CCP candidate list is derived based on a decoder; and generating a bitstream based on at least one of the inter-frame CCP mode or the intra-frame CCP mode.
[0010] In a seventh aspect, a method for storing a bitstream of video is proposed. The method includes: deriving at least one of filters or models based on decoding information; applying at least one of the filters or models to encoding / decoding at least one of the current or future blocks of the video; generating a bitstream based on at least one of the filters or models; and storing the bitstream in a non-transitory computer-readable recording medium.
[0011] In the eighth aspect, a method for storing video bitstreams is proposed. The method includes: applying at least one of inter-frame CCP modes or intra-frame CCP modes based on a Cross-Component Prediction (CCP) candidate list, wherein the CCP candidate list is derived based on a decoder; generating a bitstream based on at least one of the inter-frame CCP modes or intra-frame CCP modes; and storing the bitstream in a non-transitory computer-readable recording medium.
[0012] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0013] The above and other objects, features and advantages of exemplary embodiments of the present disclosure will become clearer from the following detailed description with reference to the accompanying drawings, in which the same reference numerals generally refer to the same parts.
[0014] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 The diagram illustrates the effect of the slope adjustment parameter "u", with the model created using the current CCLM on the left and the updated model as proposed on the right. Figure 5 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown. Figure 6 The neighbor reconstructed sample points used for the DIMD chromaticity mode are shown; Figure 7 The intra-frame template matching search area used is shown; Figure 8 The use of the IntraTMP block vector for IBC blocks is shown; Figure 9A and Figure 9B The method for dividing angle patterns is shown; Figure 10 An expanded list of MRL candidates is shown; Figure 11 An illustration of the template area is shown; Figure 12 The spatial portion of the convolution filter is shown; Figure 13 The reference region (and its padding) used to derive the filter coefficients is shown. Figure 14 Four Sobel-based gradient modes for GLM are shown; Figure 15 Candidates for the spatial GPM are shown; Figure 16The GPM template is shown; Figure 17 GPM mixing is shown; Figure 18 The possible locations of the candidate regions are shown; Figure 19 The locations of adjacent airspace candidates are shown; Figure 20 The transformation selection process for the directional plane mode is shown; Figure 21 A luminance block is shown for deriving the direct block vector; Figure 22 The diagram shows three types of reconstructed regions, comprising thirteen columns or rows of reconstructed pixels. Figure 23 The three types of filter shapes defined are shown, each with fifteen inputs and producing one output. Figures 24A-24C Examples of predictions for different positions within the current block are shown, where... Figure 21 In A, all inputs to EIP are reconstructed samples. Figure 21 In B, part of the input is reconstructed samples and part of the input is predicted samples, and... Figure 21 In C, all inputs to EIP are prediction samples; Figure 25 The proposed method is shown on the decoder; Figure 26 The luminance samples L0, ..., L5 are shown relative to the chromaticity sample C; Figure 27 The reference area of BVG-CCCM is shown; Figure 28 The spatial portion of the convolution filter is shown; Figure 29 The location used for block vector derivation from the co-position brightness block is shown; Figure 30 An example of the current template and reference template involved in CCRM encoding and decoding for the current inter-frame block is shown; Figure 31 Examples of the current template and reference template involved in CCRM encoding and decoding for the current IBC block are shown; Figure 32 An example of CCP model calculation for template samples is shown, which requires access to sample values in the current luma block; Figure 33 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 34 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and Figure 35 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0015] In all the accompanying drawings, the same or similar reference numerals usually indicate the same or similar elements. Detailed Implementation
[0016] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0017] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0018] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0019] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” and / or “having” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0021] Example Environment Figure 1This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0022] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0023] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0024] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0025] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or further standards.
[0026] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0027] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0028] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0029] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0030] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.
[0031] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0032] The mode selection unit 203 can select one of several encoding / decoding modes (intra-frame encoding / decoding or inter-frame encoding / decoding) based, for example, on the error result, and provide the resulting intra-coded or inter-coded blocks to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded blocks for use as reference images. In some examples, the mode selection unit 203 can select an intra-inter-frame joint prediction (CIIP) mode, in which prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0033] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0034] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.
[0035] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0036] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images in lists 0 and 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0037] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0038] In one example, the motion estimation unit 204 may indicate a value to the video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0039] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0040] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0041] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0042] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0043] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0044] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0045] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0046] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transformed coefficient video block, respectively, to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.
[0047] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0048] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0049] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0050] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0051] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0052] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.
[0053] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. Identifiers for interpolation filters used at sub-pixel precision can be included in the syntax elements.
[0054] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.
[0055] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be the entire image or a region of the image.
[0056] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0057] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be used to filter the decoded block to eliminate block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0058] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, although some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding and decoding steps in detail, it will be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding / decoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or a different compression bitrate.
[0059] 1. Brief Overview This disclosure relates to video encoding and decoding technologies. Specifically, this disclosure pertains to cross-component encoding and decoding in image / video encoding and decoding. This disclosure can be applied to existing video encoding and decoding standards such as HEVC, VVC, etc. It can also be applied to future video encoding and decoding standards or video codecs.
[0060] 2 Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established in 2015 by VCEG and MPEG. JVET meetings are held quarterly, and the new video codec standard was officially named Multifunctional Video Coding (VVC) at the JVET meeting in April 2018, with the first version of the VVC Test Model (VTM) also released at that time. The VVC working draft and test model VTM are updated after each meeting. The VVC project achieved Technical Completion (FDIS) at the July 2020 meeting.
[0061] 2.1 Intra-frame prediction In intra-frame prediction, the minimum chroma intra-frame prediction unit (SCIPU) constraint in the VVC is removed. Additionally, the VPDU constraint used to reduce CCLM prediction latency is also removed.
[0062] 2.1.1 Multi-Model Learning (MMLM) The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, using a threshold as the average of the luminance reconstruction neighboring samples, the reconstructed neighboring samples are classified into two categories. The linear model for each category is derived using the least mean square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model. Slope adjustment is applied to both the cross-component linear model (CCLM) and multi-model LM predictions. This adjustment is a linear function that maps luminance values to chrominance values, tilted relative to a center point determined by the average luminance values of the reference samples.
[0063] 2.1.1.1 Slope Adjustment of CCLM CCLM uses a two-parameter model to map luminance values to chrominance values. The slope parameter "a" and the bias parameter "b" are defined as follows: .
[0064] The adjustment of the slope parameter "u" is transmitted via signal to update the model in the following form:
[0065] in a' = a + u .
[0066] By making this selection, the mapping function revolves around a value with brightness y. r The points are tilted or rotated. The average value of the reference brightness samples used in model creation is used as y. r This is to allow for meaningful modifications to the model. The image below illustrates this process.
[0067] Figure 4 The diagram illustrates the effect of the slope adjustment parameter "u". Left: Model created using the current CCLM. Right: Updated model as proposed.
[0068] Implementation The slope adjustment parameter is provided as an integer between -4 and 4 (including boundary values) and is transmitted via signal in the bitstream. The unit of the slope adjustment parameter is 1 / 8 of the chroma sample value per luminance sample value (for 10-bit content).
[0069] The CCLM model (“LM_CHROMA_IDX” and “MMLM_CHROMA_IDX”) can be adjusted to use reference samples from both the top and left sides of the block, but not for “one-sided” mode. This choice is based on a trade-off between encoding / decoding efficiency and complexity.
[0070] When slope adjustment is applied to a multi-mode CCLM model, both models can be adjusted, so that up to two slope updates are transmitted via signal for a single chroma block.
[0071] Encoder method The proposed encoder method performs a SATD-based search for the optimal slope update for Cr and a similar SATD-based search for Cb. If either results in a non-zero slope adjustment parameter, the combined slope adjustment pair (SATD-based update for Cr, SATD-based update for Cb) is included in the list of RD checks for TU.
[0072] 2.1.2 Gradient PDPC In VVC, PDPC may not be applied in some scenarios due to the unavailability of secondary reference samples. In these cases, gradient-based PDPC (extended from the horizontal / vertical mode) is applied. The PDPC weights (wT / wL) and the nScale parameter, used to determine the decay of the PDPC weights relative to the distance from the left / top boundary, are set to the corresponding parameters in the horizontal / vertical mode, respectively. Bilinear interpolation is applied when the secondary reference sample is located at the fractional sample position.
[0073] 2.1.3 Secondary MPM A secondary MPM list is introduced. The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list includes 16 entries. A general MPM list with 22 entries is first constructed, then the first 6 entries in the general MPM list are included in the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the planar mode. The remaining entries consist of the intra-mode of the left (L), top (A), bottom left (BL), top right (AR), and top left (AL) neighboring blocks, the directional mode with offsets added from the first two available directional modes of the neighboring blocks, and the default mode.
[0074] If the CU block is vertically oriented, the order of the neighboring blocks is A, L, BL, AR, AL; otherwise, it is L, A, BL, AR, AL.
[0075] Figure 5 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown.
[0076] The PMPM flag is parsed first. If it is equal to 1, the PMPM index is parsed to determine which entry in the PMPM list is selected. Otherwise, the SPMPM flag is parsed to determine whether to parse the SMPM index or the remaining patterns.
[0077] 2.1.4 Reference Sample Interpolation and Smoothing for Intra-Frame Prediction The 4-tap cubic interpolation is replaced by a 6-tap cubic interpolation filter, which is used to derive the predicted samples from the reference samples.
[0078] For reference sample filtering, a 6-tap Gaussian filter is applied to larger blocks (W>= 32 and H>= 32), otherwise the existing VVC 4-tap Gaussian interpolation filter is applied. The extended intra-frame reference sample using a 4-tap interpolation filter instead of nearest-neighbor rounding is derived.
[0079] 2.1.5 Decoder-side Intra-Frame Mode Derivation (DIMD) When DIMD is applied, two intra-frame modes are derived from reconstructed neighboring samples, and these two predictions are combined with the planar mode prediction using weights derived from the gradient. Division operations in the weight derivation are performed using an integerization scheme based on the same lookup table (LUT) used by CCLM. For example, division operations in direction calculation.
[0080] It is computed using the following LUT-based scheme:
[0081] in DivSigTable
[16] = {0, 7, 6, 5,5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}.
[0082] The derived intra-frame modes are included in the main list of most probable intra-frame modes (MPMs), so the DIMD process is performed before the MPM list is built. The main derived intra-frame modes of the DIMD block are stored with the block and used for the construction of the MPM lists of neighboring blocks.
[0083] 2.1.5.1 DIMD Chroma Mode The DIMD chroma mode uses the DIMD derivation method to derive the chroma intra-prediction mode for the current block based on the reconstructed Y, Cb, and Cr samples from the second nearest neighbor row and column. Specifically, the horizontal and vertical gradients are calculated for each co-located reconstructed luma sample and the reconstructed Cb and Cr samples of the current chroma block to construct the HoG. The intra-prediction mode with the largest histogram amplitude value is then used to perform chroma intra-prediction for the current chroma block.
[0084] Figure 6 The neighbor reconstructed samples used for the DIMD chromaticity mode are shown.
[0085] When the intra-prediction mode derived from the DIMD chroma mode is the same as the intra-prediction mode derived from the DM mode, the intra-prediction mode with the second largest histogram amplitude value is used as the DIMD chroma mode. A CU level flag is transmitted via signaling to indicate whether the proposed DIMD chroma mode is applied.
[0086] 2.1.6 Fusion of Chroma Intra-Frame Prediction Modes The DM mode and the four default modes can be merged with the MMLM_LT mode, as shown below:
[0087] in These are predicted values obtained by applying a non-LM model. These are predicted values obtained by applying the MMLM_LT mode, and This is the final predicted value for the current chroma block. Two weights. and Determined by the intra-prediction mode of adjacent chroma blocks, and It is set to equal to 2. Specifically, when the upper and left adjacent blocks are both encoded and decoded using LM mode, { }={1, 3}; When the blocks above and to the left are both encoded and decoded in non-LM mode, { }={3,1}; otherwise, { }={2, 2}.
[0088] For syntax design, if a non-LM mode is selected, a flag is transmitted via signaling to indicate whether fusion is applied. This method applies only to I-stripes.
[0089] 2.1.7 Intra-frame template matching Intra-Template Matching Prediction (IntraTMP) is a special intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. The encoder then transmits the use of this mode via signal transmission, and the same prediction operation is performed on the decoder side.
[0090] The prediction signal is obtained by comparing the L-shaped causal neighbors of the current block with... Figure 7 It is generated by matching another block in a predefined search region, which consists of the following: R1: Current CTU R2: Top left CTU, R3: Above CTU, R4: Left CTU.
[0091] The sum of absolute differences (SAD) is used as the cost function.
[0092] Within each region, the decoder searches for the template with the smallest SAD relative to the current template and uses its corresponding block as the prediction block.
[0093] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is:
[0094] in' ' is a constant that controls the trade-off between gain and complexity. In practice, ' 'Equals 5.'
[0095] Figure 7 The intra-frame template matching search area used is shown.
[0096] To accelerate the template matching process, the search range of all search regions is downsampled by a factor of 2. This results in a reduction of 4 in the template matching search. After finding the best match, a refinement process is performed. Refinement is accomplished by a second template matching search around the best match with the reduced range. The reduced range is defined as min(BlkW, BlkH) / 2.
[0097] The intra-frame template matching tool is enabled for CUs with width and height dimensions less than or equal to 64. This maximum CU size for intra-frame template matching is configurable.
[0098] When DIMD is not used in the current CU, the intra-template matching prediction mode is transmitted at the CU level via a dedicated flag.
[0099] 2.1.7.1 Block Vector Candidates Derived for IntraTMP in IBC In this method, the block vector (BV) derived from IntraTMP (Intra-Temporal Matching Prediction) is used for Intra-Block Copy (IBC). The IntraTMP BV of the stored neighboring blocks, together with the IBC BV, are used as spatial BV candidates in the construction of the IBC candidate list.
[0100] The IntraTMP block vector is stored in the IBC block vector cache, and the current IBC block can use both the IBC BV and the IntraTMP BV of neighboring blocks as BV candidates for the IBC BV candidate list, such as... Figure 8 As shown.
[0101] Figure 8 The use of the IntraTMP block vector for IBC blocks is shown.
[0102] IntraTMP block vectors are added to the IBC block vector candidate list as spatial domain candidates.
[0103] 2.1.8 Fusion for Template-Based Intra-Frame Mode Derivation (TIMD) For each intra-prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The two intra-prediction modes with the smallest SATD are selected as TIMD modes. These two TIMD modes are fused using weights after applying the PDPC procedure, and this weighted intra-prediction is used for encoding and decoding the current CU. Position-dependent intra-prediction combination (PDPC) is included in the derivation of the TIMD modes.
[0104] The costs of the two selected modes are compared with a threshold, and a cost factor of 2 is applied during the test as follows: .
[0105] If the condition is true, fusion is applied; otherwise, only mode1 is used.
[0106] The weights of the patterns are calculated from their SATD costs as follows: weight1 = costMode2 / (costMode1+ costMode2), weight2 = 1 - weight1.
[0107] Division operations are performed using an integerization scheme based on the same lookup table (LUT) used by CCLM.
[0108] 2.1.9 Intra-frame prediction fusion This intra-frame prediction method derives predicted samples as a weighted combination of multiple predicted values generated from different reference rows. In this process, multiple intra-frame predicted values are generated and then fused by a weighted average. The process of deriving the predicted values to be used in the fusion process is described below: • For the intra-angle prediction mode in the single-mode case including TIMD and DIMD, the proposed method improves upon the representation as... Intra-prediction is derived by weighting the intra-prediction obtained from multiple reference rows, where It is an intra-frame prediction from the default reference line, and This is a prediction from the row above the default reference row. The weights are set to... and .
[0109] • For TIMD modes with hybrid patterns, Used in the first mode ( ),and Used in the second mode ( ).
[0110] • For DIMD patterns with mixing, the number of predicted values selected for the weighted average is increased from 3 to 6.
[0111] When the intra-frame mode has a non-integer slope (required reference sample interpolation) and the block size is greater than 16, the intra-frame prediction fusion method is applied to the luma block, used in conjunction with MRL, but not to the ISP-encoded block. In the method studied in subtest a, PDPC is applied for the intra-frame prediction mode that uses the closest reference line to the current block.
[0112] 2.1.10 Combination of CIIP with TIMD and TM Merge In CIIP mode, prediction samples are generated by weighting the inter-prediction signal using CIIP-TM Merge candidate prediction and the intra-prediction signal using the intra-prediction mode derived using TIMD. This method is only applied to codec blocks with an area of 1024 or less.
[0113] The TIMD derivation method was used to derive intra-prediction modes in CIIP. Specifically, the intra-prediction mode with the smallest SATD value in the TIMD mode list was selected and mapped to one of 67 regular intra-prediction modes.
[0114] Furthermore, it is proposed that if the derived intra-prediction mode is an angle mode, the weights (wIntra, wInter) for the two tests should be modified. For near-horizontal mode (2 <= angle mode index < 34), the current block is vertically partitioned; for near-vertical mode (34 <= angle mode index <= 66), the current block is horizontally partitioned.
[0115] For different sub-blocks (wIntra, wInter) such as Figure 9A and Figure 9B As shown, it illustrates the method for dividing angle patterns.
[0116] Table 1. Modified weights used for angle mode
[0117] Using CIIP-TM, a CIIP-TM Merge candidate list is constructed for the CIIP-TM pattern. Merge candidates are refined through template matching. CIIP-TM Merge candidates are also reordered as regular Merge candidates using the ARMC method. The maximum number of CIIP-TM Merge candidates is 2.
[0118] 2.1.11 Extended Multi-Reference Row (MRL) List The MRL list in VVC is expanded to include more reference rows for intra-frame prediction. The expanded reference row list consists of row indices {1, 3, 5, 7, 12}. For Template-Based Intra-Frame Mode Derivation (TIMD), instead of the full MRL candidate list, only the first two reference row candidates (i.e., {1, 3}) are used.
[0119] Figure 10 An expanded list of MRL candidates is shown.
[0120] 2.1.12 Template-based multi-reference row intra-frame prediction Template-based multi-reference line intra-frame prediction (TMRL) mode combines reference lines and prediction modes, and uses template matching to construct a list of candidate combinations. The indices of the candidate combination list are encoded / decoded to indicate which reference line and prediction mode to use when encoding / decoding the current block. Regular multi-reference line (MRL) for non-TIMD portions is replaced by TMRL mode.
[0121] The TMRL mode expands the reference line candidate list and the intra-prediction mode candidate list. The expanded reference line candidate list is {1, 3, 5, 7, 12}. The restriction on the top CTU line remains unchanged. The size of the intra-prediction mode candidate list is 10. The construction of the intra-prediction mode candidate list is similar to that of MPM, except that the PLANAR mode is excluded from the intra-prediction mode candidate list, the DC mode is added after the modes of the 5 neighboring PUs and the DIMD mode (if it is not included), and has a range from... arrive An angle mode with an incremental angle (compared to existing angle modes in the intra-prediction mode candidate list) has been added.
[0122] The TMRL candidates are constructed as follows. There are 5 x 10 = 50 combinations of extended reference lines and allowed intra-prediction modes for each block. Since the extended reference lines start from reference line 1, the region covered by reference line 0 is used for template matching. For template regions (see...), ... Figure 11 The SAD cost of the prediction (generated from 50 combinations) is calculated between prediction and reconstruction. The 20 combinations with the lowest SAD cost are selected in ascending order to form the TMRL candidate list.
[0123] Figure 11 An illustration of the template area is shown.
[0124] For TMR signaling, instead of directly encoding and decoding the reference line and intra-frame mode, the index of the TMRL candidate list is encoded and decoded to indicate which combination of reference line and prediction mode is used to encode and decode the current block.
[0125] 2.1.13 Convolutional Cross-Component Intra-Frame Prediction Model In this method, a convolutional cross-component model (CCCM) is applied to predict chroma samples from reconstructed luminance samples, similar in spirit to what is done by the current CCLM model. As with CCLM, when chroma downsampling is used, the reconstructed luminance samples are downsampled to match a lower-resolution chroma grid. Similar to CCLM, top, left, or top and left reference samples are used as templates for model derivation.
[0126] In addition, similar to CCLM, there are options for single-model or multi-model variants using CCCM. The multi-model variant uses two models: one model is derived for samples above the average luminance reference value, and the other model is derived for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples.
[0127] 2.1.13.1 Convolution Filter The convolutional 7-tap filter consists of a 5-tap plus-shaped spatial component, a nonlinear term, and a bias term. The input of the filter's 5-tap spatial component consists of a center (C) luminance sample that is co-located with the chrominance sample to be predicted, and its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors, as shown below.
[0128] Figure 12 The spatial portion of the convolution filter is shown.
[0129] The nonlinear term P is expressed as the square of the center luminance sample C and scaled to the range of sample values for the content: .
[0130] That is, for 10 bits of content, it is calculated as: .
[0131] The bias term B represents the scalar offset between the input and output (similar to the offset term in CCLM) and is set to an intermediate chroma value (512 for 10-bit content).
[0132] The output of the filter is calculated as the filter coefficients c. i The convolution with the input values is then limited to the range of valid chromaticity samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B.
[0133] 2.1.13.2 Calculation of Filter Coefficients Filter coefficients c iIt is calculated by minimizing the MSE between the predicted chromaticity samples and the reconstructed chromaticity samples in the reference region. Figure 13 The reference region is shown, consisting of six rows of chroma samples above and to the left of the PU. The reference region extends to the right by one PU width and below the PU boundary by one PU height. The region is adjusted to include only available samples. The expansion of the region shown in blue is necessary to support the "side samples" of the plus-shaped spatial filter and is filled in unavailable areas.
[0134] Figure 13 The reference region (and its filling) used to derive the filter coefficients is shown.
[0135] MSE minimization is performed by calculating the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chromaticity output. The autocorrelation matrix is decomposed using LDL, and the final filter coefficients are calculated using inverse substitution. This process roughly follows the calculation of ALF filter coefficients in ECM; however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations.
[0136] The autocorrelation matrix is calculated using reconstructed values from luma and chromaticity samples. These samples are full-range (e.g., between 0 and 1023 for 10-bit content), resulting in relatively large values in the autocorrelation matrix. This requires high-bit-depth operations during model parameter calculation. A proposed solution is to remove a fixed offset from the luma and chromaticity samples in each PU for each model. This reduces the magnitude of the values used in model creation and allows for a reduction in the precision required for fixed-point arithmetic. As a result, a 16-bit decimal precision is proposed instead of the 22-bit precision of the original CCCM implementation.
[0137] For simplicity, the reference sample values immediately outside the top-left corner of the PU are used as offsets (offsetLuma, offsetCb, and offsetCr). The sample values used in both model creation and final prediction (i.e., luminance and chrominance in the reference region, and luminance in the current PU) are reduced by these fixed values, as follows: C' = C – offsetLuma N' = N – offsetLuma S' = S – offsetLuma E' = E – offsetLuma W' = W – offsetLuma P' = nonLinear(C') B = midValue = 1<<(bitDepth - 1) Furthermore, the chromaticity values are predicted using the following equations, where offsetChroma for the Cr and Cb components is equal to offsetCr and offsetCb, respectively: predChromaVal = c0C' + c1N' + c2S' + c3E' + c4W' + c5P' + c6B + offsetChroma.
[0138] To avoid any additional sample-level operations, the luminance offset is removed during luminance reference sample interpolation. This can be done, for example, by replacing the rounding term used in luminance reference sample interpolation with an updated offset that includes both the rounding term and offsetLuma. The chrominance offset can be removed by directly subtracting the chrominance offset from the reference chrominance sample. Alternatively, the effect of the chrominance offset can be removed from the cross-component vector, yielding the same result. To add the chrominance offset back to the output of the convolution prediction operation, the chrominance offset is added to the bias term of the convolution model.
[0139] The calculation of CCCM model parameters requires division operations. Division operations are not always considered implementation-friendly. Division operations are replaced by multiplication (using scaling factors) and shift operations, where the scaling factor and the number of shifts are calculated based on the denominator, similar to the method used in calculating CCLM parameters.
[0140] 2.1.13.3 Gradient Linear Model For the YUV 4:2:0 color format, the Gradient Linear Model (GLM) method can be used to predict chromaticity samples from the luminance sample gradient. Two modes are supported: two-parameter GLM mode and three-parameter GLM mode.
[0141] Compared to CCLM, two-parameter GLM uses the gradient of luminance samples to derive a linear model, rather than downsampling luminance values. Specifically, when two-parameter GLM is applied, the input to the CCLM process (i.e., downsampling luminance samples) is... Gradient of brightness sample points Replacement. Other parts of CCLM (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged.
[0142]
[0143] In a three-parameter GLM, chromaticity samples can be predicted based on both the gradient of luminance samples with different parameters and the downsampled luminance values. The model parameters of the three-parameter GLM are derived from neighboring samples in 6 rows and columns using an MSE minimization method based on LDL decomposition, as used in CCCM.
[0144]
[0145] For signaling, when CCLM mode is enabled for the current CU, a flag is transmitted via signaling to indicate whether GLM is enabled for both Cb and Cr components; if GLM is enabled, another flag is transmitted via signaling to indicate which of the two GLM modes is selected, and a syntax element is further transmitted via signaling to select one of the four gradient filters for gradient calculation.
[0146] • Enable four gradient filters for GLM, such as Figure 14 As shown, it illustrates the gradient patterns based on four Sobels for GLM.
[0147] 2.1.13.4 Bitstream Signaling The use of this mode utilizes PU-level flags encoded and decoded by CABAC, which are transmitted via signaling. A new CABAC context is included to support this. When signaling is involved, CCCM is considered a sub-mode of CCLM. That is, the CCCM flag is transmitted via signaling only when the intra-frame prediction mode is LM_CHROMA.
[0148] 2.1.14 Spatial Geometric Partitioning Model (SGPM) SGPM is an intra-frame mode of inter-frame coding / decoding tools similar to GPM, where two prediction parts are generated from the intra-frame prediction process. In this mode, a candidate list is constructed, where each entry contains a segmentation partition and two intra-frame prediction modes, such as... Figure 15 As shown, 26 segmentation modes and 3 intra-frame prediction modes were used to form a combination. The length of the candidate list was set to 16. The selected candidate indices were transmitted via signaling.
[0149] Figure 15 The candidates for the spatial GPM are shown.
[0150] List using templates ( Figure 16 The templates are reordered, with the SAD between the template's prediction and reconstruction used for sorting. The template size is fixed at 1.
[0151] Figure 16 The GPM template is shown.
[0152] For each segmentation mode, the same intra-frame to inter-frame GPM list derivation is used, and the IPM list is derived for each segment. The IPM list size is set to 3. In the list, the TIMD-derived mode is replaced by two derived modes with horizontal and vertical directions.
[0153] The SGPM pattern is applied with limited block sizes: 4 <= width <= 64, 4 <= height <= 64, width < height. 8. Height < Width 8. Width Height >= 32.
[0154] Adaptive blending has also been used in spatial GPM, where Figure 17 The mixing depth τ shown is derived as follows: • If min(width, height) == 4, then 1 / 2τ is selected. • Otherwise, if min(width, height) == 8, then τ is selected. • Otherwise, if min(width, height) == 16, then 2τ is selected. • Otherwise, if min(width, height) == 32, then 4τ is selected. • Otherwise, 8τ is selected.
[0155] Figure 17 GPM mixing is shown.
[0156] 2.1.15 Nonlocal cross-component prediction Cross-component prediction (CCP), including CCLM, CCCM, and their variants, is employed by ECM to leverage cross-component correlations. With CCLM or CCCM, training samples are always adjacent to the current block. However, the cross-component relationships of the current block can be more relevant to cross-component relationships in non-local regions.
[0157] A nonlocal cross-component prediction method is proposed to improve CCP by gaining more advantages from nonlocal regions.
[0158] Method #1: A Non-Adjacent Cross-Component Prediction (NA-CCP) model is proposed. Using the NA-CCP model, samples from regions that are not adjacent to the current block can be used to derive the CCCM model for the current block. A candidate region list with six candidates is constructed by sequentially examining potential 8×8 regions. If an examined region is available, it is added to the candidate region list. The top-left position of the potential 8×8 regions is pre-determined as {(-xStep, 0), (0, -yStep), (xStep, -yStep), (-xStep, yStep), (-xStep, -yStep), (-2...} xStep, 0), (0, -2) yStep), (-2 xStep, 2 yStep), (2) xStep, -2 yStep), (-2 xStep, yStep), (xStep, -2 yStep), (-2 xStep, -yStep), (-xStep, -2 yStep), (-2 xStep, -2 yStep), (-xStep / 2, 0), (0, -yStep / 2), (xStep / 2, -yStep / 2), (-xStep / 2, yStep / 2), (-xStep / 2, -yStep / 2)}, where xStep = Max(width, 16), yStep = Max(height, 16). Figure 18 Some possible locations of the candidate regions are shown.
[0159] A flag is transmitted via signaling to indicate whether NA-CCP is applied to the chroma block. If NA-CCP is applied, an index is transmitted via signaling to indicate which candidate in the candidate region list was used to derive the CCCM model.
[0160] Method #2: A history-based cross-component prediction (H-CCP) model is proposed. Using H-CCP, similar to the HMVP table, H-CCLM table, and H-CCCM table are maintained. After decoding a block encoded / decoded via CCLM or CCCM, the corresponding table is updated. In the H-CCP implementation, the size of the H-CCLM table or H-CCCM table is 6. If the current block is encoded / decoded in CCLM or CCCM mode, a flag is signaled to indicate whether H-CCP is applied. If H-CCP is used, an index is further signaled to indicate which candidate model from the H-CCLM table or H-CCCM table is selected.
[0161] 2.1.16 Cross-component Merge Mode for Chroma Intra-Frame Coding / Decoding Cross-component prediction (CCP) using methods including Cross-component Linear Model (CCLM), Convolutional Cross-component Model (CCCM), and Gradient Linear Model (GLM) is employed by ECM to leverage cross-component correlations. The Cross-component Merge (CCMerge) mode is proposed as a new CCP mode. The cross-component model parameters of the current chroma block encoded using CCMerge can be inherited from neighboring blocks encoded using CCP. Through CCMerge, CCP can be more efficient and has less signaling overhead.
[0162] In CCMerge, the final cross-component model parameters for the current chroma block can be inherited from its spatially adjacent and non-adjacent neighbors or the default model. A list is created that includes CCP models from spatially adjacent and non-adjacent neighbors encoded and decoded in CCLM, MMLM, CCCM, GLM, chroma blending, and CCMerge modes. After including neighboring CCP models, the default model is further included to fill any remaining empty positions in the list. To avoid including redundant CCP models in the list, a deduplication operation is applied. More details are described below.
[0163] Figure 19 The locations of adjacent airspace candidates are shown.
[0164] • Airspace adjacent neighbor candidates The positions of adjacent candidates in the airspace are as follows Figure 19 As shown, the airspace candidates are included in the following order: B1 -> A1 -> B0 -> A0 -> B2.
[0165] • Airspace not adjacent to neighboring candidates After all spatially adjacent neighbors have been checked, spatially non-adjacent neighbor candidates are considered. In the current ECM design, two sets of spatially non-adjacent neighbor candidates are obtained in inter-frame merge mode. In the proposed method, the positions and inclusion order of the spatially non-adjacent neighbor candidates from the first set are used.
[0166] • CCLM candidates with default scaling parameters If the list is not full, CCLM candidates with default scaling parameters are considered after including both spatially adjacent and non-adjacent candidates. The default scaling parameters are {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8}, and the offset parameters are derived based on the selected default scaling parameters, the average neighbor reconstructed luminance sample value (Yavg), and the average neighbor reconstructed Cb / Cr sample value (Cavg).
[0167] 2.1.16.1 Merge Model Candidates When merging CCLM candidates, only the scaling parameter is inherited. The offset parameter is derived using the inherited scaling parameter, Yavg, and Cavg.
[0168] When merging MMLM candidates, the scaling parameters and classification thresholds are inherited. The offset parameters in each class are derived based on the inherited classification thresholds and the Yavg and Cavg values for each class. If no neighboring reconstructed samples are available in a class, the offset parameters are directly inherited from the candidate.
[0169] When merging CCCM candidates, all convolutional parameters, offsets (i.e., offsetLuma, offsetCb, and offsetCr), and classification thresholds are inherited.
[0170] When merging GLM candidates, if the GLM candidate is a 3-parameter GLM mode, all gradient mode indices and model parameters are inherited; otherwise, if the GLM candidate is a 2-parameter GLM mode, the offset parameters are derived by using the inherited scaling parameters, Yavg, and Cavg.
[0171] When merging chroma blending candidates, the derived MMLM parameters are inherited and used as merging MMLM candidates.
[0172] For a CCMerge block, if its merge candidate mode is CCLM, MMLM, CCCM, or GLM, the merge candidate mode is stored as the propagation mode of the current chroma block; otherwise, if its merge candidate mode is chroma blending, the propagation mode is set to MMLM. How CCP parameters are inherited or derived when merging CCMerge candidates depends on the propagation mode of the CCMerge candidate, as described in the five paragraphs above.
[0173] 2.1.16.2 Signaling Following the `cclm_mode_flag` syntax element, an additional flag is signaled indicating whether CCMerge is used. If CCMerge is used, candidate indices are additionally signaled. The signaled candidate indices are shared for the Cb / Cr color components. Currently, the maximum allowed number of candidates is set to 6 as the default. If the maximum allowed number of candidates is modified to 1, candidate indices do not need to be signaled. Each bit of the candidate index is context-encoded using a separate context.
[0174] 2.1.17 Directional Planar Mode Two additional planar modes are used, where either horizontal interpolation only or vertical interpolation only is used to obtain the predicted samples.
[0175] For the planar horizontal pattern, horizontal linear interpolation is performed only based on the left and upper right reference points to predict the current point as: .
[0176] For the planar vertical mode, vertical linear interpolation is performed only based on the upper reference point and the lower left reference point to predict the current point as: .
[0177] Transform kernel selection for horizontal and vertical planar modes, as follows: Figure 20 As shown. If the intra-prediction mode of the current block is planar vertical mode, then the horizontal intra-prediction mode is used to derive the transform kernels in the MTS set and LFNST set. Furthermore, if the intra-prediction mode of the current block is planar horizontal mode, then the vertical intra-prediction mode is used to derive the transform kernels in the MTS set and LFNST set.
[0178] Figure 20 The transformation selection process for directional planar modes is shown.
[0179] 2.1.18 Direct block vectors for chroma blocks Direct block vectors are used for chroma blocks in a dual-tree stripe. When the chroma dual-tree is activated, a flag is transmitted via signaling to indicate whether the chroma blocks are encoded / decoded using IBC mode. Figure 21 If one of the luma blocks in the five locations shown is encoded or decoded in IBC or intraTMP mode, its block vector is scaled and used as the block vector for the chroma block. Template matching is used to perform the block vector scaling.
[0180] Figure 21 The luminance block used to derive the direct block vector is shown.
[0181] 2.1.19 Intra-frame prediction mode based on extrapolation filter (EFI mode) The proposed intra-frame prediction based on extrapolation filters is processed in two steps. First, the extrapolation filter coefficients are obtained from the neighboring reconstructed pixels of the current block with a predetermined template. Second, extrapolation generates prediction values position by position within the current block from top left to bottom right.
[0182] 2.1.19.1 Searching for the mean, minimum, and maximum values Similar to CCCM mode, the mean should be removed when the input is fed to the EIP filter. The DC mode value of the current block is used as the mean for the EIP prediction. The minimum and maximum values are searched from the reconstructed pixels in the reconstructed region, which has thirteen columns and thirteen rows.
[0183] 2.1.19.2 Calculation of Filter Coefficients Three types of reconstruction regions and three filter shapes are proposed, such as Figure 22 As shown. Figure 22 The diagram illustrates the three types of reconstructed regions defined, comprising thirteen columns or rows of reconstructed pixels. When the current block is predicted using the proposed EIP pattern, the decoder decodes the relevant syntax elements to determine the selected reconstructed region type and filter shape for the current block.
[0184] Figure 23The diagram illustrates three types of filter shapes with fifteen inputs and one output.
[0185] The selected filter slides across the selected reconstruction region in a one-pixel step to collect input and output samples for the EIP. The autocorrelation matrix and cross-correlation vector are constructed while removing the mean from the input and output samples. The EIP coefficients are then obtained using the same method as in CCCM.
[0186] 2.1.19.3 Prediction of the current block EIP mode predicts the current block position by position, such as Figures 24A to 24C As shown.
[0187] For a position located at the top left of the current block, the input to the EIP filter is the reconstructed sample.
[0188] For a location located along the boundary of the current block, part of the input to the EIP filter is the reference sample, and part of the input to the EIP filter is the previously predicted sample.
[0189] For other locations in the current block, the input to the EIP filter is the previously predicted samples.
[0190] Figures 24A to 24C Examples of predictions for different positions within the current block are shown. Figure 24A In this process, all inputs to EIP are reconstructed samples. Figure 24B In this process, part of the input consists of reconstructed sample points, and part consists of predicted sample points. Figure 24C In this context, all inputs to EIP are prediction samples.
[0191] To reduce prediction error, the searched minimum and maximum values are applied to limit the output range of each predicted value.
[0192] It is the predicted value at (x, y) in the current block. It searches for the minimum and maximum values from the thirteen reconstructed columns and rows. It is the first EIP filter derived. i One coefficient, These are the reconstructed or predicted values used for predictions at the current location. It is a value calculated using the DC prediction model.
[0193] 2.2 Cross-component residual model (CCRM) for inter-frame prediction (also known as inter-frame CCCM) It is proposed that when blocks use inter-frame prediction or intra-block copy (IBC), a cross-component residual model (CCRM) is applied to predict chrominance samples from reconstructed luminance samples. Figure 25 The decoder side of the method is shown. A cross-component filter is derived using prediction blocks for both luma and chroma. The derived filter is applied to the reconstructed luma block and mixed with the chroma prediction block to produce the final chroma prediction block. During the mixing process, the filtered reconstructed luma block uses a mixing weight of 0.75, and the chroma prediction block uses a mixing weight of 0.25.
[0194] Figure 25 The proposed method on the decoder is shown.
[0195] 2.2.1 Calculation of Convolution Filter and Filter Coefficients The proposed 8-tap filter consists of 6 spatial brightness samples, a nonlinear term, and a bias term. For example... Figure 26 As shown, the spatial luminance samples (L0, ..., L5) are obtained from the luminance grid. The six luminance samples closest to the chromaticity position C are selected without downsampling. The predicted chromaticity value is obtained as follows: predChromaVal = c0L0+ c1L1 + c2L2 + c3L3 + c4L4 + c5L5 + c6nonlinear((L0+L3+1)>>1) + c7B, Where nonlinearity is the nonlinear operator of CCCM, and B is the bias.
[0196] Figure 26 The luminance samples L0, ..., L5 are shown relative to the chromaticity sample C.
[0197] The filter coefficients were derived using the division-free Gaussian elimination method of ECM, and the necessary offset was applied to the samples before the filter derivation.
[0198] For the derivation of the filter coefficients, a maximum of 256 chromaticity samples are used.
[0199] The offset of the ECM's division-free Gaussian elimination method (used to solve the CCRM filter coefficients) is obtained by averaging four points of the luminance and chrominance prediction blocks, where the four points correspond to the top left, top right, bottom left, and bottom right corners of the block.
[0200] 2.2.2 Bitstream Signaling The use of this mode utilizes the CABAC-encoded TU level flags transmitted via signaling. A new CABAC context is included to support this. The CCRM flags are transmitted via signaling only when the TU's luminance Cbf is non-zero and the CU's predMode is MODE_INTER or MODE_IBC.
[0201] 2.2.3 Encoder Operation When the luminance Cbf is non-zero and the predMode of the CU is MODE_INTER or MODE_IBC, the encoder performs RD decisions in the transform selection loop for the chrominance component.
[0202] 2.3 Block Vector Guided CCCM The Block Vector Guided CCCM (BVG-CCCM) method uses the block vectors of co-occurring luma blocks encoded and decoded in IBC or intraTMP mode to determine the reference region for calculating CCCM parameters. The reference region in the luma and the corresponding region in the chroma channel are then used to calculate the CCCM parameters. CCCM prediction is performed using the calculated model parameters and co-occurring luma samples. Figure 27 The reference region in the BVG-CCCM method is shown.
[0203] This mode is enabled only in intra-frame stripes. Additionally, an SPS level flag is introduced to enable or disable this mode.
[0204] The BVG-CCCM mode uses an 11-tap filter for cross-component prediction, as shown below: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P(C) + c6P(N) + c7P(S) +c8P(W) + c9P(E)+ c 10 B.
[0205] like Figure 28 As shown, the input of the spatial domain 5-tap component of the filter consists of the center (C) luminance sample that is co-located with the chrominance sample to be predicted, and its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors.
[0206] The nonlinear term P is represented as the square of the corresponding brightness sample, and B is the bias term.
[0207] Figure 27 The reference area of BVG-CCCM is shown.
[0208] Figure 28 The spatial portion of the convolution filter is shown.
[0209] Similar to the Direct Block Vector (DBV) mode in ECM-9.0, in the same lumen block region, such as Figure 29 The five locations shown were scanned, and the associated block vectors were subsequently used to determine the reference region for parameter calculation in the BVG-CCCM method.
[0210] This mode can use block vectors (multiple) from both IBC and intraTMP encoded blocks from the same brightness region.
[0211] Figure 29 The location used for block vector derivation from the co-position brightness block is shown.
[0212] 2.3.1 Bitstream Signaling The use of this mode utilizes PU-level flags encoded and decoded via CABAC, which are transmitted via signaling. If the co-occurrence block is encoded and decoded in IBC or intraTMP mode, and the cross-component index is LM_CHROMA_IDX or MMLM_CHROMA_IDX, then the BVG-CCCM flag is transmitted via signaling.
[0213] 2.3.2 Encoder Operation The encoder performs two additional RDs for BVG-CCCM variants of both single-model and multi-model CCCM.
[0214] 2.4 On the cross-component model for residual encoding and decoding in image and video encoding and decoding 2.4.1 Issues related to cross-component models used for residual encoding and decoding There are several problems with existing video encoding and decoding technologies, and further improvements will be made to achieve higher encoding and decoding gains.
[0215] 1. Several aspects of the video units encoded and decoded by CCRM (such as filter terms, model type, and applied block type) can be further improved.
[0216] 2. The prediction estimated by CCRM will compete with the original inter-frame prediction. The residual with the lower cost will ultimately be selected. However, the concept of fusion can be involved to achieve better results.
[0217] 3. CCRM is applied to inter-blocks as long as the luminance component has a non-zero CBF. The CCRM on / off decision can be further designed.
[0218] 4. Currently, the CCRM model is applied with luminance reconstruction as input and estimated chromaticity prediction as output. However, the estimated chromaticity prediction can be generated by adding the residual block of the CCRM estimate to the chromaticity prediction without CCRM.
[0219] 2.4.2 Related Solutions The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0220] The term "video unit" or "code-decoder unit" can refer to a picture, strip, slice, code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0221] The term "block" can refer to code-decode tree block (CTB), code-decode tree unit (CTU), code-decode block (CB), CU, PU, TU, PB, TB.
[0222] The terms "motion vector" or "block vector" can refer to the vector of horizontal and vertical displacement between the position of a reference block and the position of the current block. The reference block can be a video unit in a reference image within the RPL list. Alternatively, the reference block can be a video unit in the current image.
[0223] The term "LM" can refer to any linear regression-based method, such as CCLM, MMLM, CCCM, GL-CCCM, CCCM without downsampling, GLM, GLM with luminance values, etc. It can also be referred to as "Cross-Component Prediction (CCP)". CCP models can be used for intra-frame prediction, IBC prediction, or inter-frame prediction.
[0224] The term "CCLM" can refer to a single-model LM mode, which can be a single-model CCLM, a single-model CCCM, a single-model GL-CCCM, a single-model CCCM without subsampling, a single-model GLM, a single-model GLM with luminance values, a multi-model CCLM, a multi-model MMLM, a multi-model CCCM, a multi-model GL-CCCM, a multi-model CCCM without subsampling, a multi-model GLM, a multi-model GLM with luminance values, etc.
[0225] The term "MMLM" can refer to a multi-model LM mode, which can be multi-model CCLM, MMLM, multi-model CCCM, multi-model GL-CCCM, multi-model CCCM without downsampling, multi-model GLM, multi-model GLM with luminance values, etc.
[0226] The term "MFLM" can refer to a multi-filter LM mode, which can be MF-CCLM, MF-CCCM, MF-GLM, MF-CCRM, MF-CCCM for inter-frame use, multi-filter IBC filter, multi-filter intraTMP filter, and / or variations of the mentioned mode.
[0227] The term "CCCM" can refer to regular CCCM mode, GL-CCCM mode, or CCCM without downsampling, CCRM, etc.
[0228] The term "GL-CCCM" can refer to a CCCM mode that takes into account the gradient and location of the samples involved.
[0229] The term "CCCM without downsampling" can refer to a CCCM mode that takes into account unsampled luminance samples.
[0230] The term "CCRM" can refer to residual encoding / decoding or derivation based on cross-component models. It can also refer to inter-frame / IBC prediction based on CCCM models (such as inter-frame / IBC CCCM). Furthermore, it can refer to intra-frame prediction based on CCCM models (such as intra-frame CCCM). It can refer to the generation and application of cross-component models (such as luma-to-chroma prediction). It can also refer to the generation and application of models within the same component (such as luma-to-luma prediction).
[0231] In this document, Cross Component Prediction (CCP) can refer to any cross component prediction method, such as any kind of CCLM / CCCM / GLM / GL-CCCM.
[0232] It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0233] 1) The residuals (and / or predictions) of chroma blocks can be derived based on cross-component models.
[0234] a. For example, the cross-component model can be a specific extrapolation filter (e.g., EIP, etc.).
[0235] b. For example, the cross-component model can be a specific interpolation filter (e.g., GLM, etc.).
[0236] c. For example, the cross-component model can be a specific convolutional filter (CCCM, GL-CCCM, CCCM without downsampling, CCRM, inter-frame CCCM, intra-frame CCCM, etc.).
[0237] d. For example, the cross-component model can be a specific linear filter (e.g., CCLM, MMLM, etc.).
[0238] 2) Cross-component models used for residual coding and decoding (e.g., CCRM) may not contain nonlinear terms.
[0239] a. For example, a cross-component model used for residual encoding and decoding may contain linear terms and / or bias terms, but not nonlinear terms.
[0240] 3) CCRM can be used for intra-frame blocks or IBC blocks.
[0241] a. For example, it can be used for intra-frame blocks or IBC blocks in intra-frame (such as I) stripes.
[0242] b. For example, it can be used for intra-frame blocks or IBC blocks in inter-frame (such as B or P) stripes.
[0243] c. For example, it can also be used for single trees.
[0244] d. For example, it can also be used for two trees.
[0245] e. For example, in a single-tree I-strip, both luminance and chrominance are encoded and decoded by IBC (or intraTMP). CCRM can be generated based on reconstructed luminance and chrominance samples within a reference block retrieved / guided by block vectors, and a residual model is applied to estimate the reconstructed values of chrominance samples in the current block.
[0246] f. For example, in a dual-tree system, luminance is encoded and decoded by IBC (or intraTMP), while chrominance is encoded and decoded by intraframe. CCRM can be generated based on reconstructed luminance samples within a reference luminance block retrieved / guided by block vectors, as well as reconstructed chrominance samples that are co-located (e.g., at the same position) within that luminance block, and a residual model is applied to estimate the reconstructed values of the chrominance samples in the current block.
[0247] 4) CCRM can be used for chroma blocks encoded and decoded by DBV.
[0248] a. For example, a reference chroma block and its corresponding luma block can be identified based on the block vector of the chroma block encoded and decoded by DBV. These samples can be used as training samples for computation of the CCRM model.
[0249] b. For example, the derived CCRM model is applied to the reconstructed luminance signal of the DBV chromaticity block to produce the final chromaticity prediction.
[0250] 5) The CCRM model can be generated based on the correlation between the luminance and chrominance reconstruction values from neighboring / non-adjacent samples of the current block.
[0251] a. For example, alternatively, the CCRM model can be generated based on the correlation between the luminance and chrominance reconstruction values in a reference block of a reference image.
[0252] b. For example, alternatively, the CCRM model can be generated based on the correlation between the luminance and chrominance reconstruction values in a reference block in the current image.
[0253] 6) For example, the CCCM used for intra-frame prediction and the CCCM used for inter-frame prediction (e.g., CCRM) can share the same logic.
[0254] a. For example, both can follow the same logic to obtain training samples.
[0255] b. For example, both can follow the same logic to determine the training area.
[0256] 7) CCRM models can be generated based on unsampled luminance samples.
[0257] a. For example, CCRM model coefficients can be solved based on unsampled luminance samples from a reference region used as training samples.
[0258] b. For example, the CCRM model can be applied to chroma blocks, where the chroma prediction of the current chroma block is generated based on the unsampled luminance samples of the co-located luminance block.
[0259] 8) More than one CCRM model can be generated for a block (e.g., MM-CCRM model, MF-CCRM model, CCRMMerge model).
[0260] a. For example, the training samples of CCRM can be divided into more than one class (e.g., two classes), and each group of samples can contribute to a unique model. In this way, multiple models can be generated, each with its own filter coefficients. Each derived filter is applied to the luminance reconstruction signal of its corresponding group to produce the final predicted value for the current chroma sample belonging to the corresponding class.
[0261] i. For example, according to the multi-model CCRM (e.g., MM-CCRM) mode, training sample pairs of luminance and chrominance sample pairs of the reference block (e.g., these training samples in the reference frame) can be divided into more than one class.
[0262] ii. For example, alternatively, training sample pairs (e.g., these training samples in the reference frame) with respect to the luminance and chrominance sample pairs of neighboring samples that are adjacent / non-adjacent to the reference block can be classified into more than one category, according to a multi-model CCRM (e.g., MM-CCRM) mode.
[0263] iii. For example, alternatively, training sample pairs (e.g., those training samples in the current frame) of luminance and chrominance sample pairs of neighboring / non-adjacent samples of the current video cell can be classified into more than one class, according to a multi-model CCRM (e.g., MM-CCRM) mode.
[0264] iv. For example, in addition, by following the same criteria (e.g., by a threshold), the luminance samples in the current video unit are divided into more than one group, and for each luminance sample belonging to a category, a corresponding model can be applied to generate the model-estimated chrominance samples belonging to that group.
[0265] b. For example, multiple sets of training samples can be used to derive multiple models.
[0266] i. In one example, there are two groups where the distance between the training sample and the current sample is different.
[0267] c. For example, the threshold for separating samples into different categories (e.g., the classification threshold) may depend on the values of samples within or near the training area.
[0268] i. For example, the training region can be a reference block of the current video unit (e.g., these training samples are in the reference frame).
[0269] 1. For example, a reference block can be derived based on a block vector.
[0270] 2. For example, a reference block can be derived based on motion vectors.
[0271] ii. For example, the threshold can be derived based on samples that are adjacent to or not adjacent to the reference block of the current video unit (e.g., these training samples are in the reference frame).
[0272] iii. For example, the threshold can be derived based on samples that are adjacent to or not adjacent to the current video unit (e.g., these training samples are in the current frame).
[0273] iv. For example, the threshold can be derived based on the averaging / intermediate / intermediate operation of more than one sample point within or near the training region.
[0274] v. For example, classification thresholds can be derived based on unsampled brightness sample values.
[0275] 1. Alternatively, the classification threshold can be derived based on the downsampled brightness sample values.
[0276] a. For example, a K-tap (such as K=6) downsampling filter can be used to reduce K surrounding luminance samples to a single downsampled luminance sample value.
[0277] vi. For example, classification thresholds can be derived based on offset removal schemes.
[0278] 1. For example, the offset can be derived based on luminance samples located at fixed positions (such as the upper left or center) within the reference video unit.
[0279] 2. For example, the offset values calculated for the classification threshold derivation and the CCRM model can be the same.
[0280] vii. For example, classification thresholds can be derived at the sub-block level.
[0281] viii. For example, classification thresholds can be derived based on CU / PU / TU levels.
[0282] ix. For example, the classification threshold can be calculated based on (downsampled or non-downsampled) brightness prediction samples.
[0283] x. For example, the classification threshold can be derived based on the brightness residual sample values.
[0284] 1. For example, for a second video unit (e.g., a sub-block) that does not have a non-zero residual, the predicted samples of such a video unit may not be included in the calculation of the classification threshold for the first video unit.
[0285] a. For example, the second video unit may be a subset of the first video unit.
[0286] b. For example, the second video unit can be equal to the first video unit.
[0287] d. For example, MM-CCRM can be applied at the sub-block level.
[0288] i. For example, the size of the sub-block can be predefined.
[0289] 1. For example, the predefined sub-block size can be 16x16, or 32x32, etc.
[0290] 2. For example, predefined rules can be used to determine the sub-block size of the MM-CCRM for a specific video unit.
[0291] a. For example, the sub-block size can be adapted to the block dimensions (width and / or height) of the current video block.
[0292] b. For example, for a sub-block of a video unit encoded and decoded by MM-CCRM, a minimum number of chroma samples can be guaranteed.
[0293] ii. For example, if a video unit is larger than a predefined sub-block size, the video unit can be divided into more than one sub-block and MM-CCRM can be performed.
[0294] iii. For example, at least one sub-block of a video unit may have more than one CCRM model.
[0295] iv. For example, each sub-block (and its associated training region) can have its own classification threshold.
[0296] 1. For example, the classification threshold for a specific sub-block can be calculated based on the training sample values belonging to that sub-block.
[0297] a. For example, brightness training samples in a reference block can be used to calculate a classification threshold.
[0298] v. For example, all sub-blocks (and their associated training regions) can share the same classification threshold.
[0299] 1. For example, a classification threshold can be calculated and used for all sub-blocks.
[0300] 2. For example, the classification threshold for all sub-blocks in the current video unit can be calculated based on the training sample values of the current video unit.
[0301] 3. For example, the classification threshold for all applicable sub-blocks in the current video unit can be calculated based on the training sample values of the current video unit.
[0302] a. For example, sub-blocks that do not contain non-zero residuals may not be counted.
[0303] vi. For example, each sub-block of a video unit can have its own training samples, and the training samples of a particular sub-block can be divided into more than one category.
[0304] 1. For example, training samples in the reference video unit of a reference image can be classified based on sub-blocks.
[0305] vii. For example, training samples from the current image can be classified into more than one group, but may not be divided into sub-blocks.
[0306] e. For example, MM-CCRM / MF-CCRM / CCRM Merge can be applied at the TU level (or PU / CU level).
[0307] i. For example, MM-CCRM / MF-CCRM / CCRM Merge can be applied on the basis of TU / CU / PU (for example, for the application of MM-CCRM, TU / CU / PU may not be divided into sub-blocks).
[0308] ii. For example, whether to use a multi-model CCRM / MF-CCRM / CCRM Merge based on TU / PU / CU can be determined at the TU / PU / CU level.
[0309] 1. For example, a video unit (e.g., TU / PU / CU) may choose to use a sub-block-based CCRM (e.g., CCRM / MM-CCRM / MF-CCRM / CCRM Merge) or a TU / PU / CU-based CCRM (e.g., CCRM / MM-CCRM / MF-CCRM / CCRM Merge).
[0310] a. For example, decisions can be made at the TU / PU / CU level.
[0311] f. For example, whether and / or how to apply MM-CCRM (and / or CCRM / MF-CCRM / CCRM Merge) can be deduced based on encoding and decoding information from both the encoder and decoder sides (e.g., not through signal transmission).
[0312] i. In one example, it can be derived on the fly, for example, using information from previously encoded / reconstructed samples.
[0313] ii. For example, the determination of whether to use sub-block-based CCRM / MM-CCRM / MF-CCRM / CCRM Merge or TU / CU / PU level CCRM can be implicitly deduced based on codec information (e.g., not through signal transmission).
[0314] iii. For example, the determination of whether to use CCRM based on M1xM2 subblocks or CCRM based on N1xN2 subblocks can be implicitly deduced based on encoding and decoding information (e.g., without signal transmission).
[0315] 1. For example, M1 = 16 or 8 or 32 or TU / CU / PU.
[0316] 2. For example, M2 = 16 or 8 or 32 or TU / CU / PU.
[0317] 3. For example, N1 = 16 or 8 or 32 or TU / CU / PU.
[0318] 4. For example, N2 = 16 or 8 or 32 or TU / CU / PU.
[0319] 5. For example, M1 != N1 and / or M2 != N2.
[0320] iv. For example, the determination of whether to use sub-block-based MM-CCRM / MF-CCRM / CCRM Merge or TU / CU / PU level MM-CCRM / MF-CCRM / CCRM Merge can be implicitly deduced based on codec information (e.g., not through signal transmission).
[0321] v. For example, the determination of whether to use the MM-CCRM / MF-CCRM / CCRM Merge based on M1xM2 sub-blocks or the MM-CCRM / MF-CCRM / CCRM Merge based on N1xN2 sub-blocks can be implicitly deduced based on encoding / decoding information (e.g., without signal transmission).
[0322] 1. For example, M1 = 16 or 8 or 32 or TU / CU / PU.
[0323] 2. For example, M2 = 16 or 8 or 32 or TU / CU / PU.
[0324] 3. For example, N1 = 16 or 8 or 32 or TU / CU / PU.
[0325] 4. For example, N2 = 16 or 8 or 32 or TU / CU / PU.
[0326] 5. For example, M1 != N1 and / or M2 != N2.
[0327] vi. For example, the determination of whether to use SM-CCRM / MF-CCRM / CCRM Merge or MM-CCRM / MF-CCRM / CCRMMerge can be implicitly deduced based on codec information (e.g., not through signal transmission).
[0328] vii. For example, determining a cost-based method that can be derived from the decoder.
[0329] 1. For example, the cost of decoder derivation can be calculated based on minimizing the SAD / SATD / SSE / MSE between the model estimated sample values and the true reconstructed sample values, where a sample can refer to at least one training sample among the training samples.
[0330] 2. For example, a method with lower cost can be chosen as the final method to be applied to the current video unit.
[0331] viii. For example, information can be determined based on a reference image.
[0332] 1. For example, the POC distance between the current image and its reference image can be determined.
[0333] 2. For example, determination can be based on a reference index.
[0334] g. Alternatively, whether and / or how MM-CCRM (and / or CCRM / MF-CCRM / CCRM Merge) can be applied can be transmitted via signaling in the bitstream.
[0335] i. For example, a syntax element (e.g., a flag, an index, etc.) can be signaled based on whether the current block is CCRM-coded / decoded.
[0336] 1. For example, if a video unit is CCRM-coded / decoded, a syntax element (e.g., a flag, an index, etc.) can be further signaled to indicate whether it is MM-CCRM / MF-CCRM / CCRM Merge.
[0337] ii. For example, a syntax element (e.g., a flag, an index, etc.) can be signaled to indicate whether it is sub-block-based MM-CCRM or TU / CP / PU-based MM-CCRM / MF-CCRM / CCRM Merge.
[0338] iii. For example, a syntax element (e.g., a flag, an index, etc.) can be signaled to indicate whether it is sub-block-based CCRM or TU / CP / PU-based CCRM / MF-CCRM / CCRM Merge.
[0339] iv. For example, a syntax element can be signaled based on the block dimensions (width and / or height).
[0340] 1. For example, if W H < T (such as T = 16 or 32), it may not be signaled.
[0341] v. For example, a syntax element can be signaled based on the residual / coefficients of the current luma block.
[0342] 1. For example, whether to signal a syntax element can be based on whether there is a residual (or non-zero coefficients) in the current luma block.
[0343] 2. For example, whether to signal a syntax element can be based on the distribution / number / value of the residual (or non-zero coefficients) in the current luma block.
[0344] vi. For example, a syntax element can be signaled based on the prediction method of neighboring blocks.
[0345] 1. For example, it can be based on whether neighboring blocks (such as left and / or above neighbors) use the CCRM / MM-CCRM / MF-CCRM / CCRM Merge mode.
[0346] vii. For example, the context model of a syntax element can depend on the coding / decoding information of neighboring blocks or the current block.
[0347] 1. For example, the context model can be derived based on whether neighboring blocks (such as left and / or upper neighbors) use the CCRM / MM-CCRM / MF-CCRM / CCRM Merge mode.
[0348] 2. For example, the context model can be derived based on whether the block dimensions of the current block satisfy specific conditions.
[0349] a. For example, if the current block (e.g., TU / PU / CU) is long or wide (e.g., W > a H, and / or H > b W, where W and H are the width and height of the current block, and a and b are predefined constants, e.g., a = b = 2), then the specified context model can be used.
[0350] h. For example, block restrictions can be applied to indicate the allowance of the MM-CCRM mode.
[0351] i. In one example, assuming that the width and height of the chrominance CU / PU / TU are represented as W and H, then MM-CCRM can be allowed when at least one of the following conditions is satisfied: 1. W H > T0 or W H >= T0 (e.g., T0 = 16 or 32 or 64 or 128).
[0352] 2. W > T1, or, W >= T1.
[0353] 3. H > T2, or, H >= T2.
[0354] 4. Min (W,H) > T3, or, Min (W,H) >= T3.
[0355] 5. Max (W,H) < T4, or, Max (W,H) <= T4.
[0356] 6. W < T5 H, or, W <= T5 H.
[0357] 7. W > T6 H, or, W >= T6 H.
[0358] 8. H < T7 W, or, H <= T7 W.
[0359] 9. H > T8 W, or, H >= T8 W.
[0360] 10. W H < T9, or W H <= T9.
[0361] ii. In one example, for blocks where a specific tool is enabled (e.g., affine motion compensation is enabled), MM-CCRM can be prohibited.
[0362] i. For example, a video unit encoded / decoded by CCRM can always use multi-model CCRM.
[0363] i. Alternatively, a video unit encoded / decoded by CCRM can use single-model CCRM or multi-model CCRM.
[0364] 9) Chrominance Cb and Cr can share a CCRM.
[0365] a. Alternatively, chrominance Cb and Cr can build their own CCRM.
[0366] 10) For the filter design of the CCRM model, sample values and / or gradients and / or position information can be considered.
[0367] a. For example, at least one K-tap filter can be used for the CCRM model, which consists of (multiple) K1 sample terms, (multiple) K2 gradient terms, (multiple) K3 position / position terms, (multiple) K4 non-linear terms, (multiple) K5 bias terms, etc.
[0368] i. For example, K1 = 0 or 1 or 2 or 5 or 6.
[0369] ii. For example, K2 = 0 or 1 or 2 or 4.
[0370] iii. For example, K3 = 0 or 1 or 2 or 4.
[0371] iv. For example, K4 = 0 or 1 or 2 or 4.
[0372] v. For example, K5 = 0 or 1.
[0373] vi. For example, K = K1 + K2 + K3 + K4 + K5.
[0374] vii. For example, the sample terms can be calculated based on the luminance sample values.
[0375] viii. For example, the gradient terms can be calculated based on more than one sample adjacent to a specific luminance sample.
[0376] ix. For example, the position / location item can be calculated based on the horizontal and / or vertical coordinates of a specific brightness sample point, where the coordinates can be relative to the upper left position of a specific reference area.
[0377] x. For example, a nonlinear term can be the square of a specific value (e.g., an intermediate value related to bit depth, such as 512 or 256, or a specific brightness value).
[0378] xi. For example, a nonlinear term can be the square of the gradient value based on a specific gradient term.
[0379] xii. For example, the offset can be subtracted from the terms of the K-tap filter.
[0380] 1. For example, the offset can be derived based on predefined rules such as the value of the top-left training sample in the training region, or the average / median value of more than one sample in the training region.
[0381] xiii. For example, the coefficients of a K-tap filter can be solved using a Gaussian elimination solver.
[0382] xiv. For example, the coefficients of a K-tap filter can be solved using the LDL decomposition method.
[0383] i. For example, the coefficients of a K-tap filter can be solved using linear regression.
[0384] ii. For example, the coefficients of a K-tap filter can be solved using linear equations.
[0385] b. For example, more than one filter can be used, and the final prediction can be derived based on fusing the filtered outputs of multiple filters together.
[0386] i. For example, the weights that fuse multiple filter values can be solved using a Gaussian elimination solver.
[0387] ii. For example, the weights for fusing multiple filter values can be solved using the LDL decomposition method.
[0388] 11) For example, for a video unit encoded or decoded by CCRM, more than one filter may be allowed, and the final selection of which filter may be transmitted through the signal or divided.
[0389] a. For example, syntax elements can be transmitted via signals to indicate which filter (e.g., CCLM or CCCM) is used in CCRM mode.
[0390] b. For example, indicating which filter (e.g., CCLM or CCCM) is used in CCRM mode can be determined based on template costs from both the encoder and decoder.
[0391] c. For example, indicating which filter (e.g., CCLM or CCCM) is used in CCRM mode can be determined based on the cost derived from both the encoder and decoder.
[0392] 12) The filter output can be limited to a single value.
[0393] a. For example, it can be amplitude-limited based on the reconstructed values in the training region.
[0394] i. For example, the training region can be derived based on block vectors (or motion vectors).
[0395] ii. For example, the training region can be adjacent to the current block.
[0396] iii. For example, the training region can be the reference region of the current block.
[0397] iv. For example, the filter output can be capped within the minimum and maximum values of the reconstructed (or predicted) luminance sample values in the training region.
[0398] b. For example, it can be limited based on the reconstructed (or predicted) value in the co-located luminance block of the current chroma block.
[0399] i. For example, it can be limited within the minimum and maximum values of the current block brightness reconstruction (or prediction) value.
[0400] c. For example, if a value is outside the valid range, it can be ignored / discarded / not used.
[0401] 13) CCRM parameters can be stored in the cache and used for encoding and decoding of future blocks.
[0402] a. For example, CCRM parameters for video units (e.g., CU, PU, color components, Cb, Cr, etc.) may include model type, model coefficients, whether it is a single model or multiple models, threshold for separating samples into multiple models, etc.
[0403] b. For example, it can be stored in a local cache for encoding and decoding future blocks in the current image.
[0404] c. For example, it can be stored in the temporal domain / image / frame buffer for encoding and decoding future blocks in a future decoded image.
[0405] i. For example, the CCRM parameters of the current frame / image can be stored and referenced for the CCP process of future frames / images.
[0406] ii. For example, it can be stored in association with motion and pattern information of the video unit.
[0407] 14) Video blocks can inherit model parameters from previous filter-based codec blocks. In the sub-items below, CCRM can refer to any filter model that includes cross-component models or same-component models.
[0408] a. For example, a cross-component model can refer to a cross-component residual model or cross-component prediction model used within or between frames, where model generation and application are based on the relationship between different color components (such as luminance to chrominance, chrominance Cb to chrominance Cr).
[0409] b. For example, a common component model can refer to an inter-frame / IBC / IBC-LIC / inter-frame-LIC / intraTMP / EIF filter, where model generation and model application are based on the relationship in the common component (such as luminance to luminance S).
[0410] c. For example, video blocks can be encoded and decoded using a CCP inheritance mode.
[0411] d. For example, video blocks can be encoded and decoded using a CCP Merge (e.g., CCMerge) mode.
[0412] e. For example, video blocks can be encoded and decoded using a filter model inheritance / merge mode. For example, model parameters of previously encoded and decoded blocks using CCRM can be stored in a cache (e.g., local cache, image cache, temporal cache, history-based LUT, etc.).
[0413] f. In one example, parameters could refer to filter information, linear or nonlinear parameters of the model, model index, etc.
[0414] g. For example, at least one syntax element may be signaled at the video unit level (e.g., block level, tu / pu / cu level, etc.) to specify whether and / or how to use CCRM model inheritance mode (e.g., CCRM Merge mode, or CCP Merge mode, or IBC / intraTMP filter Merge mode).
[0415] i. For example, an indicator can be transmitted via signaling at the video unit level to specify whether the current video unit uses the regular CCRM mode or the CCRM model inheritance mode.
[0416] 1. For example, alternatively, the indicator is based on at least one type of CCRM (e.g., conventional CCRM mode) used for the current video unit being transmitted via signal.
[0417] 2. For example, a first syntax is transmitted via signal to indicate the type of CCRM mode used by the current video unit, and then a second syntax is further transmitted via signal to indicate which type of CCRM mode is being used.
[0418] a. In addition, alternatively, the second syntax may be transmitted via signaling only if there is at least one available CCRM candidate (e.g., at least one valid CCRMMerge candidate).
[0419] ii. For example, alternatively, if the CCRM model inheritance pattern is used, another syntax (e.g., indexing) can be further signaled to specify which CCRM model candidate is selected for inheritance.
[0420] 1. For example, candidate indexes can be encoded or decoded to indicate CCRM candidates from a candidate list.
[0421] 2. For example, candidate indices can be encoded or decoded using the Rounding Rice (TR) binarization process, or the Rounding Binary (TB) binarization process, or the k-th exponential Columbus (EGk) binarization process, or the fixed-length (FL) binarization process.
[0422] 3. For example, the maximum allowed CCRM candidates (e.g., the maximum length of the candidate list) can be specified in the codec (such as 12 or 8 or 4 or 2, etc.).
[0423] iii. For example, alternatively, an indicator may be transmitted via signaling at the video unit level to specify whether the current video unit uses the CCRM model inheritance mode, and / or which candidate is used for the CCRM model inheritance mode.
[0424] iv. For example, syntax elements can be transmitted via signaling based on the residuals / coefficients of the current luma block.
[0425] 1. For example, whether to transmit syntax elements via signaling can be based on whether there is a residual (or non-zero coefficient) in the current luma block.
[0426] 2. For example, whether a syntax element is transmitted via signaling can be based on the distribution / number / value of the residuals (or non-zero coefficients) in the current luma block.
[0427] v. For example, syntax elements can be transmitted via signaling based on a prediction method of neighboring blocks.
[0428] 1. For example, it can be based on whether the CCRM / CCRMMerge mode is used based on neighboring blocks (such as left and / or top neighbors).
[0429] vi. For example, the context model of a syntax element can depend on the encoding / decoding information of neighboring blocks or the current block.
[0430] 1. For example, the context model can be derived based on whether neighboring blocks (such as left and / or top neighbors) use the CCRM / MM-CCRM / MF-CCRM / CCRM Merge pattern.
[0431] 2. For example, a context model can be derived based on whether the block dimension of the current block satisfies specific conditions.
[0432] a. For example, if the current block (e.g., TU / PU / CU) is long or wide (e.g., W>a) H, and / or H>b W, where W and H are the width and height of the current block, and a and b are predefined constants (e.g., a=b=2), then the specified context model can be used.
[0433] h. For example, if the CCRM model inheritance pattern is used, a list of CCRM model candidates can be generated.
[0434] i. For example, the maximum length of the list size can be predefined in the bitstream (e.g., the size is equal to 6, 10, or 12 candidate models).
[0435] 1. In addition, the size of the history table can be predefined (e.g., size equal to 5 or 6).
[0436] ii. For example, CCRM model candidates can be obtained based on previously encoded CCRM blocks that are spatially adjacent, and / or temporally adjacent, and / or spatially non-adjacent, and / or historical CCRM candidates, and / or shifted candidates, and / or default CCRM candidates.
[0437] 1. For example, the candidate insertion order can follow predefined rules, such as spatially adjacent -> temporally adjacent -> spatially non-adjacent -> history -> shift -> default.
[0438] a. Alternatively, the candidate insertion order can follow predefined rules, such as spatially adjacent -> spatially non-adjacent (if applicable) -> history (if applicable) -> shift (if applicable) -> default (if applicable).
[0439] 2. For example, CCRM candidates can be inspected at a sub-block (e.g., 4x4) granularity.
[0440] a. For example, each consecutive sub-block within a predefined area can be inspected; for instance, all 4x4 sub-blocks above and to the left of the current video unit can be inspected.
[0441] b. Alternatively, a predefined distributed check order can be used.
[0442] 3. For example, the location of non-adjacent neighboring blocks can be based on the block dimension of the current video unit, such as a specific distance from the current video unit, where the distance is proportional to the width and / or height of the current video unit.
[0443] 4. For example, motion shifts (e.g., zero vectors or non-zero vectors) can be used to locate temporal candidates.
[0444] a. For example, motion shift can be based on the motion vector of neighboring blocks.
[0445] b. For example, temporal candidates can come from co-located images.
[0446] c. Alternative locations, temporal candidates can come from reference images that are not necessarily in the same position.
[0447] 5. For example, history-based CCRM candidates can come from a first-in-first-out history table.
[0448] a. For example, history tables can be initialized at the slice / ctu row / strip / image level.
[0449] 6. For example, temporal candidates can be derived based on motion shifts.
[0450] a. For example, the temporal candidate of a block encoded by IBC / IntraTMP / inter-frame coding can be derived based on the MV / BV of neighboring blocks.
[0451] iii. For example, it can be based on the MV / BV of the current block. For example, deduplication / redundancy / similarity checks can be applied to CCRM candidate list construction.
[0452] 1. For example, if the candidate to be inserted is different from the specified candidate that is already in the list, the candidate to be inserted is inserted into the list.
[0453] a. For example, specifying a candidate can refer to all available CCRM candidates in the list.
[0454] b. Alternatively, a candidate can refer to one or more specified CCRM candidates in the list (e.g., the last one, and / or the last one in the list – X, where X is a predefined constant).
[0455] 2. For example, the same deduplication / redundancy / similarity check rules can be applied to all types of CCRM candidates.
[0456] a. Alternatively, different deduplication / redundancy / similarity check rules can be applied to different types of CCRM candidates.
[0457] iv. For example, CCRM candidate reordering can be applied.
[0458] 1. For example, CCRM candidates in the list can be sorted based on the cost inferred by the decoder (e.g., template cost).
[0459] 2. For example, the cost can be derived based on applying CCRM candidates to the reference region / block of the current video unit.
[0460] a. For example, a reference region / block can be identified by the motion vector of the current block.
[0461] b. For example, for each CCRM candidate, the model is first applied to the reference luminance to obtain the predicted reference chromaticity, and then the cost is calculated as the absolute difference between the true reference chromaticity and the predicted reference chromaticity.
[0462] 3. For example, the cost can be derived based on applying CCRM candidates to the neighboring regions / blocks of the current video unit.
[0463] a. For example, a neighboring region / block can be the one above or to the left of the current block.
[0464] b. For example, for each CCRM candidate, the model is first applied to neighboring luminance to obtain the predicted neighboring chromaticity, and then the cost is calculated as the absolute difference between the true neighboring chromaticity and the predicted neighboring chromaticity.
[0465] 4. For example, based on the cost derived from the decoder above, CCRM candidates can be sorted from the lowest cost to the highest cost, and the one with the lowest cost is sorted first in the list.
[0466] i. For example, if the CCRM model inheritance mode is used, the specified CCRM candidate model is directly applied to the current video unit without model estimation.
[0467] i. For example, the first candidate in the CCRM list can always be used in the CCRM model inheritance pattern.
[0468] ii. Alternatively, which candidate from the CCRM list is used for blocks encoded and decoded via the CCRM model inheritance mode can be transmitted via signaling in the bitstream.
[0469] iii. For example, for an RRIBC block encoded and decoded using a CCRM model inheritance model, the inherited CCRM model can be applied based on the inherited RRIBC flip type.
[0470] 1. For example, if the inherited RRIBC flip type indicates that the inherited CCRM model comes from a block encoded and decoded by RRIBC (e.g., the inherited RRIBC flip type is non-zero), then the CCRM filter taps for the current block can be flipped according to the inherited RRIBC flip type.
[0471] a. For example, when generating CCRM filter taps from unsampled luminance samples, the filter taps of the unsampled luminance samples can be swapped / flipped.
[0472] 2. For example, suppose an 8-tap CCRM model consists of 6 spatial luminance samples, a nonlinear term, and a bias term. The spatial luminance samples (L0, ..., L5) are obtained from the luminance grid, selecting the 6 luminance samples closest to the chromaticity position C without downsampling. The predicted chromaticity value is obtained as, predChromaVal = c0 L0 + c1L1 + c2L2 + c3L3 + c4L4 + c5L5 + c6 nonlinear((L0 + L3 + 1) >> 1) + c7 B, where the nonlinearity is the nonlinear operator of CCRM, and B is the bias.
[0473] a. For example, if the inherited CCRM model comes from a block encoded with RRIBC using a horizontal flip, then when the inherited CCRM model is applied to the current block (e.g., regardless of whether the current block is encoded with RRIBC), luma samples L1 and L2 can be swapped. And luma samples L4 and L5 can be swapped.
[0474] b. For example, if the inherited CCRM model comes from a block encoded with RRIBC using a vertical flip, then when the inherited CCRM model is applied to the current block (e.g., regardless of whether the current block is encoded with RRIBC), luma samples L1 and L4 can be swapped. Furthermore, luma samples L0 and L3 can be swapped. Additionally, luma samples L2 and L5 can be swapped.
[0475] Figure 26 The luminance samples L0, ..., L5 are shown relative to the chromaticity sample C.
[0476] j. For example, the CCRM model of a block encoded and decoded by CCRM can be stored in a cache.
[0477] i. For example, the stored CCRM model information may include the following information.
[0478] 1. CCRM model coefficients / tap / parameters for Cb and Cr components.
[0479] 2. The intermediate value of the block encoded and decoded by CCRM.
[0480] 3. Bit depth of the block encoded and decoded by CCRM.
[0481] 4. The offset value of the Y component of the block encoded and decoded by CCRM.
[0482] 5. Offset values of the U and / or V components of the block encoded and decoded by CCRM.
[0483] 6. RRIBC flip type of blocks encoded and decoded by CCRM.
[0484] ii. Alternatively, model offsets may not be stored in the cache.
[0485] 1. For example, the model offset of a neighboring block may not be reused / inherited for the current block.
[0486] 2. For example, the model offset of the current block is recalculated based on specific available samples.
[0487] k. Alternatively, filter model information (e.g., model coefficients / tap / parameters / offsets of the Y component) can be stored in a cache, where the filter coefficients are calculated based on the relationship between samples in the same component (e.g., all training samples are in the luminance component domain).
[0488] i. Alternatively, model offsets may not be stored in the cache.
[0489] 1. For example, the model offset of a neighboring block may not be reused / inherited for the current block.
[0490] 2. For example, the model offset of the current block is recalculated based on specific available samples.
[0491] 15) The final prediction can be generated from a weighted sum of multiple fusion / hybrid hypotheses, where at least one hypothesis is based on predictions of CCPs (e.g., CCRM, CCRM Merge, CCP Merge, CCLM, LM, CCCM, GLM, etc.).
[0492] a. In one example, the final prediction of a block can be generated based on multiple prediction candidates from different CCPs (e.g., CCRM, CCRMMerge, CCP Merge, CCLM, LM, CCCM, GLM, etc.).
[0493] i. For example, more than one CCP prediction can be merged together.
[0494] ii. For example, the weights / coefficients of different fusion terms can be solved based on the Gaussian elimination method.
[0495] iii. For example, the weights / coefficients of different fusion terms can be solved based on the LDL decomposition method.
[0496] iv. For example, bias terms can be involved for fusion.
[0497] v. For example, nonlinear terms can be involved in fusion.
[0498] b. In one example, multiple CCP models can be derived to obtain a fused prediction.
[0499] i. Fusion predictions can refer to predictions generated through weighted summation.
[0500] ii. In one example, P0 in luminance and chrominance can be used to derive CCP model M0, P1 in luminance and chrominance can be used to derive CCP model M1, and the final chrominance prediction can be derived as wc0×Pc0+wc1×Pc1, where Pc0 and Pc1 are chrominance predictions obtained using M0 and M1, and wc0 and wc1 are weighting factors.
[0501] iii. P0 and P1 can be predictions from two different directions in a bidirectional prediction.
[0502] iv. For example, two hypothetical predictions can be generated based on the top two candidates in the CCP Merge pattern list, and the final prediction can be derived based on the weighted sum of the two hypothetical predictions.
[0503] c. In one example, the chromaticity prediction obtained through CCP can be fused with other predictions.
[0504] i. For example, chromaticity predictions obtained through CCP can be fused with intra-frame angular predictions.
[0505] ii. For example, chromaticity predictions obtained through CCP can be fused with CCLM predictions.
[0506] iii. For example, chromaticity predictions obtained through CCP can be fused with CCCM predictions.
[0507] iv. For example, chroma predictions obtained through CCP can be fused with original predictions (e.g., intra-frame, inter-frame, or IBC predictions) without CCP.
[0508] 1. For example, suppose the final chromaticity prediction can be derived as (w0×P0 + w1×P1 + offset) >> shift, where P0 represents the chromaticity prediction obtained through CCRM, and P1 represents the chromaticity prediction without CCP. a. The fusion weights w0 and w1 can be fixed and / or predefined, for example, w0=3 and w1=1, or w0=2 and w1=2.
[0509] b. The shift value can be a constant value that can be estimated based on w0 and w1, for example, shift = log2(w0 + w1).
[0510] c. The offset can be a constant value that can be estimated based on the shift and / or fusion weights, for example, offset = shift >> 1, or offset = log2(w0+w1) >> 1.
[0511] d. Alternatively, the fusion weights w0 and w1 can be adaptively determined based on encoding and decoding information (e.g., block dimension, neighboring prediction patterns, etc.).
[0512] d. In one example, the weighting factors for different hypotheses in the fusion / mixing process can be derived based on predefined rules. The final prediction P is assumed to be derived as P = w0 × P0 + w1 × P1 + w2 × P2 + ..., where w0, w1, and w2 are weighting factors.
[0513] i. For example, fixed values can be assigned to w0, w1, w2, ...
[0514] ii. For example, block-based w0, w1, w2... can be assigned.
[0515] iii. For example, a sample-based weighting factor can be assigned for each prediction (e.g., at least two different weights can be applied to different samples in a hypothetical prediction block).
[0516] iv. For example, w0, w1, w2 can be derived on the fly (e.g., based on the decoded neighboring samples, the neighboring prediction patterns, and / or template costs).
[0517] e. In one example, indicators of the weighting factors for different assumptions in the fusion / mixing process can be transmitted via signals in the bitstream.
[0518] i. For example, a lookup table containing multiple sets of weighting factors can be defined, and the index can be signaled to look up the corresponding weight.
[0519] f. For example, the sample values of the final fused / mixed prediction can be clipped to a predefined range; for instance, it can be required to be no less than T1 and no greater than T2, where T2 can depend on the bit depth. For example, Clip1(x) = Clip3(0, (1<...). <BitDepth ) 1, x).
[0520] g. For example, the proposed method can be applied to fuse more than one hypothesis, where at least one hypothesis is based on the prediction of the filter model.
[0521] i. For example, filter coefficients can be calculated based on the relationship between two sets of samples in the same component domain (e.g., one set consists of luminance samples adjacent to the current block, and the other set consists of luminance samples adjacent to the reference block).
[0522] ii. For example, a filter-based prediction can refer to a prediction generated based on applying a filter to a MC-compensated video cell.
[0523] 1. For example, filters can be based on IBC filters, intraTMP filters, LICs for inter-frame, LICs for IBC, EIF, etc.
[0524] iii. For example, filter-based predictions can be generated based on filter merge / inheritance patterns (e.g., IBC filter-based merge / inheritance patterns, intraTMP filter-based merge / inheritance patterns, inter-frame LIC-based merge / inheritance patterns, IBC LIC-based merge / inheritance patterns, and EIF-based merge / inheritance patterns).
[0525] 1. For example, a filter-based Merge pattern can refer to a pattern in which the filter model is inherited / derived from the candidate model (e.g., derived from a list of candidate models).
[0526] iv. For example, a prediction based on a filter model (e.g., encoded or decoded by prediction type A) can be fused with another prediction (e.g., encoded or decoded by prediction type B).
[0527] 1. For example, prediction type B may not be a prediction based on a filter model.
[0528] 2. For example, prediction type B can be a prediction based on a filter model, but the filter types in A and B are different.
[0529] 3. Alternatively, prediction type B can be a prediction based on a filter model, and the filter types in A and B are the same.
[0530] a. For example, two hypothetical predictions can be generated based on the top two candidates in a filter-based Merge pattern list, and the final prediction can be derived based on a weighted sum of the two hypothetical predictions.
[0531] 16) Whether CCRM predictions are fused with another prediction can be transmitted via signaling in the bitstream.
[0532] a. For example, a flag can be transmitted via signal at the video unit level (e.g., TU / PU / CU / strip header / picture header / SPS / PPS level) to indicate such a CCRM fusion mode.
[0533] b. Alternatively, CCRM predictions can always be merged with another prediction without the need for signal transmission.
[0534] 17) For CCRM, bidirectional forecasts can be managed in a different way than unidirectional forecasts. In the following discussion, assume that the forecasts from the two directions are P0 and P1, and that bidirectional forecasts are expressed as Pb = w0 × P0 + w1 × P1, where w0 and w1 are weighting factors.
[0535] a. In one example, Pb in luminance and chrominance can be used to derive the CCRM model.
[0536] b. In one example, P0 or P1 in luminance and chrominance can be used to derive the CCRM model.
[0537] c. In one example, which prediction was used to deduce that the CCRM model could be transmitted via signaling?
[0538] 18) The permission for the CCRM model may depend on at least one of the following: a. Prediction mode of video unit (e.g., MODE_INTRA, MODE_INTER, MODE_IBC, MODE_PLT, etc.).
[0539] b. Transformation type of the video unit (e.g., ACT, color transformation, transformation skip, etc.).
[0540] c. SBT (e.g., whether SBT is applied to the current video unit).
[0541] d. The number of non-zero coefficients in a video unit.
[0542] e. Luminance coefficients (e.g., luminance coefficient values, the absolute sum of all luminance coefficients, the last scan position of non-zero luminance coefficients, AC values, DC values, etc.) and segmentation tree type (e.g., single tree, dual tree).
[0543] f. Strip type (e.g., I, B, P stripes).
[0544] g. Color format (e.g., whether it is 4:0:0).
[0545] h. Availability of chromaticity components.
[0546] i. For example, CCRM may not be allowed for ACT and / or 4:0:0 color formats.
[0547] j. For example, the on / off state of the CCRM can be determined based on the last scan position of the non-zero luminance coefficient.
[0548] i. For example, if the last scan position is less than a threshold, CCRM can be presumed to be disabled for the current chroma unit, and therefore no syntax element for CCRM is transmitted via signaling.
[0549] 1. For example, the threshold can be a fixed constant (such as 1).
[0550] 2. For example, the threshold can be a variable based on encoding / decoding information such as block dimensions.
[0551] ii. For example, if the brightness is not transformed and skipped from encoding / decoding, such a condition can be checked.
[0552] iii. For example, such conditions can be checked regardless of whether the brightness is transformed and skipped from encoding / decoding.
[0553] k. For example, the on / off state of CCRM can be determined based on the absolute sum of non-zero luminance coefficients.
[0554] i. For example, it can be determined based on the absolute sum of all luminance coefficients (e.g., both AC and DC).
[0555] ii. For example, it can be determined based on the absolute sum of all luminance AC coefficients.
[0556] iii. For example, it can be determined based on the luminance DC coefficient value.
[0557] iv. For example, it can be determined based on at least one luminance coefficient value (e.g., DC and / or AC).
[0558] v. For example, if the absolute sum is less than the threshold, CCRM can be presumed to be disabled for the current chroma unit, and therefore no syntax element for CCRM is used by signal transmission.
[0559] 1. For example, the threshold can be a fixed constant value.
[0560] 2. For example, the threshold can be a variable based on encoding / decoding information such as block dimensions.
[0561] vi. For example, if the brightness is not transformed and skipped from encoding / decoding, such a condition can be checked.
[0562] vii. For example, such conditions can be checked regardless of whether the brightness is transformed and skipped from encoding / decoding.
[0563] viii. For example, such conditions can be checked together with conditions based on block size (e.g., TU width and / or width).
[0564] l. For example, if a transform skip is used on the luminance component, CCRM may not be applied to the chrominance component.
[0565] i. For example, if a transform skip is used on the luminance component, CCRM can be presumed to be disabled for the current chromaticity unit (e.g., Cb and / or Cr).
[0566] 1. Furthermore, in this case, no syntax elements for CCRM usage on that video unit are transmitted via signaling.
[0567] ii. Alternatively, if transform skip is used on the luminance component, CCRM for the current chromaticity unit (e.g., Cb and / or Cr) can be presumed to always be enabled.
[0568] 1. Furthermore, in this case, no syntax elements for CCRM usage on that video unit are transmitted via signaling.
[0569] 19) The application of CCRM can depend on template information.
[0570] a. For example, whether CCRM can be used for video units may depend on the template cost.
[0571] i. For example, if it is determined by a template cost-based method that CCRM is disabled for the current video unit (i.e., CCRM on / off is presumed rather than signaled), then no syntax element uses signaled CCRM for that video unit.
[0572] b. For example, such as Figure 30 As shown, assuming the current block is inter-frame encoded / decoded, two costs (e.g., SAD) can be calculated: the first cost is calculated based on the absolute difference between the current template predicted by the CCRM model and the actual reconstruction of the current template, and the second cost is calculated based on the difference between the reference template and the actual reconstruction of the current template. If the first cost is lower than the second cost, CCRM is presumed to be used for the current chroma unit; otherwise, the current chroma unit is encoded / decoded without CCRM.
[0573] i. For example, the CCRM model can be calculated based on the relationship between a reference luminance block (yellow) and a reference chrominance block (yellow).
[0574] ii. For example, if it is determined that CCRM is to be used, the CCRM model can be applied to the current luminance reconstruction block (i.e., the input of the CCRM model) and generate the current chromaticity prediction predicted by the CCRM model (i.e., the output of the CCRM model).
[0575] iii. For example, in this case (i.e., the on / off state of the CCRM is presumed rather than transmitted via signal), no syntax element is used for the CCRM on that video unit to be transmitted via signal.
[0576] iv. For example, such a template cost method can be applied to blocks that have undergone inter-frame encoding and decoding.
[0577] c. For example, such as Figure 31 As shown, assuming the current block is encoded / decoded using IBC, two costs (e.g., SAD) can be calculated: the first cost is calculated based on the absolute difference between the current template predicted by the CCRM model and the actual reconstruction of the current template, and the second cost is calculated based on the difference between the reference template and the actual reconstruction of the current template. If the first cost is lower than the second cost, CCRM is presumed to be used for the current chroma unit; otherwise, the current chroma unit is encoded / decoded without CCRM.
[0578] i. For example, the CCRM model can be calculated based on the relationship between a reference luminance block (yellow) and a reference chrominance block (yellow).
[0579] ii. For example, if it is determined that CCRM is to be used, the CCRM model can be applied to the current luminance reconstruction block (i.e., the input of the CCRM model) and generate the current chromaticity prediction predicted by the CCRM model (i.e., the output of the CCRM model).
[0580] iii. For example, in this case (i.e., the on / off state of the CCRM is presumed rather than transmitted via signal), no syntax element is used for the CCRM on that video unit to be transmitted via signal.
[0581] iv. For example, such a template cost method can be applied to blocks encoded and decoded by IBC.
[0582] d. For example, whether to use a template cost-based approach to determine the on / off state of CCRM may depend on whether the current video unit (e.g., TU) has residual / non-zero coefficients and / or SBT usage.
[0583] i. For example, for the zero-residual portion of the current CU encoded and decoded by SBT, the CCRM decision based on template cost may not be applied.
[0584] ii. For example, for the portion of the current CU with residuals after SBT encoding and decoding, the CCRM decision based on template cost may not be applied.
[0585] iii. For example, if the CBF flag of the current luminance TU is false, then the CCRM decision based on template cost may not be applied.
[0586] e. For example, for a TU generated from a CU encoded and decoded by SBT (e.g., the TU size is smaller than the CU size), the template can be constructed from neighboring samples outside the entire CU.
[0587] f. For example, in inter-frame / IBC modes based on sub-blocks / sub-segments, since each sub-block can have its own motion vector, the motion vectors of predefined sub-blocks can be used to locate the reference template.
[0588] i. For example, for a TU encoded with affine / sbTMVP, the MV of a specific sub-block (e.g., the top left corner or the center) can be used.
[0589] ii. For example, for a TU that has been codified by GPM inter-frame-to-inter-frame encoding, a specific segment of the MV (e.g., part 0 or part 1) can be used.
[0590] iii. For example, for a TU that has been encoded and decoded via GPM inter-intra-frame, the MV of the inter-frame portion can be used.
[0591] iv. For example, for a TU encoded and decoded by GPM, the MV after TM / MMVD can be used.
[0592] 1. Alternatively, the MV prior to TM / MMVD can be used.
[0593] v. Alternatively, if the current TU is encoded and decoded in a sub-block / sub-segment-based inter-frame / IBC mode, then such a template-based approach may not be applied to CCRM on / off decisions.
[0594] g. For example, if sub-block-based CCRM is applied, the samples of the current template predicted by the CCRM model can be constructed based on sub-blocks.
[0595] i. For example, the sample points of the current template predicted by the CCRM model can be constructed by applying multiple CCRM models with boundary sub-blocks (e.g., upper and / or left boundary sub-blocks).
[0596] ii. For example, if the boundary sub-block does not have a valid CCRM model, the corresponding template samples may not be computed for cost calculation.
[0597] 1. For example, alternative sites can be filled with real reconstructed template points.
[0598] h. For example, the samples of the current template predicted by the CCRM model can be constructed from the same CCRM model.
[0599] Figure 30 An example of the current template and reference template involved in CCRM encoding and decoding for the current inter-frame block is shown.
[0600] Figure 31 Examples of the current template and reference template involved in CCRM encoding and decoding for the current IBC block are shown.
[0601] 20) The CCRM model can be applied to the luminance residual block and output the chrominance residual block estimated by CCRM.
[0602] a. For example, the final chromaticity prediction block can be generated by adding the first candidate to the second candidate.
[0603] i. For example, the first candidate can be based on the chromaticity residual block estimated by CCRM, and the second candidate can be based on the chromaticity prediction block before / before CCRM.
[0604] ii. Alternatively, the two candidates can be mixed / fused based on a weighted sum method.
[0605] 1. For example, the weights of the two candidates can be fixed and / or based on predefined rules.
[0606] b. For example, a model can be derived from reference reconstructed samples and applied to the current residual samples.
[0607] i. For example, CCRM model coefficients can be solved / derived based on a set of training samples, where the training samples can reference luminance (non-downsampled or downsampled) and chrominance samples in a reference block.
[0608] ii. For example, the derived model coefficients can be applied to the luminance residual block (unsampled or downsampled) and output the chrominance residual block estimated by CCRM.
[0609] c. For example, different offset values can be used during the CCRM model coefficient derivation and CCRM model application processes. Assume an 8-tap CCRM model consists of 6 spatial luminance samples, a nonlinear term, and an offset term; for model coefficient derivation, the estimated reference chromaticity value is obtained as estChromaVal. ref = c0 (L0 ref – offset ref ) + c1(L1 ref–offset ref )+ c2(L2 ref – offset ref )+ c3(L3 ref – offset ref )+ c4(L4 ref – offset ref )+ c5(L5 ref – offset ref )+ c6 nonlinear((L0 ref +L3 ref +1)>>1) + c7 B ref , of which (L0 ref L5 ref () represents the six brightness reconstruction samples in the reference block, and the nonlinearity is the nonlinear CCCM operator, B ref It is a bias, and offset ref These are block-based variables; for model applications, the estimated current chromaticity residual value is obtained as estChromaResiVal. cur = c0(L0 cur – offset cur ) + c1(L1 cur – offset cur )+ c2(L2 cur – offset cur )+ c3(L3 cur – offset cur )+ c4(L4 cur – offset cur )+ c5(L5 cur – offset cur )+ c6 nonlinear((L0 cur +L3 cur +1)>>1) + c7B cur , of which (L0 cur L5 cur ) represents the six brightness residual samples in the current block, and the nonlinearity is the nonlinear CCCM operator, B cur It is a bias, and offset cur It is a block-based variable.
[0610] i. For example, the offset value can be derived based on at least one training sample.
[0611] 1. For example, the offset can be derived based on specific brightness training samples (e.g., samples at fixed positions in the brightness reference block, such as the top left sample or the center sample value).
[0612] 2. For example, the offset can be derived based on the average / median of at least two training samples (e.g., all brightness training samples).
[0613] ii. For example, the offset can be defined as a fixed constant (e.g., 0).
[0614] iii. For example, the first offset (e.g., offset) ref ) can be used in the derivation of model coefficients (e.g., c0…c7).
[0615] 1. For example, a Gaussian elimination solver can be used to minimize the difference between the reference chromaticity block estimated by the CCRM model (e.g., the model input could be a true reference luminance reconstruction block) and the true reference chromaticity reconstruction block.
[0616] 2. For example, the first offset (e.g., offset) ref The average value of all sample points in the reconstructed block can be derived based on the true reference brightness.
[0617] iv. For example, the second offset (e.g., offset) cur () can be used in the model application process.
[0618] 1. For example, the CCRM model associated with the derived model coefficients can be applied to the current luminance residual block and output an estimated chrominance residual block for the current block.
[0619] 2. For example, the second offset (e.g., offset) cur ) can be fixed to be equal to 0.
[0620] d. For example, different bias values can be used during the CCRM model coefficient derivation process and the CCRM model application process.
[0621] i. For example, the bias value can be derived based on the bit depth of the luminance and chrominance prediction / reconstruction samples in the bitstream (e.g., it can be equal to 1 << (bit depth - 1)).
[0622] 1. Alternatively, it can be equal to a fixed constant (e.g., 0).
[0623] ii. For example, the first bias (e.g., B) ref () can be used in the derivation of model coefficients.
[0624] 1. For example, B refIt can be equal to 1 << (bit depth - 1).
[0625] iii. For example, the second offset (e.g., B) cur () can be used in the model application process.
[0626] 1. For example, B cur It can be fixed to be equal to 0.
[0627] e. For example, whether the CCRM model is used to predict the current chromaticity prediction or the current chromaticity residual can be transmitted via signaling in the bitstream.
[0628] i. Alternatively, it can be implicitly inferred based on decoder information.
[0629] ii. Alternatively, the CCRM model can always be applied to predict the current chromaticity residual.
[0630] 21) The disclosed CCRM pattern can be based on one of the following filters: a. CCLM and / or its variants.
[0631] b. MMLM and / or its variants.
[0632] c. CCCM and / or its variants (e.g., GL-CCCM, non-subsampled CCCM, BVG-CCCM, inter-frame CCCM, intra-frame CCCM, etc.).
[0633] d. GLM and / or its variants.
[0634] e. Any cross-component prediction that uses information from one channel / component to predict information from another channel / component.
[0635] f. Any filter-based prediction, where the filter coefficients are solved based on the correlation between the prediction and / or reconstruction information.
[0636] 22) Block restrictions can be applied to limit the application of specific types of CCP patterns.
[0637] a. For example, CCP mode only allows block sizes that satisfy predefined rules.
[0638] b. For example, syntax elements can only be transmitted via signals when CCP mode is applicable.
[0639] c. For example, if CCP patterns are not allowed, a syntax element can be presumed to have a specific value that indicates that no such CCP pattern is used for such a block.
[0640] d. For example, at least one of the following block restrictions can be applied to the CCRM mode (assuming W represents the block width and H represents the block height): i. W < T1, or, W <= T1, ii. H < T2, or, H <= T2, iii. Min (W,H) > T3, or, Min (W,H) >= T3, iv. Max (W,H) < T4, or, Max (W,H) <= T4, v. W < T5 H, or, W <= T5 H, vi. W > T6 H, or, W >= T6 H, vii. H < T7 W, or, H <= T7 W, viii. H > T8 W, or, H >= T8 W, ix. W H < T9, or W H <= T9.
[0641] x. For example, T1, T2,... T9 can be predefined integer constants.
[0642] e. For example, the CCRM mode is only allowed for small blocks.
[0643] i. For example, it can be allowed for blocks smaller than 4x4, or 8x8, or 16x16, or 32x32.
[0644] ii. For example, it can be allowed for blocks with the number of samples less than 32, or 64, or 128.
[0645] iii. For example, it can be allowed for blocks with the number of samples less than 32, or 64, or 128.
[0646] iv. For example, it may not be allowed for 2xN blocks, where N can be greater than 4 or 8 or 16.
[0647] v. For example, it may not be allowed for Nx2 blocks, where N can be greater than 4 or 8 or 16.
[0648] 23) The disclosed method can be used for a single tree.
[0649] 24) The disclosed method can be used for two trees.
[0650] 25) The disclosed method can be used in inter-frame (such as B or P) stripes.
[0651] 26) The disclosed method can be used in intra-frame (such as I) stripes.
[0652] 27) The “block vector” in the disclosed method can be a “motion vector”.
[0653] 28) The training / reference samples in the disclosed method may refer to the predicted samples and / or reconstructed samples in the training / reference region.
[0654] 29) Whether and / or how the methods disclosed above can be applied can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0655] 30) Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.
[0656] 31) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / dual tree segmentation, color components, and stripe / picture type.
[0657] 2.5 Derivation of the Inter-Frame CCPCCM Model 2.5.1 Issues related to the derivation of the inter-frame CCPCCM model 1) For CCP modes in inter-frame stripes / pictures, copying / inheriting the CCP model from a previously encoded / decoded CCP mode is not permitted. However, inter-frame CCP models can be derived based on CCP models from previously encoded / decoded CCP modes.
[0658] a. In addition, the generated / converted CCP model can be used.
[0659] 2.5.2 Related Solutions The specific embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0660] The term "video unit" or "code-decoder unit" can refer to a picture, strip, slice, code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0661] The term "block" can refer to a codec tree block (CTB), a codec tree unit (CTU), or a codec block (CB).
[0662] The term "CCCM" can refer to intra-frame CCCM mode, inter-frame CCCM mode, IBC CCCM mode, intraTMP CCCM mode, etc. It can be a regular CCCM mode or its variants (e.g., GL-CCCM, CCCM without downsampling, CCRM, etc.).
[0663] The term "CCP" can refer to any cross-component prediction method, such as any kind of LM / CCLM / CCCM / GLM / GL-CCCM. It can be an inter-frame CCP, an intra-frame CCP, or a BV-guided CCP.
[0664] The term "CCP mode" can refer to either CCP mode or CCP Merge mode. Examples of CCP Merge modes include inter-frame CCCM Merge mode, intra-frame CCCM Merge mode, inter-frame CCP Merge mode, intra-frame CCP Merge mode, etc. CCP modes can be based on a single model or multiple models. CCP modes can be based on a single filter or multiple filters.
[0665] The term "CCP model" in CCP mode can be calculated based on decoding information such as neighboring samples or reference samples. It can also be inherited / copied / converted / generated from previously encoded / decoded blocks such as CCP Merge mode, inter-frame CCCM Merge mode, etc. The CCP model can be based on linear, non-linear, or convolutional models.
[0666] In this regard, when talking about block dimensions (such as block width and block height), "W" is used to represent block width and "H" is used to represent block height, and the block can be TU / PU / TU.
[0667] It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0668] 1) The CCP model of the current block can be derived from the previous block encoded and decoded using the CCP model, where the previous block encoded and decoded using the CCP model can be derived based on the block vector (BV).
[0669] a. For example, a BV can be associated with the current block or derived from a previously encoded / decoded block.
[0670] i. For example, a BV can be associated with the current block.
[0671] 1. For example, the current block can be encoded or decoded using IBC or intraTMP mode.
[0672] ii. For example, BV can be derived based on previously encoded or decoded blocks.
[0673] 1. For example, the current block can be encoded and decoded using inter-frame mode.
[0674] 2. For example, the current block can be encoded and decoded using intra-frame mode.
[0675] 3. For example, the current block can be encoded or decoded using IBC or intraTMP mode.
[0676] 4. For example, previously encoded or decoded blocks may be adjacent to or not adjacent to the current block.
[0677] 5. For example, a previously encoded / decoded block may be located in the sample picture / strip / subpicture / slice of the current block or in a different picture / strip / subpicture / slice.
[0678] iii. For example, BV can be derived based on the historical BV associated with the block.
[0679] 1. For example, even if a block is not encoded or decoded by BV, BV can still be stored in association with the block.
[0680] a. For example, if the reference to a block is BV encoded or decoded, its BV can be stored and / or updated in association with the block.
[0681] b. For example, at least one CCP model candidate can be derived from BV.
[0682] i. For example, a candidate CCP model guided by BV can be allowed to be used as the CCP model of the current block.
[0683] 1. For example, a BV-guided candidate can refer to a candidate block retrieved by BV and the current location.
[0684] 2. For example, a BV-guided candidate can refer to a candidate block retrieved by adding the BV to a predefined position relative to the current position.
[0685] ii. For example, a candidate guided by BV can be located in the same picture as the current block.
[0686] 1. Alternatively, BV-guided candidates can be located in the reference image of the current block.
[0687] iii. For example, more than one CCP model candidate can be derived from BV.
[0688] 1. For example, a predefined list of locations can be examined; for instance, a BV can be added as a shift vector to a predefined location. The qualified CCP models associated with the blocks at these locations can be considered as potential CCP model candidates to be used for encoding and decoding the current block.
[0689] 2) The CCP model of the current block can be derived from the previous block encoded and decoded using the CCP model, where the previous block encoded and decoded using the CCP model can be derived based on motion vectors (MV).
[0690] a. For example, an MV can be associated with the current block or derived from a previously encoded / decoded block.
[0691] i. For example, MV can be associated with the current block.
[0692] 1. For example, the current block can be encoded and decoded using inter-frame mode.
[0693] ii. For example, an MV can be derived from a previously encoded or decoded block.
[0694] 1. For example, previously encoded / decoded blocks can be encoded / decoded using inter-frame mode.
[0695] 2. For example, the current block can be encoded or decoded using inter-frame mode, IBC mode, or intra-frame mode.
[0696] 3. For example, previously encoded or decoded blocks may be adjacent to or not adjacent to the current block.
[0697] 4. For example, a previously encoded / decoded block may be located in the sample picture / strip / subpicture / slice of the current block or in a different picture / strip / subpicture / slice.
[0698] iii. For example, MV can be derived based on the historical MV associated with the block.
[0699] 1. For example, even if a block is not encoded or decoded by an MV, the MV can still be stored in association with the block.
[0700] a. For example, if the reference to a block is MV encoded, then its MV can be stored and / or updated in association with the block.
[0701] b. For example, at least one CCP model candidate can be derived from MV.
[0702] i. For example, the CCP model of a reference block in a temporally encoded or decoded image can be used as a CCP model candidate for the current block.
[0703] 1. For example, the position of a reference block can be retrieved by adding a motion shift to a predefined position relative to the current position.
[0704] a. For example, motion shifts can be derived based on the MV of the current block.
[0705] b. For example, motion shifts can be derived based on the MV of the spatial nearest neighbors of the current block.
[0706] c. For example, motion displacement can be a zero vector.
[0707] ii. Furthermore, more than one CCP model candidate can be derived from the temporal reference image.
[0708] 1. For example, at least one motion displacement can be used.
[0709] a. For example, the first available motion shift can be used, and the list of CCP model candidates can be derived by adding the motion shift to the list of predefined positions.
[0710] 2. In addition, for example, more than one motion displacement may be permitted.
[0711] a. For example, both the MV of the current block and the MV of neighboring blocks can be used as motion shifts to derive CCP model candidates, and all qualified CCP model candidates retrieved based on motion shifts are derived.
[0712] 3) The generated / transformed / shifted CCP model can be derived for the current block.
[0713] a. For example, a linear model can be generated / converted from a previously encoded / decoded nonlinear CCP model.
[0714] b. For example, shift candidates can be generated / transformed from a previously encoded / decoded CCP model by shifting at least one model parameter (such as adding an incremental value).
[0715] c. For example, the generated / transformed / shifted linear model can be inserted into the candidate list.
[0716] i. For example, more than one linear model can be generated based on a list of nonlinear model candidates already in the list.
[0717] ii. For example, assuming the maximum allowed length of the candidate list is “M”, and the number of nonlinear model candidates already in the list is “m”, then “MN” linear model candidates can be generated / transformed / shifted and subsequently inserted into the list.
[0718] iii. For example, linear model candidates can be generated / transformed by setting the nonlinear terms of available nonlinear models to zero.
[0719] d. For example, a multi-filter model can be generated / converted from a previously encoded / decoded CCP model.
[0720] i. For example, a multi-filter model can be a weighted sum of more than one single-filter model that is already in the list.
[0721] 1. For example, the weighting factor for each hypothesis can be predefined.
[0722] 2. For example, the weighting factor for each hypothesis can be calculated based on the decoded information.
[0723] 4) For example, the derived CCP model can be inserted into the CCP candidate model list.
[0724] a. For example, deduplication / redundancy checks can be performed by comparing the CCP model to be inserted with the CCP candidate models already available in the list. As long as there are no duplicate CCP models in the list, the CCP model to be inserted can be inserted as a new candidate.
[0725] 5) For example, the CCP candidate model list may include adjacent spatial candidates, and / or historical candidates, and / or generated / transformed / shifted candidates.
[0726] 6) For example, the maximum candidate list size for CCP mode can be defined in the bitstream, but the candidate list for a specific block may not be full.
[0727] a. For example, if the candidate list for a particular block is not full, the list may not be populated with additional candidates.
[0728] b. Alternatively, if the candidate list for a particular block is not full, the list can be filled with additional candidates (such as generated / transformed / shifted candidates) until the list is full.
[0729] 7) For example, the derived CCP model and other CCP model candidates can be reordered / ranked based on the derivation cost / SAD of the decoded model (e.g., based on the difference between the model-predicted chromaticity sample values after applying the candidate CCP model to a set of luminance samples and the actual reconstructed chromaticity sample values corresponding to that set of luminance samples).
[0730] 8) For example, CCP model can refer to inter-frame CCP model.
[0731] a. For example, the CCP model can be based on inter-frame LM / CCLM / CCCM / GLM / GL-CCCM.
[0732] b. Alternatively, for example, the CCP model could refer to the intra-frame CCP model.
[0733] i. For example, the CCP model can be based on intra-frame LM / CCLM / CCCM / GLM / GL-CCCM.
[0734] c. Alternatively, for example, the CCP model could refer to the BV-guided CCP model.
[0735] i. For example, the CCP model could be a BV-guided LM / CCLM / CCCM / GLM / GL-CCCM.
[0736] 9) For example, the current block can be encoded and decoded using the CCP model Merge mode.
[0737] a. For example, the model parameters of the current block can be copied / inherited / derived from previously encoded / decoded blocks, rather than being calculated on the spot.
[0738] b. For example, the current block can be encoded and decoded using the inter-frame CCCM Merge mode.
[0739] c. Alternatively, the current block can be encoded or decoded using the inter-frame CCP Merge mode.
[0740] d. Alternatively, the current block can be encoded or decoded using intra-frame CCP Merge mode.
[0741] 10) The disclosed method can be used for single trees.
[0742] 11) The disclosed method can be used for two trees.
[0743] 12) The disclosed method can be used in inter-frame (such as B or P) stripes.
[0744] 13) The disclosed method can be used in intra-frame (such as I) stripes.
[0745] 14) Whether and / or how the methods disclosed above can be applied can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0746] 15) Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.
[0747] 16) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / dual tree segmentation, color components, and stripe / picture type.
[0748] 2.6 Application of Inter-Frame CCPCCM Mode 2.6.1 Issues related to the application of inter-frame CCPCCM mode 1) The permission of a certain CCP mode may depend on decoding information, such as temporal layer, luma factor, block dimension, etc.
[0749] 2) Signaling for a certain CCP mode can depend on decoding information, such as proximity prediction mode, block dimension, etc.
[0750] 3) The determination of a specific CCP mode can be based on the cost assessment derived from the decoder.
[0751] a. In addition, a bias factor can be introduced for cost assessment.
[0752] 4) Currently, for CCP model calculation and threshold calculation in multi-model CCP mode, all samples in the corresponding luma block are used. However, for inter-frame CCCM mode, a maximum of 256 luma samples are limited to data collection for single-model inter-frame CCCM model parameter calculation. Whether or how such a limitation is applied can be redesigned.
[0753] 5) When the intra-frame CCP Merge mode is used, a candidate sorting / reordering process is applied to sort CCP candidates in ascending order based on template cost. However, diversity criteria can be applied during the sorting process.
[0754] a. If applicable, the diversity criteria can be extended to the inter-frame CCP Merge mode.
[0755] b. If applicable, the diversity standard can be extended to the inter-frame CCCM Merge mode.
[0756] 2.6.2 Related Solutions The specific embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0757] The term "video unit" or "code-decoder unit" can refer to a picture, strip, slice, code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0758] The term "block" can refer to a codec tree block (CTB), a codec tree unit (CTU), or a codec block (CB).
[0759] The term "CCCM" can refer to intra-frame CCCM mode, inter-frame CCCM mode, IBC CCCM mode, intraTMP CCCM mode, etc. It can be a regular CCCM mode or its variants (e.g., GL-CCCM, CCCM without downsampling, CCRM, etc.).
[0760] The term "CCP" can refer to any cross-component prediction method, such as any kind of LM / CCLM / CCCM / GLM / GL-CCCM. It can be an inter-frame CCP, an intra-frame CCP, or a BV-guided CCP.
[0761] The term "CCP mode" can refer to either CCP mode or CCP Merge mode. Examples of CCP Merge modes include inter-frame CCCM Merge mode, intra-frame CCCM Merge mode, inter-frame CCP Merge mode, intra-frame CCP Merge mode, etc. CCP modes can be based on a single model or multiple models. CCP modes can be based on a single filter or multiple filters.
[0762] The term "CCP model" in CCP mode can be calculated based on decoding information such as neighboring samples or reference samples. It can also be inherited / copied / converted / generated from previously encoded / decoded blocks such as CCP Merge mode, inter-frame CCCM Merge mode, etc. The CCP model can be based on linear, non-linear, or convolutional models.
[0763] In this regard, when talking about block dimensions (such as block width and block height), "W" is used to represent block width and "H" is used to represent block height, and the block can be TU / PU / TU.
[0764] It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0765] 1) Whether a specific CCP mode is allowed may depend on the decoding information.
[0766] a. For example, a specific CCP mode can be inter-frame CCCM Merge mode, intra-frame CCCM Merge mode, inter-frame CCP Merge mode, intra-frame CCP Merge mode, etc.
[0767] i. For example, the CCP model can be inherited / copied / converted / generated from previously encoded / decoded blocks.
[0768] b. For example, a specific CCP mode can be an inter-frame CCCM mode, an intra-frame CCCM mode, or LM / CCLM / CCCM / GLM / GL-CCCM, etc.
[0769] i. For example, the CCP model can be calculated based on decoded information (such as neighboring samples or reference samples).
[0770] c. For example, a specific CCP mode can be single-model-based or multi-model-based.
[0771] d. For example, a specific CCP mode can be single-filter-based or multi-filter-based.
[0772] e. For example, whether it is allowed can depend on the dimensions / width / height of a block (such as TU / PU / CU, etc.).
[0773] i. For example, the CCP mode is only allowed to be used for block sizes that meet predefined rules.
[0774] ii. For example, at least one of the following block restrictions can be applied to the CCP mode: 1. a0 W < b0 H, or, a0 W <= b0 H, 2. a1 W > b1 H, or, a1 W >= b1 H, 3. a2 H < b2 W, or, a2 H <= b2 W, 4. a3 H > b3 W, or, a3 H >= b3 W, 5. Min (W, H) > T0, or, Min (W, H) >= T0, 6. Max (W, H) < T1, or, Max (W, H) <= T1, 7. W H < T2, or W H <= T2.
[0775] 8. For example, a0, a1, a2, a3, b0, b1, b2, b3 can be predefined integers.
[0776] 9. For example, T0, T1, and T2 can be predefined values.
[0777] iii. For example, if the current TU / PU / CU satisfies W If H <= T1 (e.g., T1 = 8 or 16), then inter-frame CCCM (and / or inter-frame CCCM Merge mode, and / or multi-model inter-frame CCCM mode) may not be allowed to be used.
[0778] iv. For example, if the current TU / PU / CU satisfies "W>= T2" H" or "H>= T3" If the condition is W (e.g., T2=T3=16 or 8), then inter-frame CCCM (and / or inter-frame CCCM Merge mode, and / or multi-model inter-frame CCCM mode) may not be allowed to be used.
[0779] f. For example, whether it can be allowed depends on the time-domain layer.
[0780] i. For example, a particular CCP mode may be allowed only if the temporal layer ID of the video unit meets certain conditions (such as being less than or greater than a specific value).
[0781] ii. Alternatively, specific CCP modes may be permitted for specific predefined time-domain layers.
[0782] g. For example, whether it is allowed depends on the luminance coefficient.
[0783] i. For example, whether it can be allowed depends on the sum of the luminance coefficients.
[0784] ii. For example, whether it is allowed depends on the scan position with a non-zero luminance coefficient.
[0785] iii. For example, whether it can be allowed depends on the number of non-zero luminance coefficients.
[0786] iv. For example, whether a multi-model CCP mode is allowed can depend on the luminance coefficient.
[0787] v. For example, whether a single-model CCP mode is allowed can depend on the luminance coefficient.
[0788] h. For example, syntax elements associated with a CCP mode can only be transmitted via signals if the CCP mode is permitted.
[0789] i. For example, if a CCP pattern is not allowed, its associated syntax element can be presumed to be a specific value that indicates that no such CCP pattern is allowed / applicable / used for such a block.
[0790] 2) Whether and / or how CCP mode can be transmitted via signal can be based on decoded information.
[0791] a. For example, CCP mode can be based on inter-frame CCCM mode, or intra-frame CCCM mode, or inter-frame CCCM Merge mode, or intra-frame / inter-frame CCP Merge mode, or multi-model inter-frame / intra-frame CCP / CCCM mode.
[0792] b. For example, encoding / decoding information can refer to prediction mode / method, block width / height, luminance sample values, temporal layer of the current block, etc.
[0793] c. For example, whether or how CCP mode is transmitted via signaling may depend on the encoding and decoding information of some specific previously encoded and / or current block.
[0794] i. For example, a particular previously encoded or decoded block may refer to at least one neighboring block of the current block.
[0795] 1. For example, neighboring blocks can be adjacent to and / or not adjacent to the current block.
[0796] 2. For example, a neighboring block can be the TU / PU / CU in the M rows above and / or the N columns to the left of the current block.
[0797] a. For example, M and N can be based on CTU size and / or maximum TU / PU / CU size.
[0798] b. Alternatively, M and N can be fixed values, such as 32 or 16 or 8 or 4 or 2.
[0799] 3. For example, neighboring blocks can be defined based on a predefined inspection order (such as LUTs or rules).
[0800] 4. For example, neighboring blocks can be co-located in the time domain and / or in the same picture of the current block.
[0801] 5. For example, neighboring blocks can be defined based on a history-based table (such as a FIFO table that stores prediction methods for previously encoded / decoded blocks).
[0802] ii. For example, whether a syntax element of mode B is transmitted via signaling can be based on whether some specific previously encoded or decoded block of mode A is utilized.
[0803] 1. For example, mode A can be inter-frame CCCM, and mode B can be a sub-mode of inter-frame CCCM (such as inter-frame CCCM Merge mode or multi-model inter-frame CCCM, etc.).
[0804] 2. For example, mode A can be intra / inter-frame CCP, and mode B can be a sub-mode of intra / inter-frame CCP (such as intra / inter-frame CCP Merge mode, multi-model intra / inter-frame CCP, etc.).
[0805] 3. For example, if there are at least K blocks encoded using pattern A (e.g., adjacent to or preceding the current block), then syntax elements of pattern B can be signaled. Otherwise, pattern B can be presumed not to be used.
[0806] iii. For example, the maximum length of the candidate list size for mode B can be based on whether some specific previously encoded blocks of mode A have been utilized.
[0807] 1. For example, mode A can be inter-frame CCCM, and mode B can be a sub-mode of inter-frame CCCM (such as inter-frame CCCM Merge mode, etc.).
[0808] 2. For example, mode A can be intra / inter-frame CCP, and mode B can be a sub-mode of intra / inter-frame CCP (such as intra / inter-frame CCP Merge mode, etc.).
[0809] 3. For example, if the number of previously encoded / decoded blocks meets certain conditions (such as being less than a certain value), a smaller value can be assigned as the maximum length of the candidate list size.
[0810] a. Alternatively, a larger value may be assigned as the maximum length of the candidate list size, provided that the number of previously encoded / decoded blocks meets certain conditions (such as being greater than a certain value).
[0811] iv. For example, how to transmit a signal via a candidate index in the candidate list of mode B can be based on whether some specific previously encoded blocks of mode A are utilized.
[0812] 1. For example, mode A can be inter-frame CCCM, and mode B can be a sub-mode of inter-frame CCCM (such as inter-frame CCCM Merge mode, etc.).
[0813] 2. For example, mode A can be intra / inter-frame CCP, and mode B can be a sub-mode of intra / inter-frame CCP (such as intra / inter-frame CCP Merge mode, etc.).
[0814] 3. For example, if the number of previously encoded / decoded blocks meets certain conditions (such as being less than a certain value), fewer bits can be allocated to encode / decode the candidate index.
[0815] a. Alternatively, if the number of previously encoded / decoded blocks meets certain conditions (such as being greater than a certain value), more bits can be allocated to encode / decode the candidate index.
[0816] d. For example, whether or how CCP mode is transmitted via signaling can depend on the encoding / decoding information of the current block.
[0817] i. For example, whether a syntax element of signal transmission mode B can be conditional on whether the current mode is encoded or decoded using mode A.
[0818] 1. For example, pattern B can be a subpattern of pattern A.
[0819] 2. For example, whether the flag for inter-frame CCCM Merge mode is transmitted via signal transmission can be conditional upon the presence of the inter-frame CCCM mode flag.
[0820] a. For example, the inter-frame CCCM Merge mode can be considered a sub-mode of the inter-frame CCCM mode.
[0821] b. For example, if the inter-frame CCCM mode flag is true, the inter-frame CCCM Merge flag can be transmitted via signaling.
[0822] i. For example, if the inter-frame CCCM Merge mode flag is equal to 0 and the inter-frame CCCM mode flag is true, it indicates that regular inter-frame CCCM is used instead of inter-frame CCCM Merge mode (e.g., the inter-frame CCCM model parameters are calculated from training samples, rather than inherited).
[0823] 3. For example, whether the flag for transmitting multi-mode inter-frame CCCM mode via signal transmission can be conditional upon the presence of the inter-frame CCCM mode flag.
[0824] a. For example, multi-model inter-frame CCCM mode can be regarded as a sub-mode of inter-frame CCCM mode.
[0825] b. For example, if the inter-frame CCCM mode flag is true, the multi-model inter-frame CCCM flag can be transmitted via signaling.
[0826] i. For example, if the multi-model inter-frame CCCM mode flag is equal to 0 and the inter-frame CCCM mode flag is true, it indicates that a single-model inter-frame CCCM is used.
[0827] ii. Alternatively, the flag of mode B may be transmitted via signaling independently of the presence of mode A.
[0828] 1. For example, even if the inter-frame CCCM flag is equal to 0, the inter-frame CCCM Merge mode can still be used / transmitted via signaling.
[0829] e. For example, whether CCP mode can be transmitted via signaling can be determined at the CU / PU / TU level, CTU level, slice level, sub-picture level, strip (header) level, picture (header) level, picture group level, sequence level, etc.
[0830] f. For example, syntax elements of CCP mode can be encoded and decoded in context.
[0831] i. For example, at least one context model can be used.
[0832] ii. For example, more than one context model can be used.
[0833] 1. For example, which context model to use may depend on the width and / or height of the block (e.g., TU / PU / CU).
[0834] a. For example, if the block size satisfies the condition W>a H or H>b W, then a specific context model can be used, where a and b are constants, for example a=b=2.
[0835] 2. For example, which context model to use can depend on the prediction patterns of neighboring blocks.
[0836] a. For example, the context model of pattern B may depend on whether a particular neighbor (left and / or above) is encoded or decoded using pattern A.
[0837] i. For example, which context model is used for the inter-frame CCCM Merge mode flag can depend on whether the left and / or upper neighbors are being encoded or decoded using the inter-frame CCCM Merge mode.
[0838] ii. For example, which context model is used for the multi-model inter-frame CCCM mode flag can depend on whether the left and / or upper neighboring frames are being encoded or decoded using the multi-model inter-frame CCCM mode.
[0839] b. For example, it can depend on whether the left and / or upper neighbors are encoded or decoded using mode B.
[0840] i. For example, which context model is used for the inter-frame CCCM Merge mode flag can depend on whether the left and / or upper neighboring frames are being encoded or decoded using inter-frame CCCM mode.
[0841] ii. For example, which context model is used for the multi-model inter-frame CCCM mode flag can depend on whether the left and / or upper neighboring frames are being encoded or decoded using inter-frame CCCM mode.
[0842] 3) Bias factors can be used to determine the use of a specific CCP mode.
[0843] a. For example, assuming that the determination of the use of a particular CCP mode is based on a cost comparison derived from the decoder, a bias factor can be introduced for cost evaluation.
[0844] i. For example, the derivation cost of decoding for a specific pattern can be derived based on the difference between the true reconstructed sample value and the pattern prediction sample value on a predefined region of the sample.
[0845] 1. For example, the predefined region of a sample point can be neighboring sample points around the current block.
[0846] 2. Alternatively, the predefined region of a sample point can be a reference sample point in a reference block in the current image (e.g., a BV-guided reference block).
[0847] 3. Alternatively, the predefined area of the sample point can be a reference sample point in a reference block in a reference image (e.g., a reference block guided by an MV).
[0848] ii. For example, whether to use the first mode or the second mode can be determined by a cost evaluation derived from the decoder.
[0849] 1. For example, the values of the derivation costs of the two decoders can be calculated for each mode, such as by cost A and cost B, and whether to select the first mode can be determined by whether the following condition is true.
[0850] a. ((a Cost A + Offset >> Shift < Cost B, where a / offset / shift is the bias factor.
[0851] 2. For example, the first mode can be a multi-model CCP mode, while the second mode can be a single-model CCP.
[0852] a. For example, in addition, CCP can be inter-frame CCCM.
[0853] 3. For example, the first mode can be a multi-filter CCP mode, while the second mode can be a single-filter CCP mode.
[0854] a. For example, in addition, CCP can be inter-frame CCCM.
[0855] b. For example, in addition, CCP can be an intra-frame CCCM.
[0856] iii. For example, the value of the bias factor can depend on the decoded information.
[0857] 1. For example, it can be based on the time domain layer.
[0858] 2. Alternatively, the bias factor can be set to a predefined constant.
[0859] 4) Luminance samples used for CCP model calculations can be downsampled based on block size.
[0860] a. For example, CCP mode can be inter-frame CCCM mode, inter-frame CCP mode, or intra-frame CCP mode.
[0861] b. For example, the downsampling factor can depend on the block size.
[0862] i. For example, the downsampling factor may not depend on the chroma format.
[0863] ii. For example, a larger downsampling factor can be assigned to larger blocks.
[0864] iii. For example, the downsampling factor can be predefined based on the block width and / or height (e.g., based on LUT).
[0865] c. For example, downsampled luminance blocks can be used for CCP model calculations.
[0866] i. For example, downsampled blocks can be used for data collection to compute CCP model parameters.
[0867] ii. For example, downsampled blocks can be used to compute thresholds for multi-model sample point classification / classification.
[0868] d. For example, the maximum allowed number of samples calculated for a specific CCP model can be equal to X (where X is a predefined constant, such as X = 256 or 1024 or 4096).
[0869] i. For example, regardless of the original block size (e.g., even for 256x256, 128x128, or 64x64 blocks), the number of samples used for calculation of a particular CCP model may not exceed the value of X.
[0870] ii. For example, alternatively, different X values can be used for different CCP modes.
[0871] 1. For example, X1 can be used for single-model inter-frame CCCM mode; while X2 can be used for multi-model inter-frame CCCM mode, where X2>= X1.
[0872] 2. For example, X1 can be used for single-model inter-frame CCP mode; while X2 can be used for multi-model inter-frame CCP mode, where X2>= X1.
[0873] iii. Alternatively, no limit may be applied to the maximum number of samples calculated for a particular CCP model.
[0874] 1. For example, the complete sample points of a block (without downsampling) can be used for CCP model calculations.
[0875] 2. For example, X1 can be used for single-model inter-frame CCCM mode; while for multi-model inter-frame CCCM mode, there are no restrictions on its use.
[0876] 3. For example, X1 can be used for single-model inter-frame CCP mode; while for multi-model inter-frame CCP mode, there are no restrictions on its use.
[0877] 4. For example, no limit may be applied to the maximum number of samples that can be calculated for a particular CCP model, regardless of whether it is a multi-model or single-model CCP / CCCM model calculation.
[0878] 5) Diversity criteria can be applied to reorder / rank multiple candidate CCP patterns.
[0879] a. For example, CCP model candidates in the candidate list can be reordered / ranked according to diversity criteria.
[0880] b. For example, diversity criteria could be based on the cost difference between the current candidate and its predecessor in the list.
[0881] i. For example, cost can refer to the cost deduced by the decoder based on decoding information (e.g., template cost, reference block prediction cost, etc.).
[0882] ii. For example, diversity candidates can be placed at the beginning of the list.
[0883] 1. In addition, for example, redundant candidates can be placed after non-redundant candidates.
[0884] iii. For example, if the minimum cost difference between the current candidate and its predecessor is less than a threshold, the current candidate can be considered redundant.
[0885] 1. For example, the threshold can depend on the Lagrange parameter.
[0886] 2. For example, thresholds can be predefined according to rules.
[0887] 3. For example, the threshold can be a constant.
[0888] 4. For example, the threshold can depend on the decoding information (such as the temporal layer, block dimension, etc.).
[0889] c. For example, the final CCP model can be selected from a reordered / sorted list of candidates for encoding and decoding the current block.
[0890] i. In addition, for example, indicators (e.g., indexes) can be transmitted via signals in the bitstream to specify the final candidate to be selected from the entire candidate list.
[0891] d. For example, CCP Merge modes (e.g., intra-frame CCP Merge mode, inter-frame CCP Merge mode, inter-frame CCCM Merge mode, etc.) can be applied based on a reordered / sorted candidate list.
[0892] 6) Whether to use a model calculated on the fly or a model derived from it can be determined based on the method of derivation by the decoder (rather than encoder selection).
[0893] a. For example, the CCP model can be calculated on the fly (e.g., the CCP model can be calculated from reference luminance samples and reference chrominance samples), or it can be derived / inherited from the CCP model associated with a previously encoded / decoded CCP block.
[0894] b. For example, whether to use a computed CCP model or a deduced / inherited CCP model for the current block (or, which CCP model will be used) may depend on the cost evaluation method deduced by the decoder.
[0895] i. For example, cost evaluation methods derived from decoder derivation can be applied based on decoder-side information such as reference samples.
[0896] ii. For example, given a candidate model, predicted chromaticity samples can be obtained by applying the candidate model to reference luminance samples, and the cost can be calculated by summing the differences between the predicted and reconstructed chromaticity samples. In this way, each candidate model can calculate its own cost.
[0897] iii. For example, comparisons can be made based on comparing the cost of the model computed on the fly with the cost of the model derived / inherited, and the one with the lower cost can be determined as the final model for the current block.
[0898] iv. For example, comparisons can be made based on comparing the cost of the first derivation / inheritance model and the cost of the second derivation / inheritance model, and the one with the lower cost can be determined as the final model for the current block.
[0899] c. For example, no syntax element may be transmitted via signaling to indicate whether a computed CCP model or a deduced / inherited CCP model is used for the current CCP-encoded block (or, which CCP mode will be used).
[0900] i. For example, it can be determined on the decoder side.
[0901] ii. For example, rate-distortion optimization (RDO)-based encoder search may not be necessary.
[0902] d. For example, the CCP model for a block encoded and decoded via inter-frame CCCM can be calculated from a reference sample or derived / inherited from a previous block encoded and decoded via inter-frame CCCM, but it is not necessary to know where the signal transmission model comes from.
[0903] i. For example, only one flag (e.g., the inter-frame CCCM flag) can be signaled to specify whether the current block is encoded or decoded using inter-frame CCCM. However, it is not necessary to signal additional flags to specify whether it is an on-the-fly computation model or an inherited / derived model.
[0904] 7) The presence (or how syntax elements are signaled) of syntax elements (e.g., pattern flags, candidate indices, etc.) can depend on whether there are neighbors (or how many neighbors) that are encoded / decoded using a particular prediction pattern / method.
[0905] a. For example, whether or not a signal transmission mode flag is used can depend on this.
[0906] i. For example, if no neighboring code is used for predictive mode / method A, the mode flag for predictive mode / method B of the current block may not be transmitted via signaling (e.g., it is presumed to be equal to 0).
[0907] ii. For example, if the number of adjacent blocks above (and / or to the left) of the block encoded using prediction mode / method A is less than a certain value, the mode flag for prediction mode / method B of the current block may not be transmitted via signaling (e.g., it is presumed to be equal to 0).
[0908] b. For example, whether the prediction mode / method for the current block is allowed may depend on this.
[0909] i. For example, if the neighboring codecs of prediction mode / method A are not utilized, prediction mode / method B may not be applied to the current block.
[0910] ii. For example, if the number of adjacent blocks above (and / or to the left) of the block encoded using prediction mode / method A is less than a certain value, prediction mode / method B may not be applied to the current block.
[0911] c. For example, how candidate indices are transmitted via signals may depend on this.
[0912] i. For example, the maximum allowed number of Merge candidates can depend on this.
[0913] ii. For example, if the number of above (and / or left) neighbors encoded using a specific prediction mode / method is equal to K0, then the maximum allowed number of merge candidates for the current block can be equal to L0. Otherwise, if the number of above (and / or left) neighbors encoded using a specific prediction mode / method is equal to K1, then the maximum allowed number of merge candidates for the current block can be equal to L1, where K0, K1, L0, and L1 are predefined values.
[0914] iii. For example, the binarization process may depend on this.
[0915] d. For example, the context model for syntax signaling can depend on this.
[0916] i. For example, which context model is used can depend on how many neighbors are used to encode / decode the prediction pattern / method A.
[0917] e. For example, the size of the Merge list can depend on this.
[0918] i. For example, if the number of above (and / or left) neighbors encoded using a specific prediction mode / method is equal to K0, then the Merge list size of the current block can be equal to L0. Otherwise, if the number of above (and / or left) neighbors encoded using a specific prediction mode / method is equal to K1, then the Merge list size of the current block can be equal to L1, where K0, K1, L0, and L1 are predefined values.
[0919] f. For example, the neighbors to be checked can be based on the neighbors in the M rows above and / or the N columns to the left of the current TU / PU / CU.
[0920] g. For example, the neighbors to be examined can be based on spatial adjacency and / or non-adjacency and / or candidates based on temporal domain.
[0921] i. For example, neighbors can be checked based on the candidate check order of the Merge list.
[0922] 1. For example, deduplication may not be necessary.
[0923] 2. For example, no reconstructed samples may be inspected.
[0924] 3. For example, no candidates may be checked.
[0925] h. For example, the presence of syntax flags for pattern B can depend on the encoding / decoding information of neighboring blocks (such as whether neighbors are encoded / decoded using pattern A).
[0926] i. For example, pattern B can be a subpattern of pattern A.
[0927] ii. For example, pattern B can be equivalent to pattern A.
[0928] iii. For example, the presence of the CCP Merge / submode flag may depend on whether neighboring blocks are encoded or decoded using CCP mode.
[0929] iv. For example, the presence of the inter-frame CCCM Merge / submode flag can depend on whether neighboring blocks are encoded or decoded using inter-frame CCCM mode.
[0930] v. For example, the presence of the CIIP TM / submode flag can depend on whether neighboring blocks are encoded or decoded using inter-frame / CIIP / TM mode.
[0931] vi. For example, the presence of the DIMD Merge / submode flag can depend on whether neighboring blocks are encoded or decoded using intra-frame / DIMD / TIMD modes.
[0932] vii. For example, the presence of the TIMD Merge / submode flag can depend on whether neighboring blocks are encoded or decoded using intra-frame / DIMD / TIMD modes.
[0933] 8) Multiple assumptions: CCP can be applied to video blocks.
[0934] a. For example, the prediction of chroma blocks can be generated based on the mixing / fusion of the first CCP prediction and the second CCP prediction.
[0935] i. For example, the first CCP prediction and the second CCP prediction can be different.
[0936] b. For example, chroma block predictions can be generated based on a mixture / fusion of regular inter-frame (or intra-frame) predictions and CCP predictions.
[0937] c. For example, CCP prediction can be generated based on inter-frame CCCM mode, inter-frame CCP mode, intra-frame CCP mode, etc.
[0938] d. For example, CCP predictions can be generated based on CCP models of previously encoded and decoded CCP blocks (such as spatially adjacent, non-adjacent, history-based, and time-based CCP candidates).
[0939] e. For example, CCP predictions can be generated based on a CCP model calculated using neighboring sample information (or reference sample information).
[0940] f. For example, a weighted sum of multiple hypotheses can be considered as the final prediction of a video block.
[0941] i. For example, average weighting (e.g., equal weights) can be applied.
[0942] ii. For example, fixed weighting factors (e.g., predefined unequal weights) can be used.
[0943] iii. For example, weights can be determined based on reconstruction / decoding information of neighboring samples / blocks.
[0944] 1. For example, if the energy of the residual signal in the decoded region is low, a higher weight can be assigned to the hypothesis (e.g., the energy of the residual signal in the decoded region can be calculated based on the sum of the absolute differences of neighboring sample regions with and without the hypothesis).
[0945] iv. For example, weights can be derived based on syntactic information.
[0946] 9) The disclosed method can be used for single trees.
[0947] 10) The disclosed method can be used for two trees.
[0948] 11) The disclosed method can be used in inter-frame (such as B or P) stripes.
[0949] 12) The disclosed method can be used in intra-frame (such as I) stripes.
[0950] 13) Whether and / or how the methods disclosed above can be applied to be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0951] 14) Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.
[0952] 15) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / dual tree segmentation, color components, and stripe / picture type.
[0953] 3 questions In the current ECM, intra-blocks can be encoded and decoded using the intra-CCP model, where intra-CCP model parameters can be inherited from neighboring intra-codec blocks or calculated on the fly from a reference region consisting of luminance and chrominance reconstructed samples adjacent to the current block.
[0954] Furthermore, inter-frame or IBC blocks can be encoded and decoded using an inter-frame CCCM model, where the parameters of the inter-frame CCCM model are calculated in real time from a reference region consisting of luminance and chrominance samples in the inter-frame or IBC prediction block.
[0955] Furthermore, due to the BV-guided CCCM, IBC or intraTMP blocks can be encoded and decoded using the BV-guided CCCM model, where the BV-guided CCCM model parameters are calculated in real time from the reference area pointed to by the block vector of the co-position luminance block.
[0956] The above design is insufficient. Theoretically, CCP model parameters from previous codec blocks can be reused / inherited for subsequent codec blocks, regardless of whether the subsequent codec block is intra-frame, inter-frame, or IBC-encoded. Furthermore, given a codec block and a reference region, different types of CCP models can be computed from the reference region, regardless of whether the codec block is intra-frame, inter-frame, or IBC-encoded.
[0957] In the current ECM, CCP model computation can use samples from the current block, and CCP model application can use samples from the reference region. The use of samples from the current block and the use of the reference region can be decoupled.
[0958] 4 Specific Solutions The specific embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0959] The term "video unit" or "code-decoder unit" can refer to a picture, strip, slice, code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0960] The term "block" can refer to code-decode tree block (CTB), code-decode tree unit (CTU), code-decode block (CB), CU, PU, TU, PB, TB.
[0961] The term "CCP" can refer to any cross-component prediction method, such as any kind of LM / CCLM / MMLM / CCCM / GLM / GL-CCCM. It can be an inter-frame CCP, an intra-frame CCP, or a BV-guided CCP.
[0962] The term "CCP mode" can refer to either CCP mode or CCP Merge mode. Examples of CCP Merge modes include inter-frame CCCM Merge mode, intra-frame CCCM Merge mode, inter-frame CCP Merge mode, intra-frame CCP Merge mode, etc. CCP mode can be based on a single model or multiple models. CCP mode can be based on a single filter or multiple filters. CCP mode can be applied to inter-frame blocks, intra-frame blocks, or IBC blocks.
[0963] The term "CCP model" in CCP mode can be computed based on decoded information such as neighboring samples or reference samples. It can also be inherited / copied / transformed / generated from previously encoded / decoded blocks. CCP models can be based on linear, nonlinear, or convolutional models.
[0964] The term "CCCM" can refer to intra-frame CCCM mode, inter-frame CCCM mode, BVG CCCM mode, IBC CCCM mode, intraTMP CCCM mode, etc. It can be a regular CCCM mode or its variants (e.g., GL-CCCM, CCCM without downsampling, CCRM, etc.). CCCM mode can be based on a single model or multiple models.
[0965] In this regard, when talking about block dimensions (such as block width and block height), "W" is used to represent block width and "H" is used to represent block height, and the block can be TU / PU / TU.
[0966] It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0967] 1) Given the current inter-frame (or IBC) block, a specific type of CCP model can be used, and the CCP model parameters can be calculated from the reference region.
[0968] a. For example, the CCP model can be in the style of the intraCCP model, such as LM / CCLM / MMLM / GLM / CCCM / intraCCCM and / or its variant models.
[0969] b. For example, the CCP model could be inter-frame CCCM.
[0970] c. For example, CCP model parameters can be calculated on the fly from the reference region.
[0971] d. For example, the reference region may consist of luminance and chrominance reconstructed samples that are adjacent to the current block (e.g., within the current frame).
[0972] i. For example, the reference area can consist of up to M columns and N rows of luminance and chromaticity reconstructed samples from the top and left sides of the current block, such as M=N=6.
[0973] e. For example, a reference region may consist of luminance and chrominance reconstructed samples based on block vector retrieval (e.g., within the current frame).
[0974] i. For example, block vectors can be derived based on previously encoded or decoded IBC or intraTMP blocks.
[0975] ii. For example, a reference region may consist of luminance and chrominance samples of a reference block to which the block vector points.
[0976] f. For example, the reference region may consist of luminance and chromaticity reconstructed samples based on motion vector retrieval (e.g., within the reference frame).
[0977] i. For example, motion vectors can be derived based on the current block or previously encoded inter-frame blocks.
[0978] ii. For example, the reference region may consist of luminance and chromaticity samples of a reference block to which the motion vector points.
[0979] g. For example, the calculated CCP model may be based on at least one of the following models (or variations thereof): i. Models based on linear regression models / filters ii. Models based on convolutional models / filters iii. Model based on extrapolation filter iv. Single-model LM v. Multi-model LM vi. Single-model GLM (e.g., considering gradient information) vii. Multi-model GLM (e.g., considering gradient information) viii. Single-model intra-frame - CCCM ix. Multi-model intra-frame - CCCM x. Single-model intra-frame CCCM without lumen downsampling xi. Multi-model intra-frame CCCM without lumen downsampling xii. Single-model intra-frame GL-CCCM (e.g., considering gradient and position information) xiii. Multi-model intra-GL-CCCM (e.g., considering gradient and position information) xiv. Single-model intra-frame MDF-CCCM (e.g., considering multiple filters) xv. Multi-model intra-frame MDF-CCCM (e.g., considering multiple filters) xvi. Single-model inter-frame - CCCM xvii. Multi-model Inter-frame - CCCM xviii. Single-model inter-frame without lumen downsampling - CCCM xix. Multi-model inter-frame without lumen downsampling - CCCM xx. Single-model inter-frame GL-CCCM (e.g., considering gradient and position information) xxi. Multi-model inter-frame GL-CCCM (e.g., considering gradient and position information) xxii. Single-model inter-frame MDF-CCCM (e.g., considering multiple filters) xxiii. Multi-model inter-frame - MDF - CCCM (e.g., considering multiple filters) xxiv. The model above can be based on both the upper and left neighboring samples / templates (e.g., a TL-based model). xxv. The model above can be based on the upper neighboring sample / template or the left neighboring sample / template (e.g., a T-based pattern or an L-based pattern). xxvi. The chroma predictions from the model above can be further fused / blended with the second prediction to form the final prediction block for encoding and decoding. a. For example, CCP model parameters can be derived / inherited from previously encoded / decoded inter-frame blocks (e.g., TU / CU / PU).
[0980] i. For example, the previously encoded / decoded block could be encoded / decoded via inter-frame CCCM.
[0981] ii. For example, a previously encoded block may not be encoded or decoded by inter-frame CCCM, but it has a reference block encoded or decoded by inter-frame CCCM or BVG CCCM.
[0982] iii. For example, a previously encoded block may not be encoded or decoded by inter-frame CCCM, but it has a reference block encoded or decoded by intra-frame CCP (such as LM / CCLM / MMLM / GLM / CCCM).
[0983] b. For example, CCP model parameters can be derived / inherited from previously encoded / decoded intra-blocks.
[0984] i. For example, previously encoded / decoded blocks can be encoded / decoded using a type of intra-frame CCP (such as LM / CCLM / MMLM / GLM / intraCCCM or variants thereof).
[0985] c. For example, CCP model parameters can be derived / inherited from previously encoded / decoded IBC blocks (e.g., TU / CU / PU).
[0986] i. For example, the previously encoded / decoded block could be encoded / decoded using BVG CCCM.
[0987] ii. For example, the previously encoded / decoded block may have been encoded / decoded via inter-frame CCCM.
[0988] iii. For example, a previously encoded block may not be encoded or decoded by BVG or inter-frame CCCM, but it has a reference block encoded or decoded by inter-frame CCCM or BVG CCCM or intra-frame CCP (such as LM / CCLM / MMLM / GLM / CCCM).
[0989] d. For example, CCP model parameters can be derived / inherited based on spatially adjacent neighboring blocks or spatially non-adjacent neighboring blocks.
[0990] e. For example, CCP model parameters can be derived / inherited based on time-domain blocks.
[0991] i. For example, a temporal block can be pointed to by a motion shift (e.g., the current MV, neighboring MVs, etc.).
[0992] ii. For example, a time-domain block can be in the same position as the current block.
[0993] f. For example, CCP model parameters can be derived / inherited based on a history-based lookup table.
[0994] i. For example, historical candidates could be inter-frame CCCM models.
[0995] ii. For example, historical candidates can be intra-frame CCP models (such as LM / CCLM / MMLM / GLM / intraCCCM or their variants).
[0996] g. For example, the derived / inherited CCP model can be applied using the current inter-frame (or IBC) luminance reconstruction block as input, and then output an estimated / predicted inter-frame (or IBC) chrominance prediction block.
[0997] 3) Given the current intra-frame block, the inter-frame CCCM model can be used, and the parameters of the inter-frame CCCM model can be derived / inherited from the previous inter-frame CCCM model.
[0998] a. For example, the current intra-frame block may belong to a B / P stripe.
[0999] b. For example, the previous inter-frame CCCM model could have come from inter-frame codec blocks.
[1000] c. For example, previous inter-frame CCCM models could be derived based on spatially adjacent neighboring blocks or spatially non-adjacent neighboring blocks.
[1001] d. For example, previous inter-frame CCCM models could be derived based on temporal blocks.
[1002] i. For example, a temporal block can be pointed to by a motion shift (e.g., the current MV, neighboring MVs, etc.).
[1003] ii. For example, a time-domain block can be in the same position as the current block.
[1004] e. For example, previous inter-frame CCCM models could be derived based on history-based lookup tables.
[1005] i. For example, the inter-frame CCCM model in the history table can come from previously encoded and decoded inter-frame blocks.
[1006] f. For example, the derived / inherited inter-frame CCCM model can be applied using the current intra-frame luma reconstruction block as input, and then outputs an estimated / predicted intra-frame chroma prediction block.
[1007] 4) For the current inter-frame / IBC / intra-frame block, more than one type of CCP model can be calculated / inherited / derived.
[1008] a. For example, models of the type LM / CCLM and / or MMLM and / or their variants can be computed.
[1009] b. For example, models of the type GLM and / or its variants can be computed.
[1010] c. For example, type intra-frame CCCM and / or multi-model intra-frame CCCM and / or its variant models can be computed.
[1011] d. For example, inter-frame CCCM and / or its variant models can be computed.
[1012] e. For example, BVG CCCM and / or its variant models can be computed.
[1013] f. For example, the CCP model can be derived / inherited from a previously encoded / decoded block, and the CCP model is associated with that block.
[1014] g. For example, the CCP model can be calculated (instantly) from the reference region.
[1015] h. For example, the same / uniform reference region can be used to calculate model parameters for these different types of CCP models.
[1016] i. For example, all LM, CCLM, MMLM, GLM, intraCCCM, interCCCM, and BVGCCCM model parameters can be calculated based on the same reference region based on inter-frame (or IBC) prediction.
[1017] ii. For example, all LM, CCLM, MMLM, GLM, intraCCCM, interCCCM, and BVGCCCM model parameters can be calculated based on the same reference region based on adjacent samples.
[1018] i. Alternative locations: Different reference regions can be used for calculating parameters of different types of CCP models.
[1019] i. For example, the inter-frame CCCM model can use a reference region based on inter-frame prediction, the BVG CCCM model can use a reference region based on IBC prediction, and the LM / CCLM / MMLM / GLM / intraCCCM model can use a reference region based on neighboring samples.
[1020] 5) Different types of CCP model candidates for the current block can be inserted into the candidate model list.
[1021] a. For example, only inherited / derived CCP models can be inserted into the candidate list (i.e., these on-the-fly computed candidates may not be inserted into such a candidate list).
[1022] b. For example, only CCP models computed on the spot can be inserted into the candidate list (i.e., these inherited / derived candidates may not be inserted into such a candidate list).
[1023] c. For example, both inherited / derived CCP models and just-in-time computed CCP models can be inserted into a single candidate list.
[1024] d. For example, one of these candidates in the candidate list can be selected for encoding and decoding the current block.
[1025] e. For example, which type of CCP model will ultimately be used for the current block can be transmitted via signaling in the bitstream.
[1026] i. For example, model indexes can be transmitted via signals.
[1027] ii. For example, some pattern flags can be transmitted via signals.
[1028] f. For example, which CCP model will ultimately be used for the current block can be deduced at both the encoder and decoder sides (e.g., without signal transmission).
[1029] i. For example, costs / rules / standards based on decoded information can be used.
[1030] g. For example, which CCP model is ultimately used for the current block can be based on predefined rules (e.g., without signaling).
[1031] i. For example, a list of candidate models can be constructed, sorted according to rules, and then candidates at a fixed order in the list (e.g., first order) can be ultimately selected.
[1032] 6) CCP model candidates can be sorted according to rules.
[1033] a. For example, the cost of decoder derivation can be computed for each unranked model candidate.
[1034] i. For example, the cost of sorting can be derived based on a template.
[1035] 1. For example, a template can refer to the brightness reconstruction samples in columns M and rows N at the top and left of the current block, such as M=N=1 or 2 or 3.
[1036] ii. Alternatively, the cost for ordering can be derived based on inter-frame / IBC prediction blocks.
[1037] iii. For example, each candidate to be ranked can follow the same cost evaluation criteria to calculate its own cost.
[1038] iv. For example, among all these candidates to be sorted, the model candidate with the lowest cost can be selected for the next step of processing.
[1039] b. For example, only inherited / derived CCP model candidates can be sorted (i.e., inter-frame CCCM models computed on the fly can be left unsorted).
[1040] c. For example, only CCP models that are computed on the fly can be sorted.
[1041] d. For example, both inherited / derived CCP models and just-in-time computed CCP models can be sorted.
[1042] i. For example, inherited / derived CCP models can be sorted according to rule A, while these just-in-time computed CCP models can be sorted according to rule B.
[1043] 1. For example, rule A can calculate the cost based on a reference region consisting of inter-frame / IBC prediction blocks.
[1044] 2. For example, rule B can calculate the cost based on a reference region composed of templates.
[1045] ii. For example, both inherited / derived CCP models and just-in-time computed CCP models can be sorted according to the same rules.
[1046] 1. For example, the cost of all candidates can be calculated based on the same reference region (e.g., inter-frame prediction block, IBC prediction block, template with adjacent neighboring samples, etc.).
[1047] e. For example, a sorting process with more than one round can be used in the CCP pattern.
[1048] i. For example, in order to derive the final CCP model for a block that has been encoded and decoded by inter-frame CCP, more than one round of sorting process can be applied.
[1049] ii. For example, suppose CCP model candidates are classified into two groups. The first group of CCP model candidates is derived based on the motion information of the current block, while the second group of CCP model candidates is derived without the motion information of the current block. The first group of CCP model candidates may not be sorted / reordered together with the second group of CCP model candidates.
[1050] a. For example, the first group of CCP model candidates may include at least one of the following CCP model candidates: i. Real-time computed CCP models (e.g., inter-frame CCCM, intra-frame CCCM, intra-frame CCLM, intra-frame GLM, etc.) ii. CCP model derived from spatially adjacent blocks iii. CCP model derived based on spatially non-adjacent blocks iv. CCP model derived from history-based lookup table derivation v. CCP models derived from temporal blocks using motion information from neighboring blocks (e.g., temporal shift candidates). vi. Default CCP model (e.g., generated from existing CCP candidates using predefined rules, based on CCLM, not based on the motion information of the current block, etc.) b. For example, the second group of CCP model candidates may include CCP models derived from time-domain blocks through motion information of the current block.
[1051] c. For example, the first group of CCP model candidates can be sorted together in the first round of sorting.
[1052] i. For example, the M candidates with the lowest cost after the first round of ranking can be considered as input and further ranked together with the second group of CCP model candidates in the second round of ranking.
[1053] d. For example, the second group of CCP model candidates can be ranked together in the first round of ranking.
[1054] i. For example, the N candidates with the lowest cost after the first round of sorting can be considered as input and further sorted together with the first group of CCP model candidates in the second round of sorting.
[1055] f. For example, a temporal CCP model candidate for CCP mode can be derived without motion information of the current block.
[1056] i. For example, CCP mode can refer to inter-frame CCP mode.
[1057] g. For example, the real-time computed inter-frame CCCM model may not be ordered together with those CCP models derived from previously encoded / decoded blocks.
[1058] h. For example, real-time computed inter-frame CCCM models may not be sorted together with these real-time computed intra-frame CCP models.
[1059] i. Alternatively, the real-time computed inter-frame CCCM model may be ordered together with other CCP models (e.g., the real-time computed intra-frame CCP model and / or the intra-frame CCP model derived from previously encoded / decoded blocks and / or the intra-frame CCP model, etc.). 7) The CCP model of the current block can be stored in a cache (e.g., image cache, history table) and can be used as a model candidate for encoding and decoding the next one.
[1060] a. For example, the current block can be intra-frame encoded or decoded.
[1061] i. For example, the stored CCP model can be an intra-frame CCP (such as an LM / CCLM / MMLM / GLM / CCCM-based model).
[1062] ii. For example, the stored CCP model can be an inter-frame CCCM or a BV-guided CCCM model.
[1063] b. For example, the current block can be non-intra-frame (e.g., inter-frame or IBC) encoded / decoded.
[1064] i. For example, the stored CCP model can be an inter-frame CCCM or a BV-guided CCCM model.
[1065] ii. For example, the stored CCP model may be a model based on intra-frame CCP (such as LM / CCLM / MMLM / GLM / CCCM).
[1066] c. For example, each stored piece of information may include CCP model information and its index.
[1067] d. For example, each stored piece of information may include CCP model information and its corresponding block / sub-block location / index / ID.
[1068] e. For example, information in the cache can be accessed for encoding and decoding in the next step.
[1069] i. For example, the next one could belong to the current image.
[1070] ii. For example, the next one could be an inter-frame codec block.
[1071] iii. For example, the next one could be an intra-frame codec block.
[1072] f. For example, information in the cache can be accessed for encoding and decoding of the next image.
[1073] i. For example, an intra-frame CCP model can be used as a temporal model candidate for encoding and decoding the next picture patch.
[1074] 1. For example, the next image block could be an inter-frame codec block.
[1075] 2. For example, the next image block can be an intra-frame codec block.
[1076] ii. For example, the inter-frame CCCM model can be used as a temporal model candidate for encoding and decoding the next picture patch.
[1077] 1. For example, the next image block could be an inter-frame codec block.
[1078] 2. For example, the next picture block can be an intra-frame codec block in a B / P stripe.
[1079] 8) For example, given a block that is encoded and decoded via inter-frame / IBC, the CCP model can be stored in association with such a block even if it is not encoded and decoded by the CCP model.
[1080] a. For example, models of inter-frame-MV (or IBC-BV) propagation can be stored in association with such blocks.
[1081] i. For example, models of inter-frame-MV propagation can be retrieved based on the motion vectors of such inter-frame blocks.
[1082] ii. For example, the model of IBC-BV propagation can be retrieved based on the block vector of such IBC blocks.
[1083] iii. For example, the CCP model of a reference block pointed to by an inter-frame-MV (or IBC-BV) (e.g., assuming the reference block is encoded and decoded using the CCP model) can be stored in association with such a block.
[1084] b. For example, the stored CCP model can be an inter-frame CCCM or a BV-guided CCCM model.
[1085] c. For example, the stored CCP model can be an intra-frame CCP (such as an LM / CCLM / MMLM / GLM / CCCM-based model).
[1086] 9) For a given block (such as CU, PU, or TU), different components can be encoded and decoded using different encoding and decoding modes.
[1087] a. For example, a given block can be encoded and decoded using a single-tree structure.
[1088] b. For example, for a given block (such as CU, PU, or TU), the luminance component can be encoded and decoded using a first mode, and the chrominance component can be encoded and decoded using a second mode.
[1089] i. For example, the first mode can be AMVP inter-frame, Merge inter-frame, Merge skip, AMVP affine inter-frame, affine Merge inter-frame, affine Merge skip, sbTMVP, IBC AMVP, IBC Merge, IBC Merge skip, angular intra-frame prediction, MIP intra-frame prediction, CCLM or its variants, MM-CCLM or its variants, CCCM or its variants, MM-CCCM or its variants, CCPMerge mode, GPM, CIIP, TM-Merge, SGPM, intra-frame template, etc.
[1090] ii. For example, the second mode can be AMVP inter-frame, Merge inter-frame, Merge skip, AMVP affine inter-frame, affine Merge inter-frame, affine Merge skip, sbTMVP, IBC AMVP, IBC Merge, IBC Merge skip, angular intra-frame prediction, MIP intra-frame prediction, CCLM or its variants, MM-CCLM or its variants, CCCM or its variants, MM-CCCM or its variants, CCPMerge mode, GPM, CIIP, TM-Merge, SGPM, intra-frame template, etc.
[1091] 10) In one example, at least one SE can be transmitted via signal to indicate whether different components can be encoded or decoded using different encoding / decoding modes.
[1092] a. In one example, if SE is not transmitted via signaling, it can be implicitly interpreted as "false".
[1093] b. In one example, the SE can only be transmitted via signal if the luminance component is encoded or decoded using a specific mode or a specific mode.
[1094] c. In one example, the SE can only be transmitted via signal if the width and / or height of the block meets a specific condition or condition.
[1095] 11) In one example, at least one SE can be signaled to indicate a mode for chroma encoding / decoding that is different from the mode signaled for luminance.
[1096] a. In one example, the SE can only be transmitted via signal if the luminance and chrominance components can be encoded and decoded using different encoding and decoding modes.
[1097] b. In one example, a set of candidate patterns can be determined, and the SE can indicate one of them.
[1098] i. For example, a set of candidate patterns can be determined in a predefined manner.
[1099] ii. For example, a set of candidate patterns can be determined on the spot for the current block.
[1100] 1. This group can depend on the brightness mode.
[1101] 2. This group can depend on the width and / or height of the block.
[1102] 3. This group may depend on neighboring information.
[1103] c. Alternatively, the mode used for chroma encoding and decoding can be determined at the decoder without signal transmission.
[1104] i. In one example, different patterns can be applied to the template of the current block to calculate the cost, and the pattern with the lowest cost can be selected for chroma.
[1105] 12) The model candidate checking rules for inter-frame CCP model inheritance / Merge mode and intra-frame CCP model inheritance / Merge mode can be the same.
[1106] a. For example, the same CCP model checklist can be defined for CCP model inheritance / Merge modes, regardless of the prediction mode of the current block (e.g., whether the current block is inter-frame or intra-frame encoded).
[1107] b. For example, the CCP model checklist may contain candidates in a predefined order from at least one of the following locations: i. Spatial adjacency candidates.
[1108] ii. Temporal candidates at predefined locations.
[1109] iii. Temporal candidates with shifted MVs (e.g., current MV, neighboring MVs, generated MVs, etc.).
[1110] iv. Non-adjacent candidates at predefined locations.
[1111] v. Non-adjacent candidates with shifted BVs (e.g., current BV, neighboring BVs, generated BVs, etc.).
[1112] vi. A history-based lookup table.
[1113] vii. Default candidate.
[1114] viii. Model candidates calculated on the fly (e.g., calculated from a reference region).
[1115] c. For example, a model candidate can be inserted into the list if at least one of the following conditions is met: i. The candidate is intra-frame CCP encoded / decoded.
[1116] ii. The candidates are encoded and decoded via inter-frame CCP.
[1117] iii. The candidate is different from the candidate already inserted in the list.
[1118] d. Alternatively, different model candidate checking rules can be applied to inter-frame CCP model inheritance / Merge mode and intra-frame CCP model inheritance / Merge mode.
[1119] i. For example, different types of models can be examined.
[1120] ii. For example, inter-frame CCP model inheritance / Merge mode may not check temporal candidates.
[1121] iii. For example, inter-frame CCP model inheritance / Merge mode may not check the default candidate.
[1122] iv. For example, inter-frame CCP model inheritance / Merge mode can examine candidates for on-the-fly computation.
[1123] 13) The final chromaticity prediction of a block can be a mixture of chromaticity predictions based on a first CCP model and chromaticity predictions based on a second non-CCP model.
[1124] a. For example, the final chroma prediction for an intra-block can be fused from intraCCP chroma prediction and angle / plane / DC / LM-based chroma prediction.
[1125] i. For example, the intraCCP model can be inherited from previous blocks.
[1126] ii. For example, the intraCCP model can be calculated from the reference region.
[1127] b. For example, the final chroma prediction for an inter-frame block can be fused from interCCP chroma prediction and chroma prediction based on inter-frame motion compensation.
[1128] i. For example, the interCCP model can be inherited from previous blocks.
[1129] ii. For example, the interCCP model can be computed from a reference region.
[1130] 14) For CCP model computation and / or application, access to / use of samples in the reference region (training region) and access to / use of samples in the current block can be decoupled.
[1131] a. For example, for chroma samples in the reference region (training region), the CCP model input may not include samples within the current block.
[1132] b. For example, for chroma samples in the reference region (training region), the CCP model input can include only the samples within the reference region (training region).
[1133] c. For example, for chroma samples within the current block, the CCP model input may not include samples from the reference region (training region).
[1134] d. For example, for chroma samples within the current block, the CCP model input may only include samples within the current block (e.g., luminance samples within the current block).
[1135] 15) The CCP model of the current chroma block can be calculated based on reference samples only (e.g., without accessing internal samples from the current luma block).
[1136] a. For example, model calculations for the CCP / CCCM model for encoding and decoding the current inter-frame (or IBC) chroma block may not use samples from the current luma block.
[1137] i. For example, it can occur during the inter-frame CCP Merge mode process.
[1138] ii. For example, it can occur during a regular inter-frame CCP mode process.
[1139] b. For example, model calculations for the CCP / CCCM model encoding / decoding the current intra-frame chroma block may not use samples from the current luma block.
[1140] i. For example, it can occur during an intra-frame CCP Merge mode process.
[1141] ii. For example, it can occur during a regular intra-frame CCP mode process.
[1142] c. For example, if the CCP model computation for the current chroma block requires access to samples within the current luma block (e.g., in the case where the training region (reference region) contains the upper sample closest / adjacent to the current chroma block), the south term of the CCCM model input for such reference samples (e.g.) Figure 12 As shown, it may require samples within the current chroma block; or, for example, if the training region (reference region) contains the leftmost sample closest to / adjacent to the current chroma block, the CCCM model input for such reference samples (e.g., Figure 12 (As shown) If samples within the current luminance block are required, then one of the following methods can be applied: i. For example, you can use replacement sample values (instead of the sample values of the current luma block).
[1143] 1. Default values can be used. a. For example, default values can be predefined. b. For example, the default value could be based on the bit depth of the samples in the luminance or chrominance array. 2. Sample values of neighboring brightness samples (in the reference area) can be used.
[1144] a. For example, in the case where the training region (reference region) contains the upper sample point that is closest / adjacent to the current chroma block, the southern term of the CCCM model input for such reference sample points (e.g.) Figure 12 (As shown) can be set to the value of the brightness sample above.
[1145] b. For example, if the training region (reference region) contains the leftmost / adjacent sample point to the current chroma block, the Eastern term input to the CCCM model for such reference sample points (e.g.) Figure 12 (As shown) can be set to the value of the left-hand brightness sample point.
[1146] ii. For example, if a specific sample in the training region (reference region) needs to access a sample within the current luminance block, such a training sample can be excluded / discarded from the collection of training samples computed for such a CCP model.
[1147] 16) The cost calculation process for the decoder derivation used to evaluate the CCP model can be calculated based solely on reference samples (e.g., without accessing internal samples from the current block).
[1148] a. For example, it can occur during the candidate ranking / re-ranking process in the CCP model.
[1149] b. For example, it can occur when computing a CCP model on a template (e.g., the template includes at least one sample point that is adjacent to and outside the current block).
[1150] c. For example, the cost calculation for the decoder derivation used to evaluate the CCP model can be calculated based on inter-frame prediction blocks in a reference image.
[1151] d. For example, the cost calculation for the decoder derivation used to evaluate the CCP model can be calculated based on a BV-guided reference block in the current image.
[1152] i. In addition, the distance between the current block and the reference block guided by BV can be greater than N samples, such as N=1.
[1153] e. For example, if the CCP model calculation for template samples requires access to samples inside the current luma block (e.g., in the case where the template sample is adjacent to the current block above, the southern term of the CCCM model input for such template samples (e.g.) Figure 12 As shown, it may require samples within the current luminance block; for example, when the template sample is adjacent to the left side of the current block, the Eastern term (e.g., ...) input to the CCCM model for such template samples may be required. Figure 12 (As shown) If samples within the current luminance block are required, then one of the following methods can be applied: i. For example, you can use replacement sample values (instead of the sample values of the current luma block).
[1154] 1. Default values can be used.
[1155] a. For example, default values can be predefined.
[1156] b. For example, the default value can be based on the bit depth of the samples in the luminance or chrominance array.
[1157] 2. Sample values of neighboring brightness samples (in the reference area) can be used.
[1158] a. For example, when the template sample is adjacent to the current block, the southern term (e.g., ...) input to the CCCM model for such template sample... Figure 12 (As shown) can be set to the value of the brightness sample above.
[1159] b. When the template sample point is adjacent to the left side of the current block, the Eastern term (e.g., ...) is input to the CCCM model for such template sample points. Figure 12 (As shown) can be set to the value of the left-hand brightness sample point.
[1160] ii. For example, if the CCP model computation for template samples requires access to samples within the current luma block, such training samples can be excluded / discarded from the training sample collection for such CCP model computation.
[1161] iii. For example, if the CCP model calculation for template samples requires accessing at least one sample within the current luma block (e.g. Figure 32 As shown, assuming that the first row and first column of the current luminance block, which are painted yellow, will be accessed, and (x, y) represents the sample point located in the x-th row and y-th column of the current luminance block, then these sample points to be accessed in the current luminance block can be filled with other values.
[1162] 1. For example, other values can be based on neighboring samples (e.g., in the reference region, training region, decoding region, etc.).
[1163] 2. For example, the first row of samples to be accessed within the current luminance block (e.g., in...). Figure 32 The value (0, 0)...(0, 7) can be filled with neighboring samples on the row above the current luminance block (e.g., as shown in the image). Figure 32 The neighboring sample points are represented as (-1, 0)...(-1, 7).
[1164] a. For example, additionally, the upper left sample point (e.g., in...) Figure 32 The value represented as (0, 0) can be filled in another way (e.g., not from sample values from (-1, 0)).
[1165] 3. For example, the first column of the samples to be accessed within the current luminance block (e.g., in...). Figure 32 The value (0, 0)...(7, 0) can be filled with neighboring samples on the left column outside the current luminance block (e.g., as shown in the image). Figure 32 The neighboring sample points are represented as (0, -1)...(7, -1)).
[1166] a. For example, additionally, the upper left sample point (e.g., in...) Figure 32The value represented as (0, 0) can be filled in another way (e.g., not from sample values from (0, -1)).
[1167] 4. For example, the upper left sample of the sample to be accessed within the current luminance block (e.g., in...). Figure 32 The value (0, 0) can be filled with the upper-left neighboring sample outside the current brightness block (e.g., in the...). Figure 32 The expression is represented as (-1, -1)).
[1168] a. Alternatively, it can be filled with the left-side neighboring samples outside the current brightness block (e.g., in...). Figure 32 In this case, it is represented as (0, -1).
[1169] b. Alternatively, it may be filled with the upper neighboring samples outside the current brightness block (e.g., in...). Figure 32 The value is represented as (-1, 0).
[1170] c. Alternative sites, which can be filled with the values of sample points at predefined locations.
[1171] d. Alternatively, it may be filled with predefined values (e.g., 1 << (bit depth - 1), where the bit depth is the bit depth of the sample points of the luminance or chrominance array).
[1172] 5. For example, the first row of samples to be accessed within the current luminance block can be filled first, followed by the filling process of the first column of samples to be accessed within the current luminance block.
[1173] a. Alternatively, the first column of the samples to be accessed within the current luminance block can be filled first, followed by the filling process of the first row of the samples to be accessed within the current luminance block.
[1174] 6. For example, the fill value of the current brightness sample can be used as an intermediary for CCP model calculation or CCP model cost calculation for the template.
[1175] a. For example, the fill value of the current brightness sample may not be used in subsequent processing of the current block (e.g., CCP model application for the current block, prediction / reconstruction derivation for the current block, etc.).
[1176] Figure 32 The diagram illustrates an example where CCP model calculations for template samples (dark gray) require access to sample values in the current luminance block (light gray, note that each grid represents a sample), where (x, y) in the figure represents the sample located in the x-th row and y-th column of the current luminance block.
[1177] 17) The CCP model application for the current block / region can be applied based only on samples within the current block / region (e.g., without needing to access samples outside the current block / region).
[1178] a. For example, if the CCP model calculation for the current block / region requires access to samples outside the current block / region (e.g., for chromaticity samples located at the top boundary of the current block / region, the northern term of the CCCM model input for such samples (e.g.) Figure 12 As shown, it may require samples outside the current block / region; for example, for chromaticity samples located at the left boundary of the current block / region, the Western term of the CCCM model input for such samples (e.g.) Figure 12 (As shown) If samples outside the current block / region are required, then one of the following methods can be applied: i. For example, you can use replacement sample values (instead of sample values outside the current block / region).
[1179] 1. Default values can be used. a. For example, default values can be predefined. b. For example, the default value could be based on the bit depth of the samples in the luminance or chrominance array. 2. The sample values of the current luminance block can be used.
[1180] a. For example, for chromaticity samples located at the top boundary of the current block / region, the northern term of the CCCM model input for such samples (e.g.) Figure 12 (As shown) can be set to the value equal to the current brightness sample.
[1181] b. For example, for a chromaticity sample located at the left boundary of the current block / region, the Western term of the CCCM model input for such a sample (e.g.) Figure 12 (As shown) can be set to the value equal to the current brightness sample.
[1182] 18) Inter-frame CCP mode may be permitted to be applied to video blocks when at least one of the following conditions is met: a. The current stripe type (e.g., whether it is a B / P stripe, etc.). b. Two trees or one tree (e.g., whether it is a single tree, etc.) c. Merge type (e.g., whether to use Merge mode, regular Merge mode, or sub-block Merge mode for encoding / decoding) d. AMVP type (e.g., whether it is not encoded or decoded using AMVP, etc.) e. Reference picture information (e.g., whether it is traditional bidirectional prediction coding / decoding where two reference pictures come from different directions, whether it is unidirectional prediction coding / decoding, whether it is a low-latency picture where all reference pictures are before the current picture in display order, etc.) i. Additionally, whether it is a low-latency picture / strip and Merge coding / decoding f. cbf information (e.g., whether it has non-zero luminance coefficients, etc.) g. Residual / coefficient information (e.g., the absolute value of luminance residual / coefficient, the number of non-zero luminance coefficients, etc.) h. TB and CB sizes of the current video block (e.g., whether TB is equal to CB, etc.) i. Non-SBT coding / decoding j. Prediction mode (e.g., whether it is an inter-frame mode and / or an IBC mode, etc.) k. Picture resolution (e.g., whether it is non-4K) i. For example, whether the picture height is no greater than 1080 ii. For example, if the height of the current picture is greater than T (e.g., T = 1080, etc.), the inter-frame CCP mode may not be allowed iii. For example, if the width of the current picture is greater than R (e.g., R = 4096, etc.), the inter-frame CCP mode may not be allowed l. Block dimensions, such as block height H and / or block width W i. a0 W < b0 H, or, a0 W <= b0 H ii. a1 W > b1 H, or, a1 W >= b1 H iii. a2 H < b2 W, or, a2 H <= b2 W iv. a3 H > b3 W, or, a3 H >= b3 W v. Min (W, H) > T0, or, Min (W, H) >= T0 vi. Max (W, H) < T1, or, Max (W, H) <= T1 vii. W H < T2, or W H <= T2 viii. For example, a0, a1, a2, a3, b0, b1, b2, b3 can be predefined integers.
[1183] ix. For example, T0, T1, T2 can be predefined values.
[1184] x. For example, for W For a transform block (TB) where H > 1024 or 516, the inter-frame CCP mode may not be allowed xi. For example, for W For a transform block (TB) where H <= 8 or 16 or 32, the inter-frame CCP mode may not be allowed xii. For example, for 16 For a transform block (TB) where W <= H, the inter-frame CCP mode may not be allowed xiii. For example, for 16 For a transform block (TB) where H <= W, the inter-frame CCP mode may not be allowed m. The temporal layer of the video unit is less than a threshold (such as less than 4 or 5) 19) Whether to insert the instantaneously computed CCP model into the CCP candidate list can depend on at least one of the following conditions: a. The temporal layer of the current video unit i. For example, if the temporal layer of the current video unit is less than a threshold (such as less than 4 or 5), it may not be inserted.
[1185] b. Reference picture information (e.g., whether it is conventional bidirectional prediction coding and decoding where two reference pictures are from different directions, whether it is unidirectional prediction coding and decoding, whether it is a low-latency picture where all reference pictures are before the current picture in display order, etc.) i. For example, if the current slice is not a low-latency picture / slice, it may not be inserted c. cbf information (e.g., whether it has non-zero luminance coefficients, etc.) d. The slice type of the current slice (e.g., whether it is a B / P slice, etc.) i. For example, if the current slice is an I slice, it may not be inserted e. Dual-tree or single-tree (e.g., whether it is a single-tree, etc.) i. For example, if the current slice is a dual-tree, it may not be inserted f. Merge type (e.g., whether it is encoded and decoded using the Merge mode, or the conventional Merge mode, or the sub-block Merge mode, etc.) g. AMVP type (e.g., whether it is not encoded / decoded by AMVP, etc.) h. Residual / coefficient information (e.g., absolute value of luminance residual / coefficient, number of non-zero luminance coefficients, etc.) i. TB and CB sizes of the current video block (e.g., whether TB is equal to CB, etc.) j. Non-SBT encoding / decoding k. Prediction mode (e.g., whether it is an inter-frame mode and / or an IBC mode, etc.) l. Picture resolution (e.g., whether it is non-4K) i. For example, whether the picture height is not greater than 1080 ii. For example, if the current picture height is greater than T (e.g., T = 1080, etc.), the inter-frame CCP mode may not be allowed iii. For example, if the current picture width is greater than R (e.g., R = 4096, etc.), the inter-frame CCP mode may not be allowed m. Block dimensions, such as block height H and / or block width W i. a0 W < b0 H, or, a0 W <= b0 H ii. a1 W > b1 H, or, a1 W >= b1 H iii. a2 H < b2 W, or, a2 H <= b2 W iv. a3 H > b3 W, or, a3 H >= b3 W v. Min (W, H) > T0, or, Min (W, H) >= T0 vi. Max (W, H) < T1, or, Max (W, H) <= T1 vii. W H < T2, or W H <= T2 viii. For example, a0, a1, a2, a3, b0, b1, b2, b3 can be predefined integers.
[1186] ix. For example, T0, T1, and T2 can be predefined values.
[1187] x. For example, for W Transform blocks (TB) with H > 1024 or 516 may not be allowed in inter-frame CCP mode. xi. For example, for W For transform blocks (TB) with H <= 8, 16, or 32, inter-frame CCP mode may not be allowed. xii. For example, for 16 Transform blocks (TBs) where W <= H may not be allowed in inter-frame CCP mode. xiii. For example, for 16 Transform blocks (TBs) where H <= W, inter-frame CCP mode may not be allowed. 20) On-the-fly CCP model candidates can be inserted before spatially adjacent candidates, and / or temporal candidates, and / or spatially non-adjacent candidates, and / or history-based candidates, and / or default candidates.
[1188] a. In addition, redundancy checks (e.g., full or partial deduplication) can be applied, and an immediate model candidate to be inserted can only be finally inserted into the list if it is not the same as / similar to a predefined (or all) existing candidate in the list.
[1189] 21) Instantaneously computed CCP model candidates can be inserted after spatially adjacent candidates, and / or temporal candidates, and / or spatially non-adjacent candidates, and / or history-based candidates.
[1190] a. In addition, redundancy checks (e.g., full or partial deduplication) can be applied, and an immediate model candidate to be inserted can only be finally inserted into the list if it is not the same as / similar to a predefined (or all) existing candidate in the list.
[1191] 22) At least one default CCP candidate can be inserted into the CCP candidate list.
[1192] a. For example, a CCP candidate list can be used in inter-frame CCP mode.
[1193] b. For example, a CCP candidate list can be used in intra-frame CCP mode.
[1194] c. For example, default CCP candidates can be derived based on CCP models computed on the fly.
[1195] i. For example, the default CCP candidate can be derived based on the CCLM parameters of the on-the-fly computed CCLM model.
[1196] 1. For example, the scaling factor of the default CCP model can be derived based on the scaling factor of the CCLM model calculated on the fly.
[1197] a. In addition, the scaling factor of the default CCP model can be derived based on a predefined fixed value.
[1198] b. Alternatively, the scaling factor of the default CCP model can be derived based on both the shift and fixed values of the on-the-fly calculated CCLM model.
[1199] 2. For example, the shift value of the default CCP model can be derived based on the shift value of the CCLM model calculated on the fly.
[1200] a. In addition, the shift value of the default CCP model can be derived based on predefined fixed values.
[1201] b. Alternatively, whether the shift value is equal to the predefined fixed value may depend on the relationship between the shift value and the fixed value of the CCLM model calculated on the fly (e.g., less than or greater than, etc.).
[1202] ii. For example, it can be derived based on the first available, instant-computable CCLM model.
[1203] d. For example, if the CCP candidate list is not full, at least one default CCP candidate based on an on-the-spot computational CCP model can be inserted into the list.
[1204] Derived CCP / non-CCP candidate correlation 23) The above items can be applied to non-CCP model / filter calculation / prediction / application.
[1205] a. For example, non-CCP filter / model coefficients can be derived based on the same component prediction (e.g., luminance-to-luminance prediction, etc.).
[1206] b. For example, non-CCP models / filters can be based on EIF filters, LIC models, IBC filters, intraTMP filters, reference sample filters, template filters, DIMD filters, TIMD filters, etc.
[1207] 24) Filters / models can be computed from the decoded information and used for encoding and decoding the current block or future blocks.
[1208] a. For example, a filter / model can contain a set of coefficients.
[1209] b. For example, the filter / model can be an inter-frame / intra-frame CCP model.
[1210] i. For example, filter / model coefficients can be derived based on cross-component predictions.
[1211] c. For example, the filter / model can be a non-CCP model (such as EIF filter, LIC model, IBC filter, intraTMP filter, template filter, DIMD filter, TIMD filter, etc.).
[1212] i. For example, filter / model coefficients can be derived based on the same component prediction.
[1213] d. For example, a filter / model (e.g., a CCP model) can be derived based on minimizing the difference between the luminance sample values and the chrominance sample values of the reference / training region / block.
[1214] i. For example, the model can be applied to the luminance reconstruction samples of the reference / training region / block and the resulting model estimate samples are obtained. Then, minimization processing is performed based on the obtained model estimate samples and the chromaticity reconstruction sample values of the reference / training region / block.
[1215] ii. For example, in addition, the reference / training region / block can be adjacent / non-adjacent / temporally / co-located with the current block.
[1216] iii. For example, the reference / training region / block can also be derived based on the block vector.
[1217] iv. For example, in addition, the reference / training region / block can be derived based on motion vectors.
[1218] e. For example, filters / models (e.g., non-CCP models) can be derived based on minimizing the difference between predicted and reconstructed samples of the current block.
[1219] i. For example, the model can be applied to the predicted samples of the current block and the resulting model estimated samples are obtained. Then, minimization processing is performed based on the obtained model estimated samples and the reconstructed sample values of the current block.
[1220] ii. For example, the predicted samples of the current block can be derived based on the block vector.
[1221] iii. For example, the predicted samples of the current block can be derived based on motion vectors.
[1222] f. For example, filters / models (e.g., non-CCP models) can be derived based on minimizing the difference between non-adjacent reconstructed samples and the current reconstructed sample.
[1223] i. For example, non-adjacent reconstructed samples can be located by block vectors.
[1224] ii. For example, non-adjacent reconstructed samples can be located by predefined positions (e.g., depending on the width / height of the current block).
[1225] iii. For example, the model can be applied to non-adjacent reconstructed samples and the resulting model-estimated samples are obtained. Then, a minimization process is performed based on the obtained model-estimated samples and the reconstructed sample values of the current block.
[1226] g. For example, filters / models (e.g., non-CCP models) can be derived based on minimizing the difference between reference image samples and the current reconstructed samples.
[1227] i. For example, reference image samples can be located using motion vectors.
[1228] ii. For example, reference image samples can be located by predefined positions (e.g., co-located with the current block).
[1229] iii. For example, the model can be applied to reference image samples and the resulting model-estimated samples are obtained. Then, a minimization process is performed based on the obtained model-estimated samples and the reconstructed sample values of the current block.
[1230] h. For example, filters / models (e.g., non-CCP models, CCP models, etc.) can be derived based on minimizing the difference between the reference template and the current template.
[1231] i. For example, a reference template can be adjacent to or not adjacent to a reference block.
[1232] ii. For example, the current template can be adjacent to or not adjacent to the current block.
[1233] iii. For example, the model can be applied to reference template samples and the resulting model estimate samples can be obtained. Then, minimization processing is performed based on the obtained model estimate samples and the reconstructed current template sample values.
[1234] i. For example, the above filter / model can be computed based on a training set constructed from a set of samples from {the current block, the reference block guided by BV / MV, the template, the non-adjacent block, the adjacent block, the temporal co-occurrence block, the temporal block adjacent / non-adjacent to the co-occurrence block, etc.}.
[1235] i. For example, more than one type of sample point can be used.
[1236] ii. For example, for the computation of a particular model, the training samples used for the calculation of model coefficients can be derived from previously encoded blocks encoded by such a particular model.
[1237] 1. For example, a specific model can be based on CCLM, intra / inter / BVG CCCM, CCCM with MDF, GL-CCCM, GLM, etc.
[1238] 2. For example, a specific model can be based on EIF filters, LIC models, IBC filters, intraTMP filters, template filters, DIMD filters, TIMD filters, etc.
[1239] j. For example, the derived filter / model can be a linear model, a nonlinear model, or a convolution-based model, etc.
[1240] i. For example, the derived filter / model can be similar to CCLM, similar to intraCCCM, similar to BVG-CCCM, similar to interCCCM, similar to LIC model, similar to EIF, etc.
[1241] k. For example, the calculated model can be used for the luma and / or chroma encoding and decoding of the current block.
[1242] For example, the computed model can be stored in a cache for encoding and decoding of future blocks.
[1243] 25) Inter-frame CCP / intra-frame CCP modes can be applied based on a CCP candidate list derived by the decoder.
[1244] a. For example, the CCP list may contain at least one CCP model candidate that is derived / computed on the fly.
[1245] i. For example, the CCP model can be based on CCLM, intra / inter / BVG CCCM, CCCM with MDF, GL-CCCM, GLM, LBCCP, etc.
[1246] ii. For example, the CCP model can be computed on the fly from training samples that are adjacent to the current block.
[1247] iii. For example, the CCP model can be computed on the fly from training samples in the reference block.
[1248] b. For example, all candidates in the list are derived / computed on the fly.
[1249] i. For example, alternatively, the CCP list may contain more than one CCP model candidate derived / computed on the fly.
[1250] c. For example, LBCCP-based model candidates can be inserted into the list.
[1251] i. For example, whether a model candidate has been encoded or decoded using LBCCP can be inherited from previous blocks.
[1252] ii. For example, whether a model candidate is LBCCP encoded or decoded can be derived / computed / determined based on the template (e.g., by comparing the template cost with and without a low-pass filter).
[1253] d. For example, alternatively, both candidates for instantaneous derivation / computation and candidates for inheritance can be included in the list.
[1254] i. For example, additionally, candidates derived / computed on the fly can be placed before inherited candidates.
[1255] e. For example, all CCP models in the list can be sorted / reordered.
[1256] i. For example, sorting can be processed based on template cost.
[1257] ii. For example, candidate indices can be signaled to indicate which candidate model is ultimately selected for encoding and decoding the current block.
[1258] 1. Alternatively, the candidate model with the lowest cost after sorting can be used by default for encoding and decoding the current block (e.g., without indexing via signaling).
[1259] f. For example, alternatively, which CCP model in the list is selected can be determined on the encoder side and transmitted via signal in the bitstream.
[1260] g. For example, such a mode can be transmitted via signaling as a separate mode independent of the interCCCM / interCCP Merge / intraCCPMerge / intraCCLM / intraCCCM modes.
[1261] i. Alternatively, such a mode can be transmitted via signal as a sub-mode of interCCCM / interCCP Merge / intraCCP Merge / intraCCLM / intraCCCM mode.
[1262] h. For example, such a pattern can be used for chroma inter-frame encoding and decoding.
[1263] i. For example, such a pattern can be used for chroma intra-frame encoding and decoding.
[1264] General aspects : 1) The disclosed method can be used for single trees.
[1265] 2) The disclosed method can be used for two trees.
[1266] 3) The disclosed method can be used for chroma encoding and decoding.
[1267] 4) The disclosed method can be used in inter-frame (such as B or P) stripes.
[1268] 5) The disclosed method can be used in intra-frame (such as I) stripes.
[1269] 6) Whether and / or how the methods disclosed above can be applied to be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[1270] 7) Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.
[1271] 8) Whether and / or how to apply the methods disclosed above may depend on the encoded / decoded information, such as block size, color format, single / dual tree segmentation, color components, and stripe / picture type.
[1272] In some embodiments, items 1-22 above may be applicable to at least one of the following: non-CCP models, non-CCP filters, non-CCP computation, non-CCP prediction, or non-CCP applications. For example, the term "CCP" mentioned in items 1-22 may be replaced with the term "non-CCP". That is, the example embodiments described with reference to items 1-22 may be applicable to non-CCP scenarios. In some embodiments, the coefficients of at least one of the non-CCP filters or non-CCP models may be derived based on component prediction. For example, component prediction may be lumen-to-lumen prediction. In some other embodiments, the non-CCP filters and / or non-CCP models may be based on at least one of the following: extended information filter (EIF), local illumination compensation (LIC) model, intra-block copy (IBC) filter, intra-template matching prediction (intra TMP) filter, reference sample filter, template filter, decoder-side intra-mode derivation (DIMD) filter, or template-based intra-mode derivation (TIMD) filter.
[1273] Figure 33 A flowchart of a method 3300 for video processing according to an embodiment of the present disclosure is shown. Method 3300 is implemented during the conversion between video units of a video and a bitstream of a video.
[1274] At box 3310, for the conversion between the current block of the video and the bitstream of the video, at least one filter or model is derived based on the decoding information.
[1275] At box 3320, at least one of the filters or models is applied to the encoding and decoding of at least one of the current block or future blocks.
[1276] At block 3330, the transformation is performed based on at least one of a filter or model. In some embodiments, the transformation may include encoding the current block into a bitstream. Alternatively, the transformation may include decoding the current block from the bitstream.
[1277] Method 3300 enables the derivation and application of filters or models. Compared to traditional solutions, encoding / decoding efficiency and performance can be significantly improved.
[1278] In some embodiments, at least one of the filters or models may include a set of coefficients. In some other embodiments, at least one of the filters or models may include an inter-frame cross-component prediction (CCP) model or an intra-frame CCP model. For example, the coefficients of at least one of the filters or models may be derived based on cross-component prediction. Alternatively, at least one of the filters or models may include a non-CCP model. For example, a non-CCP model may include at least one of the following: an extended information filter (EIF), a local illumination compensation (LIC) model, an intra-frame block copy (IBC) filter, an intra-frame template matching prediction (intraTMP) filter, a template filter, a decoder-side intra-frame pattern derivation (DIMD) filter, or a template-based intra-frame pattern derivation (TIMD) filter. In some examples, the coefficients of at least one of the filters or models may be derived based on cross-component prediction.
[1279] In some embodiments, at least one of the filters or models can be derived based on a process of minimizing the difference between the value of a luminance sample and the value of a chrominance sample in one of the reference, training, or block regions. For example, the model may include a CCP model. In some embodiments, the model can be applied to luminance reconstructed samples in one of the reference, training, or block regions to obtain the resulting model-estimated samples. Additionally, the minimization process can be based on the values of the resulting model-estimated samples and the chrominance reconstructed samples in one of the reference, training, or block regions. In some embodiments, one of the reference, training, or block regions may be adjacent to the current block. In some other embodiments, one of the reference, training, or block regions may not be adjacent to the current block. Alternatively, one of the reference, training, or block regions may be temporally co-located with the current block. In some embodiments, one of the reference, training, or block regions can be derived based on a block vector. Alternatively, one of the reference, training, or block regions can be derived based on a motion vector.
[1280] In some embodiments, at least one of the filters or models can be derived based on a process of minimizing the difference between the predicted samples of the current block and the reconstructed samples of the current block. For example, the model may include a non-CCP model. In some embodiments, the model can be applied to the predicted samples of the current block to obtain the resulting model-estimated samples. Alternatively, the minimization process can be based on the values of the resulting model-estimated samples and the reconstructed samples of the current block. In some embodiments, the predicted samples of the current block can be derived based on the block vector. Alternatively, the predicted samples of the current block can be derived based on the motion vector.
[1281] In some embodiments, at least one of the filters or models can be derived based on a process of minimizing the difference between non-adjacent reconstructed samples and the current reconstructed sample. For example, the model may include a non-CCP model. In some embodiments, non-adjacent reconstructed samples may be located by block vectors. In some other embodiments, non-adjacent reconstructed samples may be located by predetermined positions. For example, the predetermined positions may depend on the width and / or height of the current block. In some embodiments, the model may be applied to non-adjacent reconstructed samples to obtain the resulting model-estimated samples. Additionally, the minimization process may be based on the values of the obtained model-estimated samples and the reconstructed samples of the current block.
[1282] In some embodiments, at least one of the filters or models can be derived based on a process of minimizing the difference between reference image samples and current reconstructed samples. For example, the model may include a non-CCP model. In some embodiments, the reference image samples may be located by block vectors. In some other embodiments, the reference image samples may be located at predetermined locations. For example, the predetermined location may be co-located with the current block. In some embodiments, the model may be applied to the reference image samples to obtain the resulting model-estimated samples. Additionally, the minimization process may be based on the values of the obtained model-estimated samples and the reconstructed samples of the current block.
[1283] In some embodiments, at least one of the filters or models can be derived based on a process of minimizing the difference between the reference template and the current template. For example, the model may include a non-CCP model or a CCP model. In some embodiments, the reference template may be adjacent to a reference block. In some other embodiments, the reference template may not be adjacent to the reference block. In some embodiments, the current template may be adjacent to the current block. Alternatively, the current template may not be adjacent to the current block. In some embodiments, the model may be applied to reference template samples to obtain the resulting model estimate samples. Additionally, the minimization process may be based on the values of the obtained model estimate samples and the reconstructed current template samples.
[1284] In some embodiments, at least one of the filters or models may be computed based on a training set. For example, the training set may be determined based on a set of samples from at least one of the following: the current block, a block vector (BV)-guided reference block, a motion vector (MV)-guided reference block, a template, a non-adjacent block, an adjacent block, a temporal co-occurrence block, a temporal block adjacent to a co-occurrence block, or a temporal block not adjacent to a co-occurrence block. In some embodiments, multiple samples may be used to determine the training set. In some other embodiments, training samples used to compute the coefficients of the model may be derived from previously encoded / decoded blocks. In this case, the previously encoded / decoded blocks may be encoded / decoded by the model. For example, the model may be based on at least one of the following: Cross-Component Linear Model (CCLM), Convolutional Cross-Component Model (CCCM), Inter-Frame CCCM, Block Vector Guided CCCM (BVG-CCCM), CCCM with Multiple Downsampling Filters (MDF), Gradient Linear-Convolutional Cross-Component Model (GL-CCCM), or Gradient Linear Model (GLM). Alternatively, the model may be based on at least one of the following: EIF, Local Illumination Compensation (LIC) model, Intra-Block Copy (IBC) filter, Intra-Template Matching Prediction (intraTMP) filter, Template Filter, Decoder-Side Intra-Mode Derivation (DIMD) Filter, or Template-Based Intra-Mode Derivation (TIMD) Filter.
[1285] In some embodiments, at least one of the derived filters or derived models may include at least one of linear models, nonlinear models, or convolution-based models. For example, at least one of the derived filters or derived models may include at least one of models similar to CCLM, intra-frame CCCM, BVG-CCCM, inter-frame CCCM, LIC, or EIF. In some other embodiments, the derived model may be used for at least one of luma encoding / decoding or chroma encoding / decoding of the current block. Alternatively, the derived model may be stored in a cache for encoding / decoding of future blocks.
[1286] In some embodiments, at least one of the derived filters or derived models may be used in at least one of single-tree or dual-tree architectures. In some other embodiments, at least one of the derived filters or derived models may be used for chroma encoding / decoding. In some embodiments, at least one of the derived filters or derived models may be used in inter-frame stripes. For example, an inter-frame stripe may be a B-strip or a P-strip. In some other embodiments, at least one of the derived filters or derived models may be used in intra-frame stripes. For example, an intra-frame stripe may be an I-strip.
[1287] In some embodiments, an indication of whether and / or how to derive at least one of the filters or models based on decoded information may be indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level. For example, an indication of whether and / or how to derive at least one of the filters or models based on decoded information may be indicated at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[1288] In some embodiments, an indication of whether at least one of the filters or models is derived based on the decoded information and / or how at least one of the filters or models is derived based on the decoded information may be included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[1289] In some embodiments, method 3300 may further include: determining, based on the encoded and decoded information of the video unit of the video, whether to derive at least one of the filters or models based on the decoded information and / or how to derive at least one of the filters or models based on the decoded information, wherein the encoded and decoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type or picture type.
[1290] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by a video process apparatus. The method includes: deriving at least one of filters or models based on decoding information; applying at least one of the filters or models to at least one of the current or future blocks of the video for encoding / decoding; and generating a bitstream based on at least one of the filters or models.
[1291] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: deriving at least one of filters or models based on decoding information; applying at least one of the filters or models to encoding / decoding at least one of the current or future blocks of the video; generating a bitstream based on at least one of the filters or models; and storing the bitstream in a non-transitory computer-readable recording medium.
[1292] Figure 34 A flowchart of a method 3400 for a video process according to an embodiment of the present disclosure is shown. Method 3400 is implemented during the conversion between video units of a video and a bitstream of a video.
[1293] At box 3410, for the conversion between the current block of the video and the video bitstream, at least one inter-frame CCP mode or intra-frame CCP mode is applied based on a cross-component prediction (CCP) candidate list. In this case, the CCP candidate list is derived based on the decoder.
[1294] At block 3420, the conversion is performed based on at least one of inter-frame CCP mode or intra-frame CCP mode. In some embodiments, the conversion may include encoding the current block into a bitstream. Alternatively, the conversion may include decoding the current block from the bitstream.
[1295] Method 3400 enables the application of inter-frame CCP mode and / or intra-frame CCP mode based on a CCP candidate list. Compared to traditional solutions, encoding / decoding efficiency and performance can be significantly improved.
[1296] In some embodiments, the CCP candidate list may include at least one CCP model candidate derived or computed in real time. In some embodiments, the CCP model may be based on at least one of the following: Cross-Component Linear Model (CCLM), Convolutional Cross-Component Model (CCCM), Inter-Frame CCCM, Block Vector Guided CCCM (BVG-CCCM), CCCM with Multiple Downsampling Filters (MDF), Gradient Linear-Convolutional Cross-Component Model (GL-CCCM), Gradient Linear Model (GLM), or Local-Boosting Cross-Component Prediction (LBCCP). In some other embodiments, the CCP model may be computed in real time from training samples adjacent to the current block. Alternatively, the CCP model may be computed in real time from training samples in a reference block.
[1297] In some embodiments, all candidates in the CCP candidate list may be derived or computed in real time. Alternatively, the CCP candidate list may include multiple CCP model candidates derived or computed in real time. In some embodiments, LBCCP-based model candidates may be inserted into the CCP candidate list. In some embodiments, whether a model candidate is LBCCP-encoded can be inherited from a previous block. In some embodiments, whether a model candidate is LBCCP-encoded can be derived based on a template. In some other embodiments, whether a model candidate is LBCCP-encoded can be computed based on a template. Alternatively, whether a model candidate is LBCCP-encoded can be determined based on a template. For example, whether a model candidate is LBCCP-encoded can be determined by comparing the template cost using a low-pass filter with the template cost without a low-pass filter.
[1298] In some embodiments, candidates derived or computed in real time and inherited candidates can be included in the CCP candidate list. Furthermore, candidates derived or computed in real time can precede inherited candidates. In some embodiments, CCP models in the CCP candidate list can be sorted or reordered. In some embodiments, the sorting process can be based on template cost. In some embodiments, candidate indices can be signaled to indicate candidate models. In this case, a candidate model is selected for encoding / decoding the current block. Alternatively, the candidate model with the lowest cost after the sorting process can be used for encoding / decoding the current block. For example, the candidate model with the lowest cost can be used without signaling candidate indices.
[1299] In some embodiments, CCP models in the CCP candidate list can be selected based on the encoder side. In this case, CCP models in the CCP candidate list can be transmitted via signaling in the bitstream. In some embodiments, at least one of inter-frame CCP mode or intra-frame CCP mode can be transmitted via signaling as a separate mode. In this case, the model can be independent of the following modes: inter-frame CCCM mode, inter-frame CCPmerge mode, intra-frame CCPmerge mode, intra-frame CCLM mode, or intra-frame CCCM mode. Alternatively, at least one of the inter-frame CCP mode or intra-frame CCP mode can be transmitted via signaling as a sub-mode of one of the following: inter-frame CCCM mode, inter-frame CCPmerge mode, intra-frame CCPmerge mode, intra-frame CCLM mode, or intra-frame CCCM mode. In some embodiments, at least one of the inter-frame CCP mode or intra-frame CCP mode can be used for chroma inter-frame encoding and decoding. Alternatively, at least one of the inter-frame CCP mode or intra-frame CCP mode can be used for chroma intra-frame encoding and decoding.
[1300] In some embodiments, at least one of inter-frame CCP mode or intra-frame CCP mode may be used in at least one of single-tree or dual-tree configurations. In some other embodiments, at least one of inter-frame CCP mode or intra-frame CCP mode may be used for chroma encoding / decoding. In some embodiments, at least one of inter-frame CCP mode or intra-frame CCP mode may be used in inter-frame stripes. For example, an inter-frame stripe may be a B-strip or a P-strip. In some other embodiments, at least one of inter-frame CCP mode or intra-frame CCP mode may be used in intra-frame stripes. For example, an intra-frame stripe may be an I-strip.
[1301] In some embodiments, an indication of whether to apply at least one of inter-frame CCP mode or intra-frame CCP mode based on a CCP candidate list and / or how to apply at least one of inter-frame CCP mode or intra-frame CCP mode based on a CCP candidate list may be indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level. For example, an indication of whether to apply at least one of inter-frame CCP mode or intra-frame CCP mode based on a CCP candidate list and / or how to apply at least one of inter-frame CCP mode or intra-frame CCP mode based on a CCP candidate list may be indicated at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[1302] In some embodiments, an indication of whether at least one of inter-frame CCP mode or intra-frame CCP mode is applied based on a CCP candidate list and / or how at least one of inter-frame CCP mode or intra-frame CCP mode is applied based on a CCP candidate list may be included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[1303] In some embodiments, method 3400 may further include: determining, based on the encoded and decoded information of the video units of the video, whether to apply at least one of the inter-frame CCP mode or the intra-frame CCP mode based on the CCP candidate list and / or how to apply at least one of the inter-frame CCP mode or the intra-frame CCP mode based on the CCP candidate list, wherein the encoded and decoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type or picture type.
[1304] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by a video process apparatus. The method includes: applying at least one of an inter-frame CCP mode or an intra-frame CCP mode based on a cross-component prediction (CCP) candidate list, wherein the CCP candidate list is derived based on a decoder; and generating a bitstream based on at least one of the inter-frame CCP mode or the intra-frame CCP mode.
[1305] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: applying at least one of an inter-frame CCP mode or an intra-frame CCP mode based on a Cross-Component Prediction (CCP) candidate list, wherein the CCP candidate list is derived based on a decoder; generating a bitstream based on at least one of the inter-frame CCP mode or the intra-frame CCP mode; and storing the bitstream in a non-transitory computer-readable recording medium.
[1306] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[1307] Item 1. A method for video processing, comprising: for a conversion between a current block of video and a bitstream of the video, deriving at least one of a filter or a model based on decoding information; applying the filter or at least one of the models to the encoding and decoding of at least one of the current block or a future block; and performing the conversion based on the filter or at least one of the models.
[1308] Item 2. The method according to Item 1, wherein at least one of the filters or the model comprises a set of coefficients.
[1309] Item 3. The method according to Item 1, wherein at least one of the filters or the model includes at least one of the following: an inter-frame cross-component prediction (CCP) model or an intra-frame CCP model.
[1310] Item 4. The method according to Item 3, wherein the coefficients of at least one of the filters or the model are derived based on cross-component prediction.
[1311] Item 5. The method according to Item 1, wherein at least one of the filters or the models includes a non-CCP model.
[1312] Item 6. The method according to Item 5, wherein the non-CCP model includes at least one of the following: extended information filter (EIF), local illumination compensation (LIC) model, intra block copy (IBC) filter, intra template matching prediction (intraTMP) filter, template filter, decoder-side intra-mode derivation (DIMD) filter, or template-based intra-mode derivation (TIMD) filter.
[1313] Item 7. The method according to Item 4 or 5, wherein the coefficients of at least one of the filters or the model are derived based on the cross-component prediction.
[1314] Item 8. The method according to Item 1, wherein at least one of the filters or the model is derived based on a process of minimizing the difference between the values of luminance samples and chrominance samples of one of the reference, training regions or blocks.
[1315] Item 9. The method according to Item 8, wherein the model includes the CCP model.
[1316] Item 10. The method according to Item 8, wherein the model is applied to brightness reconstruction samples of one of the reference, the training region, or the block to obtain the resulting model estimation samples.
[1317] Item 11. The method according to Item 10, wherein the minimization process is based on the obtained model to estimate the values of the chromaticity reconstructed samples of one of the reference, the training region, or the block.
[1318] Item 12. The method according to Item 8, wherein one of the reference, the training region, or the block is adjacent to the current block; or wherein one of the reference, the training region, or the block is not adjacent to the current block; or wherein one of the reference, the training region, or the block is temporally co-located with the current block.
[1319] Item 13. The method according to Item 8, wherein one of the reference, the training region, or the block is derived based on a block vector.
[1320] Item 14. The method according to Item 8, wherein one of the reference, the training region, or the block is derived based on a motion vector.
[1321] Item 15. The method according to Item 1, wherein at least one of the filters or the model is derived based on a process of minimizing the difference between the predicted samples of the current block and the reconstructed samples of the current block.
[1322] Item 16. The method according to Item 15, wherein the model includes a non-CCP model.
[1323] Item 17. The method according to Item 15, wherein the model is applied to the predicted samples of the current block to obtain the resulting model estimate samples.
[1324] Item 18. The method according to Item 17, wherein the minimization process is based on the values of the sample points estimated by the obtained model and the reconstructed sample points of the current block.
[1325] Item 19. The method according to Item 15, wherein the predicted sample points of the current block are derived based on the block vector.
[1326] Item 20. The method according to Item 15, wherein the predicted samples of the current block are derived based on motion vectors.
[1327] Item 21. The method according to Item 1, wherein at least one of the filters or the model is derived based on a process of minimizing the difference between non-adjacent reconstructed samples and the current reconstructed sample.
[1328] Item 22. The method according to Item 21, wherein the model includes a non-CCP model.
[1329] Item 23. The method according to Item 21, wherein the non-adjacent reconstructed samples are located by block vectors.
[1330] Item 24. The method according to Item 21, wherein the non-adjacent reconstructed samples are located by a predetermined position.
[1331] Item 25. The method according to Item 24, wherein the predetermined position depends on the width and / or height of the current block.
[1332] Item 26. The method according to Item 21, wherein the model is applied to the non-adjacent reconstructed samples to obtain the resulting model estimate samples.
[1333] Item 27. The method according to Item 26, wherein the minimization process is based on the obtained model to estimate the values of the sample points and the reconstructed sample points of the current block.
[1334] Item 28. The method according to Item 1, wherein at least one of the filters or the model is derived based on a process of minimizing the difference between the reference image sample and the currently reconstructed sample.
[1335] Item 29. The method according to Item 28, wherein the model includes a non-CCP model.
[1336] Item 30. The method according to Item 28, wherein the reference image sample points are located by block vectors.
[1337] Item 31. The method according to Item 28, wherein the reference image sample is positioned by a predetermined location.
[1338] Item 32. The method according to Item 31, wherein the predetermined position is co-located with the current block.
[1339] Item 33. The method according to Item 28, wherein the model is applied to the reference image samples to obtain the resulting model estimation samples.
[1340] Item 34. The method according to Item 33, wherein the minimization process is based on the obtained model to estimate the values of the sample points and the reconstructed sample points of the current block.
[1341] Item 35. The method according to Item 1, wherein at least one of the filters or the model is derived based on a process of minimizing the difference between the reference template and the current template.
[1342] Item 36. The method according to Item 35, wherein the model includes either a non-CCP model or a CCP model.
[1343] Item 37. The method according to Item 35, wherein the reference template is adjacent to the reference block; or wherein the reference template is not adjacent to the reference block.
[1344] Item 38. The method according to Item 35, wherein the current template is adjacent to the current block; or wherein the current template is not adjacent to the current block.
[1345] Item 39. The method according to Item 35, wherein the model is applied to reference template samples to obtain the resulting model estimate samples.
[1346] Item 40. The method according to Item 39, wherein the minimization process is based on the obtained model to estimate the sample points and reconstruct the values of the current template sample points.
[1347] Item 41. The method according to Item 1, wherein at least one of the filters or the model is computed based on the training set.
[1348] Item 42. The method according to Item 41, wherein the training set is determined based on a set of samples from at least one of the following: the current block, a block vector (BV) guided reference block, a motion vector (MV) guided reference block, a template, a non-adjacent block, an adjacent block, a temporal co-location block, a temporal block adjacent to a co-location block, or a temporal block not adjacent to a co-location block.
[1349] Item 43. The method according to Item 42, wherein multiple samples are used to determine the training set.
[1350] Item 44. The method according to Item 42, wherein training samples for calculating the coefficients of the model are derived from previously encoded / decoded blocks, wherein the previously encoded / decoded blocks are encoded / decoded by the model.
[1351] Item 45. The method according to Item 44, wherein the model is based on at least one of the following: a cross-component linear model (CCLM), an intra-convolutional cross-component model (CCCM), an inter-frame CCCM, a block vector guided CCCM (BVG-CCCM), a CCCM with multiple downsampling filters (MDF), a gradient linear-convolutional cross-component model (GL-CCCM), or a gradient linear model (GLM).
[1352] Item 46. The method according to Item 44, wherein the model is based on at least one of the following: EIF, Local Illumination Compensation (LIC) model, Intra-Block Copy (IBC) filter, Intra-Template Matching Prediction (intraTMP) filter, template filter, Decoder-Side Intra-Mode Derivation (DIMD) filter, or Template-Based Intra-Mode Derivation (TIMD) filter.
[1353] Item 47. The method according to Item 1, wherein at least one of the derived filter or the derived model includes at least one of a linear model, a nonlinear model, or a convolution-based model.
[1354] Item 48. The method according to Item 47, wherein at least one of the derived filter or the derived model includes at least one of the following: similar to CCLM, similar to intra-frame CCCM, similar to BVG-CCCM, similar to inter-frame CCCM, similar to LIC model, or similar to EIF.
[1355] Item 49. The method according to Item 1, wherein the derived model is used for at least one of the luma encoding / decoding or the chroma encoding / decoding of the current block.
[1356] Item 50. The method according to Item 1, wherein the derived model is stored in a cache for encoding and decoding of future blocks.
[1357] Item 51. The method according to Item 1, wherein at least one of the derived filter or the derived model is used in at least one of a single tree or a dual tree.
[1358] Item 52. The method according to Item 1, wherein at least one of the derived filter or the derived model is used for chroma encoding / decoding.
[1359] Item 53. The method according to Item 1, wherein at least one of the derived filter or the derived model is used in the inter-frame stripe.
[1360] Item 54. The method according to Item 53, wherein the inter-frame stripe is a B stripe or a P stripe.
[1361] Item 55. The method according to Item 1, wherein at least one of the derived filter or the derived model is used in an intra-frame stripe.
[1362] Item 56. The method according to Item 55, wherein the intra-frame stripe is an I-strip.
[1363] Item 57. The method according to any one of items 1-56, wherein an indication of whether at least one of the filter or the model is derived based on the decoding information and / or how the filter or at least one of the model is derived based on the decoding information is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[1364] Item 58. The method according to any one of items 1-56, wherein an indication of whether at least one of the filter or the model is derived based on the decoding information and / or how the filter or at least one of the model is derived based on the decoding information is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice header.
[1365] Item 59. The method according to any one of items 1-56, wherein an indication of whether at least one of the filter or the model is derived based on the decoding information and / or how the filter or at least one of the model is derived based on the decoding information is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[1366] Item 60. The method according to any one of items 1-56 further comprises: determining, based on the encoded and decoded information of the video units of the video, whether to derive at least one of the filter or the model based on the decoded information and / or how to derive at least one of the filter or the model based on the decoded information, wherein the encoded and decoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[1367] Item 61. A method for video processing, comprising: for a current block of video and a bitstream of the video, applying at least one of an inter-frame CCP mode or an intra-frame CCP mode based on a cross-component prediction (CCP) candidate list, wherein the CCP candidate list is derived based on a decoder; and performing the conversion based on at least one of the inter-frame CCP mode or the intra-frame CCP mode.
[1368] Item 62. The method according to Item 61, wherein the CCP candidate list includes at least one CCP model candidate derived or calculated in real time.
[1369] Item 63. The method according to Item 62, wherein the CCP model is based on at least one of the following: a cross-component linear model (CCLM), an intra-convolutional cross-component model (CCCM), an inter-frame CCCM, a block vector guided CCCM (BVG-CCCM), a CCCM with multiple downsampling filters (MDF), a gradient linear-convolutional cross-component model (GL-CCCM), a gradient linear model (GLM), or a locally enhanced cross-component prediction (LBCCP).
[1370] Item 64. The method according to Item 62, wherein the CCP model is computed in real time from training samples adjacent to the current block.
[1371] Item 65. The method according to Item 62, wherein the CCP model is computed in real time from training samples in the reference block.
[1372] Item 66. The method according to Item 61, wherein all candidates in the CCP candidate list are derived or calculated in real time.
[1373] Item 67. The method according to Item 66, wherein the CCP candidate list includes multiple real-time derived or computed CCP model candidates.
[1374] Item 68. The method according to Item 61, wherein LBCCP-based model candidates are inserted into the CCP candidate list.
[1375] Item 69. The method according to Item 68, wherein whether the model candidate is LBCCP encoded or decoded is inherited from the previous block.
[1376] Item 70. The method according to Item 68, wherein whether a model candidate is LBCCP encoded or decoded is derived based on a template; or wherein whether a model candidate is LBCCP encoded or decoded is calculated based on a template; or wherein whether a model candidate is LBCCP encoded or decoded is determined based on a template.
[1377] Item 71. The method according to Item 70, wherein whether the model candidate is LBCCP encoded or decoded is determined by comparing the template cost with a low-pass filter and the template cost without a low-pass filter.
[1378] Item 72. The method according to Item 61, wherein candidates derived or computed in real time and inherited candidates are included in the CCP candidate list.
[1379] Item 73. The method according to Item 72, wherein the candidate derived or computed in real time precedes the inherited candidate.
[1380] Item 74. The method according to Item 61, wherein the CCP models in the CCP candidate list are sorted or reordered.
[1381] Item 75. The method described in Item 74, wherein the sorting process is based on template cost.
[1382] Item 76. The method according to Item 74, wherein a candidate index is signaled to indicate a candidate model, wherein the candidate model is selected for encoding and decoding of the current block.
[1383] Item 77. The method according to Item 74, wherein the candidate model with the lowest cost after the sorting process is used for encoding and decoding the current block.
[1384] Item 78. The method according to Item 77, wherein the candidate model with the lowest cost is used without signaling the candidate index.
[1385] Item 79. The method according to Item 61, wherein the CCP model in the CCP candidate list is selected based on the encoder side, and wherein the CCP model in the CCP candidate list is transmitted via signal in the bitstream.
[1386] Item 80. The method according to Item 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is transmitted via signaling as a separate mode, wherein the mode is independent of the following modes: inter-frame CCCM mode, inter-frame CCPmerge mode, intra-frame CCPmerge mode, intra-frame CCLM mode, or intra-frame CCCM mode.
[1387] Item 81. The method according to Item 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is transmitted via signaling as a sub-mode of one of the following: inter-frame CCCM mode, inter-frame CCPmerge mode, intra-frame CCPmerge mode, intra-frame CCLM mode, or intra-frame CCCM mode.
[1388] Item 82. The method according to Item 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used for chroma inter-frame encoding and decoding.
[1389] Item 83. The method according to Item 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used for chroma intra-frame encoding and decoding.
[1390] Item 84. The method according to Item 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used in at least one of a single tree or a dual tree.
[1391] Item 85. The method according to Item 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used for chroma encoding / decoding.
[1392] Item 86. The method according to Item 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used in an inter-frame stripe.
[1393] Item 87. The method according to Item 86, wherein the inter-frame stripe is a B stripe or a P stripe.
[1394] Item 88. The method according to Item 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used in an intra-frame stripe.
[1395] Item 89. The method according to Item 88, wherein the intra-frame stripe is an I-strip.
[1396] Item 90. The method according to any one of items 61-89, wherein an indication of whether at least one of the inter-frame CCP mode or the intra-frame CCP mode is applied based on the CCP candidate list and / or how to apply at least one of the inter-frame CCP mode or the intra-frame CCP mode based on the CCP candidate list is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[1397] Item 91. The method according to any one of items 61-89, wherein an indication of whether at least one of the inter-frame CCP mode or the intra-frame CCP mode is applied based on the CCP candidate list and / or how at least one of the inter-frame CCP mode or the intra-frame CCP mode is applied based on the CCP candidate list is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[1398] Item 92. The method according to any one of items 61-89, wherein an indication of whether at least one of the inter-frame CCP mode or the intra-frame CCP mode is applied based on the CCP candidate list and / or how at least one of the inter-frame CCP mode or the intra-frame CCP mode is applied based on the CCP candidate list is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[1399] Item 93. The method according to any one of items 61-89 further comprises: determining, based on encoded and decoded information of the video units of the video, whether to apply at least one of the inter-frame CCP mode or the intra-frame CCP mode based on the CCP candidate list and / or how to apply at least one of the inter-frame CCP mode or the intra-frame CCP mode based on the CCP candidate list, wherein the encoded and decoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[1400] Item 94. The method according to any one of items 1-93, wherein at least one of the following is applicable: a non-CCP model, a non-CCP filter, a non-CCP computation, a non-CCP prediction, or a non-CCP application.
[1401] Item 95. The method according to Item 94, wherein the coefficients of at least one of the non-CCP filter or non-CCP model are derived based on component prediction.
[1402] Item 96. The method according to Item 95, wherein the component prediction is a luminance-to-luminance prediction.
[1403] Item 97. The method according to Item 94, wherein the non-CCP filter or at least one of the non-CCP models is based on at least one of the following: EIF, Local Illumination Compensation (LIC) model, Intra-Block Copy (IBC) filter, Intra-Template Matching Prediction (intraTMP) filter, Reference Sample Filter, Template Filter, Decoder-Side Intra-Mode Derivation (DIMD) Filter, or Template-Based Intra-Mode Derivation (TIMD) Filter.
[1404] Item 98. The method according to any one of items 1-97, wherein the conversion includes encoding the current block into the bitstream.
[1405] Item 99. The method according to any one of items 1-97, wherein the conversion includes decoding the current block from the bitstream.
[1406] Item 100. An apparatus for video processing, comprising a processor and a non-transitory memory having i...
Claims
1. A method for video processing, comprising: For the conversion between the current block of the video and the bitstream of the video, derive at least one of a filter or model based on the decoding information; Applying at least one of the filters or the model to the encoding / decoding of at least one of the current block or future blocks; and The transformation is performed based on at least one of the filters or the models.
2. The method according to claim 1, wherein at least one of the filter or the model comprises a set of coefficients.
3. The method of claim 1, wherein at least one of the filters or the models comprises at least one of the following: an inter-frame cross-component prediction (CCP) model or an intra-frame CCP model.
4. The method of claim 3, wherein the coefficients of at least one of the filters or the model are derived based on cross-component prediction.
5. The method of claim 1, wherein at least one of the filters or the models comprises a non-CCP model.
6. The method according to claim 5, wherein the non-CCP model comprises at least one of the following: an extended information filter (EIF), a local illumination compensation (LIC) model, an intra-block copy (IBC) filter, an intra-template matching prediction (intraTMP) filter, a template filter, a decoder-side intra-mode derivation (DIMD) filter, or a template-based intra-mode derivation (TIMD) filter.
7. The method of claim 4 or 5, wherein the coefficients of at least one of the filters or the model are derived based on the cross-component prediction.
8. The method of claim 1, wherein at least one of the filters or the model is derived based on a process of minimizing the difference between the values of luminance samples and chrominance samples in a reference, training region, or block.
9. The method of claim 8, wherein the model comprises a CCP model.
10. The method of claim 8, wherein the model is applied to brightness reconstruction samples of one of the reference, the training region, or the block to obtain the resulting model estimation samples.
11. The method of claim 10, wherein the minimization process is based on the obtained model to estimate the values of the chromaticity reconstructed samples of one of the reference, the training region, or the block.
12. The method of claim 8, wherein one of the reference, the training region, or the block is adjacent to the current block; or The reference, the training region, or the block is not adjacent to the current block; or The reference, the training region, or the block is co-located with the current block in the time domain.
13. The method of claim 8, wherein one of the reference, the training region, or the block is derived based on a block vector.
14. The method of claim 8, wherein one of the reference, the training region, or the block is derived based on a motion vector.
15. The method of claim 1, wherein at least one of the filters or the model is derived based on a process of minimizing the difference between the predicted samples of the current block and the reconstructed samples of the current block.
16. The method of claim 15, wherein the model comprises a non-CCP model.
17. The method of claim 15, wherein the model is applied to the predicted samples of the current block to obtain the resulting model estimate samples.
18. The method of claim 17, wherein the minimization process is based on the values of the sample points estimated by the obtained model and the reconstructed sample points of the current block.
19. The method of claim 15, wherein the predicted samples of the current block are derived based on the block vector.
20. The method of claim 15, wherein the predicted samples of the current block are derived based on motion vectors.
21. The method of claim 1, wherein at least one of the filters or the model is derived based on a process of minimizing the difference between non-adjacent reconstructed samples and the current reconstructed sample.
22. The method of claim 21, wherein the model includes a non-CCP model.
23. The method of claim 21, wherein the non-adjacent reconstructed samples are located by block vectors.
24. The method of claim 21, wherein the non-adjacent reconstructed samples are located by a predetermined position.
25. The method of claim 24, wherein the predetermined position depends on the width of the current block and / or the height of the current block.
26. The method of claim 21, wherein the model is applied to the non-adjacent reconstructed samples to obtain the resulting model estimation samples.
27. The method of claim 26, wherein the minimization process is based on the obtained model to estimate the values of the sample points and the reconstructed sample points of the current block.
28. The method of claim 1, wherein at least one of the filters or the model is derived based on a process of minimizing the difference between reference image samples and currently reconstructed samples.
29. The method of claim 28, wherein the model includes a non-CCP model.
30. The method of claim 28, wherein the reference image sample points are located by block vectors.
31. The method of claim 28, wherein the reference image sample is positioned by a predetermined location.
32. The method of claim 31, wherein the predetermined position is co-located with the current block.
33. The method of claim 28, wherein the model is applied to the reference image samples to obtain the resulting model estimation samples.
34. The method of claim 33, wherein the minimization process is based on the obtained model to estimate the values of the sample points and the reconstructed sample points of the current block.
35. The method of claim 1, wherein at least one of the filters or the model is derived based on a process of minimizing the difference between the reference template and the current template.
36. The method of claim 35, wherein the model comprises either a non-CCP model or a CCP model.
37. The method of claim 35, wherein the reference template is adjacent to the reference block; or The reference template and the reference block are not adjacent.
38. The method of claim 35, wherein the current template is adjacent to the current block; or The current template is not adjacent to the current block.
39. The method of claim 35, wherein the model is applied to reference template samples to obtain the resulting model estimation samples.
40. The method of claim 39, wherein the minimization process is based on the obtained model to estimate the sample points and reconstruct the values of the current template sample points.
41. The method of claim 1, wherein at least one of the filters or the model is computed based on a training set.
42. The method of claim 41, wherein the training set is determined based on a set of samples from at least one of the following: the current block, a block vector (BV) guided reference block, a motion vector (MV) guided reference block, a template, a non-adjacent block, an adjacent block, a temporal co-location block, a temporal block adjacent to the co-location block, or a temporal block not adjacent to the co-location block.
43. The method of claim 42, wherein a plurality of samples are used to determine the training set.
44. The method of claim 42, wherein training samples for calculating the coefficients of the model are derived from previously encoded / decoded blocks, wherein the previously encoded / decoded blocks are encoded / decoded by the model.
45. The method of claim 44, wherein the model is based on at least one of the following: a cross-component linear model (CCLM), an intra-convolutional cross-component model (CCCM), an inter-frame CCCM, a block vector guided CCCM (BVG-CCCM), a CCCM with multiple downsampling filters (MDF), a gradient linear-convolutional cross-component model (GL-CCCM), or a gradient linear model (GLM).
46. The method of claim 44, wherein the model is based on at least one of the following: EIF, Local Illumination Compensation (LIC) model, Intra-Block Copy (IBC) filter, Intra-Template Matching Prediction (intraTMP) filter, template filter, Decoder-Side Intra-Mode Derivation (DIMD) filter, or Template-Based Intra-Mode Derivation (TIMD) filter.
47. The method of claim 1, wherein at least one of the derived filter or the derived model comprises at least one of a linear model, a nonlinear model, or a convolution-based model.
48. The method of claim 47, wherein at least one of the derived filter or the derived model comprises at least one of the following: similar to CCLM, similar to intra-frame CCCM, similar to BVG-CCCM, similar to inter-frame CCCM, similar to LIC model, or similar to EIF.
49. The method of claim 1, wherein the derived model is used for at least one of the luma encoding / decoding or the chroma encoding / decoding of the current block.
50. The method of claim 1, wherein the derived model is stored in a cache for encoding and decoding of future blocks.
51. The method of claim 1, wherein at least one of the derived filter or the derived model is used in at least one of a single tree or a dual tree.
52. The method of claim 1, wherein at least one of the derived filter or the derived model is used for chroma encoding / decoding.
53. The method of claim 1, wherein at least one of the derived filter or the derived model is used in inter-frame stripes.
54. The method of claim 53, wherein the inter-frame stripe is a B-strip or a P-strip.
55. The method of claim 1, wherein at least one of the derived filter or the derived model is used in an intra-frame stripe.
56. The method of claim 55, wherein the intra-frame stripe is an I-strip.
57. The method according to any one of claims 1-56, wherein an indication of whether at least one of the filter or the model is derived based on the decoding information and / or how the filter or at least one of the model is derived based on the decoding information is indicated at one of the following: sequence level, Image group level, Image quality, strip level, or Film series level.
58. The method according to any one of claims 1-56, wherein an indication of whether at least one of the filter or the model is derived based on the decoding information and / or how the filter or at least one of the model is derived based on the decoding information is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
59. The method according to any one of claims 1-56, wherein an indication of whether at least one of the filter or the model is derived based on the decoding information and / or how the filter or at least one of the model is derived based on the decoding information is included in one of the following: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Virtual Pipeline Data Unit (VPDU). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or This includes regions containing more than one sample point or pixel.
60. The method according to any one of claims 1-56, further comprising: Based on the encoded and decoded information of the video unit of the video, it is determined whether and / or how to derive at least one of the filter or the model based on the decoded information, wherein the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.
61. A method for video processing, comprising: For the conversion between the current block of the video and the bitstream of the video, at least one of inter-frame CCP mode or intra-frame CCP mode is applied based on a cross-component prediction (CCP) candidate list, wherein the CCP candidate list is derived based on the decoder; as well as The conversion is performed based on at least one of the inter-frame CCP mode or the intra-frame CCP mode.
62. The method of claim 61, wherein the CCP candidate list includes at least one CCP model candidate derived or calculated in real time.
63. The method of claim 62, wherein the CCP model is based on at least one of the following: a cross-component linear model (CCLM), an intra-convolutional cross-component model (CCCM), an inter-frame CCCM, a block vector guided CCCM (BVG-CCCM), a CCCM with multiple downsampling filters (MDF), a gradient linear-convolutional cross-component model (GL-CCCM), a gradient linear model (GLM), or a locally enhanced cross-component prediction (LBCCP).
64. The method of claim 62, wherein the CCP model is computed in real time from training samples adjacent to the current block.
65. The method of claim 62, wherein the CCP model is computed in real time from training samples in the reference block.
66. The method of claim 61, wherein all candidates in the CCP candidate list are derived or calculated in real time.
67. The method of claim 66, wherein the CCP candidate list includes a plurality of real-time derived or computed CCP model candidates.
68. The method of claim 61, wherein LBCCP-based model candidates are inserted into the CCP candidate list.
69. The method of claim 68, wherein whether the model candidate is LBCCP encoded or decoded is inherited from the previous block.
70. The method of claim 68, wherein whether the model candidate is LBCCP encoded / decoded is derived based on the template; or Whether a model candidate is encoded or decoded using LBCCP is calculated based on a template; or Whether a model candidate is encoded or decoded using LBCCP is determined based on a template.
71. The method of claim 70, wherein whether the model candidate is LBCCP encoded or decoded is determined by comparing the template cost with a low-pass filter and the template cost without the low-pass filter.
72. The method of claim 61, wherein the candidates derived or computed in real time and the inherited candidates are included in the CCP candidate list.
73. The method of claim 72, wherein the candidate derived or computed in real time precedes the inherited candidate.
74. The method of claim 61, wherein the CCP models in the CCP candidate list are sorted or reordered.
75. The method of claim 74, wherein the sorting process is based on template cost.
76. The method of claim 74, wherein a candidate index is signaled to indicate a candidate model, wherein the candidate model is selected for encoding and decoding of the current block.
77. The method of claim 74, wherein the candidate model with the lowest cost after the sorting process is used for the encoding and decoding of the current block.
78. The method of claim 77, wherein the candidate model with the lowest cost is used without transmitting the candidate index via signal transmission.
79. The method of claim 61, wherein the CCP model in the CCP candidate list is selected based on the encoder side, and wherein the CCP model in the CCP candidate list is transmitted via signal in the bitstream.
80. The method of claim 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is transmitted via signaling as a separate mode, wherein the mode is independent of the following modes: inter-frame CCCM mode, inter-frame CCP Merge mode, intra-frame CCP Merge mode, intra-frame CCLM mode, or intra-frame CCCM mode.
81. The method of claim 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is transmitted via signal as a sub-mode of one of the following: inter-frame CCCM mode, inter-frame CCP Merge mode, intra-frame CCP Merge mode, intra-frame CCLM mode, or intra-frame CCCM mode.
82. The method of claim 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used for chroma inter-frame encoding and decoding.
83. The method of claim 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used for chroma intra-frame encoding and decoding.
84. The method of claim 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used in at least one of a single tree or a dual tree.
85. The method of claim 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used for chroma encoding / decoding.
86. The method of claim 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used in an inter-frame stripe.
87. The method of claim 86, wherein the inter-frame stripe is a B-strip or a P-strip.
88. The method of claim 61, wherein at least one of the inter-frame CCP mode or the intra-frame CCP mode is used in an intra-frame stripe.
89. The method of claim 88, wherein the intra-frame stripe is an I-strip.
90. The method according to any one of claims 61-89, wherein an indication of whether at least one of the inter-frame CCP mode or the intra-frame CCP mode is applied based on the CCP candidate list and / or how at least one of the inter-frame CCP mode or the intra-frame CCP mode is applied based on the CCP candidate list is indicated at one of the following: sequence level, Image group level, Image quality, strip level, or Film series level.
91. The method according to any one of claims 61-89, wherein an indication of whether to apply at least one of the inter-frame CCP mode or the intra-frame CCP mode based on the CCP candidate list and / or how to apply at least one of the inter-frame CCP mode or the intra-frame CCP mode based on the CCP candidate list is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
92. The method according to any one of claims 61-89, wherein an indication of whether at least one of the inter-frame CCP mode or the intra-frame CCP mode is applied based on the CCP candidate list and / or how at least one of the inter-frame CCP mode or the intra-frame CCP mode is applied based on the CCP candidate list is included in one of the following: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Virtual Pipeline Data Unit (VPDU). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or This includes regions containing more than one sample point or pixel.
93. The method according to any one of claims 61-89, further comprising: Based on the encoded and decoded information of the video unit of the video, it is determined whether to apply at least one of the inter-frame CCP mode or the intra-frame CCP mode based on the CCP candidate list and / or how to apply at least one of the inter-frame CCP mode or the intra-frame CCP mode based on the CCP candidate list, wherein the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.
94. The method according to any one of claims 1-93, wherein at least one of the following is applicable: non-CCP model, non-CCP filter, non-CCP computation, non-CCP prediction, or non-CCP application.
95. The method of claim 94, wherein the coefficients of at least one of the non-CCP filter or non-CCP model are derived based on component prediction.
96. The method of claim 95, wherein the component prediction is a luminance-to-luminance prediction.
97. The method of claim 94, wherein at least one of the non-CCP filter or the non-CCP model is based on at least one of the following: EIF, Local Illumination Compensation (LIC) model, Intra-Block Copy (IBC) filter, Intra-Template Matching Prediction (intraTMP) filter, Reference Sample Filter, Template Filter, Decoder-Side Intra-Mode Derivation (DIMD) Filter, or Template-Based Intra-Mode Derivation (TIMD) Filter.
98. The method according to any one of claims 1-97, wherein the conversion comprises encoding the video unit into the bitstream.
99. The method according to any one of claims 1-97, wherein the conversion comprises decoding the video unit from the bitstream.
100. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-99.
101. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-99.
102. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Derive at least one of the filters or models based on the decoded information; At least one of the filters or the models is applied to the encoding and decoding of at least one of the current or future blocks of the video; as well as The bitstream is generated based on at least one of the filters or the model.
103. A method for storing a bitstream of video, comprising: Derive at least one of the filters or models based on the decoded information; At least one of the filters or the models is applied to the encoding and decoding of at least one of the current or future blocks of the video; The bitstream is generated based on at least one of the filters or the models; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
104. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Based on a cross-component prediction (CCP) candidate list, at least one of an inter-frame CCP mode or an intra-frame CCP mode is applied, wherein the CCP candidate list is derived based on the decoder. as well as The bitstream is generated based on at least one of the inter-frame CCP mode or the intra-frame CCP mode.
105. A method for storing a bitstream of video, comprising: Based on a cross-component prediction (CCP) candidate list, at least one of an inter-frame CCP mode or an intra-frame CCP mode is applied, wherein the CCP candidate list is derived based on the decoder. The bitstream is generated based on at least one of the inter-frame CCP mode or the intra-frame CCP mode; as well as The bitstream is stored in a non-transitory computer-readable recording medium.