Method and apparatus for improving video coding and decoding through stored information and implicit derivation

By using cross-component model-related patterns, the problem of insufficient efficiency and quality in chroma component encoding and decoding in existing technologies is solved, enabling more efficient film data storage and reconstruction, and improving encoding and decoding performance and quality.

CN121444451APending Publication Date: 2026-01-30MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480045680.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-05
Filing Date
2024-07-05
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing multi-functional video encoding and decoding technologies are insufficient in terms of efficiency and quality when processing cross-component information, especially chroma components. This results in significant damage to video data during storage and reconstruction, affecting encoding and decoding performance and complexity.

Method used

Cross-component model-related modes are adopted, including cross-component linear model (CCLM), convolutional cross-component model (CCCM), gradient linear model (GLM), and cross-component residual model (CCRM). By using the target CCP model to generate predictions for the current block, the encoding and decoding efficiency and quality of chroma blocks are improved. Furthermore, model parameters are stored and inherited at the encoder and decoder ends to optimize the encoding and decoding process.

Benefits of technology

It improves the encoding and decoding performance of cross-component information, reduces the loss of video data during storage and reconstruction, reduces encoding and decoding complexity, and improves video quality and encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121444451A_ABST
    Figure CN121444451A_ABST
Patent Text Reader

Abstract

A method and apparatus for encoding and decoding a color picture or movie including one or more cross-component model related modes using an encoding and decoding tool are disclosed. According to this method, input data associated with a current block is received, including a first color patch and a second color patch, where the input data includes pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side. The method comprises the following steps of: determining a target cross-component prediction (CCP) model of the current block according to the current block of the current block, and determining the target cross-component prediction (CCP) model of the current block according to the target CCP model. And storing the target CCP model. The second color patch is encoded or decoded using a target prediction generated for the current block according to the target CCP model.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is a non-provisional application and claims priority to U.S. Provisional Patent Application No. 63 / 511,921, filed July 5, 2023. This U.S. Provisional Patent Application is hereby incorporated by reference in its entirety. TECHNICAL FIELD

[0002] The present disclosure relates to video coding systems. In particular, the present disclosure relates to coding chroma components with stored cross-component information. BACKGROUND

[0003] Versatile Video Coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) which is composed of the Video Coding Experts Group (VCEG) of the International Telecommunication Union - Telecommunication Standardization Sector (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC). The standard has been published as an ISO standard: ISO / IEC 23090-3:2021, Information technology - Coding of audio-visual

[0004] Figure 1AAn exemplary adaptive inter / intra video coding system with in loop processing is shown. For intra prediction 110, prediction data is derived based on previously coded video data in the current picture. For inter prediction 112, motion estimation (ME) is performed at the encoder side, and motion compensation (MC) is performed based on the result of ME to provide prediction data derived from other pictures and motion data. A switch 114 selects either intra prediction 110 or inter prediction 112, and provides the selected prediction data to a summer 116 to form prediction error, also known as residues. The prediction error is then processed by transform (T) 118, followed by quantization (Q) 120. The transformed and quantized residues are then encoded by an entropy encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream related to the transform coefficients is then packed with side information related to intra prediction (Intra Pred.) and inter prediction (Inte Pred.) such as motion and coding mode, and other information related to parameters of in loop filter (ILPF) applied to the underlying image region. The side information related to intra prediction 110, inter prediction 112, and in loop filter (ILPF) 130, as shown, is provided to the entropy encoder 122. When inter prediction mode is used, the reference picture or pictures also have to be reconstructed at the encoder side. Therefore, the transformed and quantized residues are processed by inverse quantization (IQ) 124 and inverse transform (IT) 126 to recover the residues. The residues are then added back to the prediction data 136, and the video data is reconstructed at reconstruction (REC) 128. The reconstructed video data can be stored in a reference picture buffer 134, and used for prediction of other frames. Figure 1A

[0005] As Figure 1A ​As shown, the incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 can be subject to various impairments due to the series of processing. Therefore, before storing the reconstructed video data in the reference picture buffer (Ref. Pic. Buffer) 134, an in-loop filter (ILPF) 130 is usually applied to the reconstructed video data to improve the video quality. For example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF) can be used. The in-loop filter information can need to be included in the bitstream so that the decoder can correctly recover the required information. Therefore, the in-loop filter information is also provided to the entropy encoder 122 for inclusion in the bitstream. In Figure 1A the in-loop filter 130 is applied to the reconstructed video, and then the reconstructed samples are stored in the reference picture buffer 134. Figure 1A The system in

[0006] As shown in Figure 1B , the decoder can use the same or some of the same functional blocks as the encoder, except for the transform 118 and the quantization 120, because the decoder only needs the inverse quantization 124 and the inverse transform 126. The decoder uses an entropy decoder (Entropy Decoder) 140 instead of the entropy encoder 122 to decode the video bitstream (Input Bitstream) into the quantized transform coefficients and the required coding information (e.g., ILPF information, intra prediction information, and inter prediction information). The intra prediction 150 at the decoder side does not need to do mode search. Instead, the decoder only needs to generate the intra prediction based on the intra prediction information received from the entropy decoder 140. Furthermore, for inter prediction, the decoder only needs to do motion compensation (MC 152) based on the inter prediction information received from the entropy decoder 140, without doing motion estimation.

[0007] To improve the coding performance and / or reduce the complexity of a system using a cross-component model, methods and apparatuses for implicit derivation of a cross-component model for storage information and / or chroma blocks are disclosed herein. SUMMARY

[0008] A method and apparatus for coding a color picture or video using a coding tool that includes one or more cross-component model related modes is disclosed. According to the method, input data associated with a current block is received, including a first color block and a second color block, wherein the input data includes pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side. The current block is coded in a non-intra mode. A target cross-component prediction (CCP) model for the current block is determined. The target CCP model is stored. The second color block is encoded or decoded using a target prediction generated for the current block according to the target CCP model.

[0009] In an embodiment, the target CCP model includes cross-component mode information, cross-component model (CCM) information, or both.

[0010] In an embodiment, the target CCP model corresponds to a self-derived cross-component model or an inherited cross-component model.

[0011] In an embodiment, the target CCP model includes CCM information and model parameters of an inherited cross-component model. In an embodiment, the CCM information of the inherited cross-component model includes a cross-component linear model (CCLM), a convolutional cross-component model (CCCM), a CCCM with different filters, or a combination thereof. In another embodiment, the CCM information is refined according to previously stored CCM information.

[0012] In an embodiment, inherited model parameters associated with the CCM information are refined. In an embodiment, the inherited model parameters are refined using different types of templates and / or different number of lines.

[0013] In an embodiment, the target CCP model corresponds to an inherited cross-component model from a chroma intra-merge mode. In an embodiment, the chroma intra-merge mode is derived by merging an intra prediction of non-cross-component coding and an intra prediction of cross-component coding. In an embodiment, when CCM information is inherited from a block or a position coded by the chroma intra-merge mode, model parameters used to obtain the intra prediction of cross-component coding are inherited and further refined.

[0014] In an embodiment, the stored target CCP model is used or referred to by one or more subsequent coded blocks. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1A An exemplary adaptive Inter / Intra video coding system incorporating in-loop processing is shown.

[0016] Figure 1B An example of a decoder corresponding to the encoder in Figure 1A

[0017] Figure 2 The 16 gradient modes of GLM are shown.

[0018] Figure 3 An exemplary system block diagram of a Cross-component residual model (CCRM) is shown.

[0019] Figure 4 An example of a template used in TIMD and its reference samples is shown.

[0020] Figure 5 Five neighboring blocks used to derive VVC spatial merge candidates are shown.

[0021] Figure 6 An exemplary pattern of spatial merge candidates is shown.

[0022] Figure 7 An example of temporal candidate derivation is shown, where a scaled motion vector is derived according to POC (Picture Order Count) distance.

[0023] Figure 8 The position of the temporal candidate selected between candidates CO and CI is shown.

[0024] Figure 9 An example of weight setting proposed according to an embodiment of the present invention is shown.

[0025] Figure 10 An example of inheriting temporal neighboring model parameters is shown.

[0026] Figure 11A Two search patterns for inheriting non-neighboring spatial neighboring models are shown.

[0027] Figure 12 An example is shown, limiting temporal candidates to refer to CCM information located in the same CTU, Area 1, Area 2, or Area 3.

[0028] Figure 13 An example is shown, storing inter-coded or CCM information from CTU-level buffer to picture-level buffer, where the top-left position marked in each 2x2 grid is saved to the picture-level buffer.​

[0029] Figure 14 A flowchart of an exemplary video encoding / decoding system is shown. According to one embodiment of the invention, the system stores cross-component model information for use or reference by one or more subsequent encoding / decoding blocks.

Detailed Implementation Methods

[0030] It will be readily understood that the components of the present invention, as generally described and depicted in the figures herein, can be arranged and designed in a variety of different configurations. Therefore, the following more detailed description of embodiments of the systems and methods of the present invention, as illustrated in the figures, is not intended to limit the scope of the invention as claimed, but merely represents selected embodiments of the invention. References to “one embodiment,” “an embodiment,” or similar language throughout this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of the invention. Therefore, the phrases “in one embodiment” or “in a kind of embodiment” appearing throughout this specification do not necessarily refer to the same embodiment.

[0031] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. However, those skilled in the art will recognize that the invention can be practiced without using one or more specific details, or using other methods, components, etc. In other instances, well-known structures or operations are not shown or described in detail to avoid obscuring aspects of the invention. Embodiments of the invention can be best understood by referring to the drawings, wherein identical parts are designated by the same numbers throughout. The following description is illustrative only and simply illustrates apparatus and methods of certain selected embodiments consistent with the invention declared herein.

[0032] Cross-Component Linear Model (CCLM) Prediction To reduce redundancy across components, VVC uses a cross-component linear model (CCLM) prediction mode, where chroma samples are predicted based on reconstructed luminance samples from the same coding unit (CU) using a linear model, as shown below: (1) in Represents the predicted chromaticity samples in CU. This represents the downsampled reconstructed brightness sample from the same CU.

[0033] CCLM parameters ( and The chroma block is derived from at most four neighboring chroma samples and their corresponding downsampled luminance samples. Assuming the current chroma block size is W×H, then W' and H' are set to... W' = W, H' = H when LM_LA mode is applied; W' = W + H when LM_A mode is applied; H' = H + W when the LM_L mode is applied.

[0034] Multi-model CCLM (MMLM) In JEM (J. Chen, E. Alshina, GJ Sullivan, J.-R. Ohm, and J. Boyce, Algorithm Description of Joint Exploration Test Model 7, document JVET-G1001, ITU-T / ISO / IEC Joint Film Exploration Team (JVET), Jul. 2017), a multi-model CCLM (MMLM) mode was proposed to predict chromaticity samples of the entire CU from luminance samples using two models. In MMLM, neighboring luminance samples and neighboring chromaticity samples of the current block are divided into two groups, each group serving as a training set to derive a linear model (i.e., deriving specific α and β for a specific group). Furthermore, samples of the current luminance block are also classified according to the classification rules for neighboring luminance samples.

[0035] The threshold is calculated as the average of the neighboring reconstructed brightness samples. A neighboring sample is defined as Rec′. L If [x,y] <= the threshold, it is classified as group 1; while if the neighboring samples are Rec′ L If [x,y]>threshold, then it is classified as group 2.

[0036] (2) Local Illumination Compensation (LIC) Local Illumination Compensation (LIC) is a method for mutual prediction using neighboring samples of the current block and a reference block. It is based on a linear model using a scaling factor a and an offset b. It derives the scaling factor a and offset b by referencing neighboring samples of the current block and the reference block. Furthermore, it adaptively enables or disables it for each CU.

[0037] For more details about LIC, please refer to document JVET-C1001 (Jianle Chen et al., “Algorithm description of JointExploration Test Model 3”, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG11, 3rd Meeting: Geneva, CH, 26 May to 1 June 2016, document: JVET-C1001).

[0038] Convolutional Cross-Component Model (CCCM) In CCCM, a convolutional model is applied to improve chromaticity prediction performance. The convolutional model has a 7-tap filter, consisting of a 5-tap plus shapespace component, a nonlinear term, and a bias term.

[0039] The output of the filter is the convolution calculation between the filter coefficients and the input values, clipped to the range of valid chromaticity samples.

[0040] The filter coefficients are calculated by minimizing the MSE between the predicted and reconstructed chromaticity samples in the reference region.

[0041] MSE minimization is performed by calculating the autocorrelation matrix of the luminance input and the cross-correlation vector between the luminance input and the chrominance output. The autocorrelation matrix is ​​decomposed using LDL, and the final filter coefficients are calculated using back-substitution. This process roughly follows the calculation of ALF filter coefficients in the Enhanced Compression Model (ECM); however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations.

[0042] Gradient Linear Model (GLM) Compared to the Cross-Component Linear Model (CCLM), the Gradient Linear Model (GLM) does not use downsampled luminance values; instead, it derives a linear model using the gradient of luminance samples. Specifically, when applying GLM, the data input to the CCLM process is the downsampled luminance samples. The gradient of the brightness sample Replaced. Other parts of CCLM (e.g., parameter derivation, linear transformation of predicted samples) remain unchanged: .

[0043] In terms of signal transmission, when the current codec unit (CU) enables CCLM mode, two flags are sent to the Cb and Cr components respectively to indicate whether GLM is enabled for each component; if GLM is enabled for a component, a syntax element is further sent to select one of 16 gradient filters (2^10-2^40) for gradient calculation, such as... Figure 2 As shown. GLM can be combined with an existing CCLM by sending an additional flag in the bitstream. When this combination is applied, the filter coefficients for the input luminance samples used to derive the linear model are calculated as a combination of the gradient filter chosen by GLM and the downsampling filter of CCLM.

[0044] Intra-block copying Intra-block copy (IBC) is a tool used in HEVC Extended Screen Content Coding (SCC). It is well-known for significantly improving the encoding and decoding efficiency of screen content material. Since IBC mode is implemented as a block-level encoding / decoding mode, block matching (BM) is performed at the encoder end to find the optimal block vector (or motion vector) for each CU. Here, the block vector indicates the displacement from the current block to a reference block that has been reconstructed within the current image. The luma block vector of an IBC-encoded CU is integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-pel and 4-pel motion vector precision. IBC-encoded CUs are considered a third prediction mode in addition to intra-frame or inter-frame prediction modes. IBC mode is suitable for CUs with a width and height of 64 luma samples or less.

[0045] Cross-component residual model (CCRM) As described in JVET-AD0108 (Pekka Astola et al., “AHG12: Cross-component residual model (CCRM) for inter-frame prediction”, ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, 30th meeting of the Joint Film Experts Group (JVET), Antalya, TR, 21-28 April 2023, document: JVET-AD0108), when blocks use inter-frame prediction or intra-block copying (IBC), a cross-component residual model (CCRM) is applied to predict chroma samples from reconstructed luminance samples. Figure 3 The decoder side of the method is shown. Cross-component filters are derived using the predicted signals for both luminance and chrominance. The derived filters are applied to the reconstructed luminance signal to produce the final chrominance prediction. In step 320, filter coefficients for each chrominance component are derived using the predicted signals (i.e., predY310 and predCb 312 or predCr 314), and as shown... Figure 3As shown, in step 330, a filter is applied to the reconstructed luminance signal. The reconstructed luminance signal is formed by combining the luminance prediction (PredY) 310 and the residual luminance signal (resY) using adder 322. After applying the filter, step 330 generates the filtered predicted Cb 340 and the filtered predicted Cr 350. The reconstructed Cb signal is formed by combining the filtered predicted Cb 340 and the residual Cb signal (i.e., resCb) using adder 342. Similarly, the reconstructed Cr signal is formed by combining the filtered predicted Cr 350 and the residual Cr signal (i.e., resCr) using adder 352.

[0046] Chromaticity DM mode For chroma DM mode, the intra-prediction mode that directly inherits the corresponding (same position) luma block covering the center position of the current chroma block is directly inherited.

[0047] Decoder-Side Intra Mode Derivation (DIMD) For intra-frame prediction modes of implicitly derived blocks, texture gradient analysis is performed at both the encoder and decoder ends. This process begins with an empty gradient histogram (HoG) corresponding to 65 angular modes. The magnitudes of these entries are determined during texture gradient analysis.

[0048] Template-based Intra Mode Derivation (TIMD) Template-based intra-mode derivation (TIMD) modes implicitly derive intra-prediction modes (CUs) from neighboring templates in the encoder and decoder, instead of sending the exact intra-prediction mode bits to the decoder. For example... Figure 4 As shown, reference samples of the template are used to generate predicted samples for each candidate mode. The SATD between the predicted and reconstructed samples of the template is calculated as the cost. The intra-prediction mode with the lowest cost is selected as the TIMD mode (similar to a derivative method of the DIMD mode) and used for intra-prediction of the CU. The candidate modes may be one of the 67 intra-prediction modes in the VVC or extended to 131 intra-prediction modes. Generally, the MPM can provide clues to indicate the orientation information of the CU. Therefore, in order to reduce the intra-mode search space and utilize the characteristics of the CU, the intra-prediction modes are implicitly derived from the MPM list. Figure 4 As shown, the template's reference samples (420 and 422) are used to generate the template's predicted samples (412 and 414) for each candidate pattern of the current block (Current CU) 410.

[0049] Intra-template matching prediction (IntraTMP) is a special intra-frame prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. The encoder then issues a signal to use this mode, and the same prediction operation is performed at the decoder.

[0050] Inter-frame prediction overview For each coding unit (CU) in inter-frame prediction, motion parameters consist of motion vectors, reference picture indices, reference picture list usage indices, and additional information required for generating inter-frame prediction samples using the new encoding / decoding features of VVC. Motion parameters can be sent explicitly or implicitly. When a CU is encoded in skip mode, the CU is associated with a prediction unit (PU), with no significant residual coefficients, no encoded motion vector difference, and no reference picture index. A merging mode is specified where the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates, and additional scheduling introduced in VVC. The merging mode can be applied to any inter-frame prediction CU, not just skip mode. An alternative to the merging mode is explicit transmission of motion parameters, where the motion vectors for each reference picture list, the corresponding reference picture indexes, and reference picture list usage flags, along with other necessary information, are explicitly sent to each CU.

[0051] In addition to the inter-frame encoding and decoding capabilities in HEVC, VVC includes several new and improved inter-frame predictive encoding and decoding tools, listed below: Extended merge forecast Merging Mode with MVD (Motion Vector Difference) (MMVD) Symmetrical MVD (SMVD) signal Affine Motion Compensation Prediction Subblock-based temporal motion vector prediction (SbTMVP) Adaptive motion vector resolution (AMVR) Sports field storage: 1 / 16 th Luminance sample MV storage and 8x8 motion field compression Bi-prediction with CU-level weight (BCW) Bidirectional optical flow (BDOF) Decoder-side motion vector refinement (DMVR) Geometric partitioning mode (GPM) Combined inter and intra prediction (CIIP) The following text provides detailed or concise information on some inter-frame prediction methods.

[0052] Extended merge forecast In VVC, the list of merge candidates is constructed by including the following five types of candidates in sequence: Space MVP from Space Neighbor CU Time MVP from the same CU Historical MVPs come from the FIFO table Pair average MVP Zero motion vectors (Zero MVs).

[0053] Space candidate derivatives The derivation of spatial merge candidates in VVC is the same as in HEVC, except that the positions of the first two merge candidates are swapped. Figure 5 In the positions shown, select up to four merge candidates for the current CU 510 (B 0, A 0, B1 and A1). The order of derivation is B. 0, A 0, B 1, A1 and B2. Position B2 is considered only when one or more neighboring CU positions B0, A0, B1, and A1 are unavailable (e.g., belonging to another slice or tile) or are intra-frame encoded. After a candidate for position A1 is added, the addition of remaining candidates is subject to redundancy checks to ensure that candidates with the same motion information are excluded from the list, thereby improving encoding and decoding efficiency.

[0054] In addition to the aforementioned spatial candidates, the non-adjacent spatial merging candidate in JVET-L0399 (Yu Han et al., “CE4.4.6: Improvement to the merge / skip mode,” ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 of the Joint Video Exploration Team (JVET), 12th meeting: Macau, China, 3-12 October 2018, document: JVET-L0399) is inserted after the TMVP in the regular merging candidate list. Figure 6 This shows an example of a pattern for space merging candidates. The distance between a non-adjacent space candidate and the current decode block is based on the width and height of the current decode block. No line buffer constraints are applied.

[0055] Time candidate derivative In this step, only one candidate is added to the list. Specifically, when deriving this temporal merging candidate for the current CU (Curr_CU) 710, a scaled motion vector is derived based on the co-located CU (Col_CU) 720 belonging to the co-located reference image, such as... Figure 7 As shown. The list of reference images and reference indices used to derive the corresponding CUs are explicitly sent in the slice header. The scaled motion vectors of the time merging candidates are shown in Figure 730. Figure 7 As shown by the dashed lines, the motion vector 740 from the co-located CU is scaled using the POC (Picture Order Count) distances tb and td, where tb is defined as the POC difference between the reference image and the current image, and td is defined as the POC difference between the reference image and the co-located image. The reference image index of the temporal merge candidate is set to zero.

[0056] like Figure 8 As shown, the position of the time candidate is selected between candidate C0 and C1. If the CU at position C0 is unavailable, is intra-coded, or is outside the current CTU line, then position C1 is used. Otherwise, position C0 is used when deriving time merging candidates.

[0057] Derivation of historical merger candidates Historically based MVP (HMVP) merge candidates are added to the merge list after the spatial MVP and TMVP. In this approach, motion information from previously coded blocks is stored in a table and used as the MVP for the current CU. A table containing multiple HMVP candidates is maintained during encoding / decoding. The table is reset (cleared) when a new CTU row is encountered. Whenever a non-sub-block inter-coded CU is encountered, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0058] Pair average merge candidate derivation Pairwise averaging candidates are generated by averaging predefined candidate pairs from an existing list of merged candidates, using the first two merged candidates. The first merged candidate is defined as p0Cand, and the second merged candidate can be defined as p1Cand. The average motion vector is calculated for each reference list based on the availability of motion vectors for p0Cand and p1Cand. If both motion vectors are available in a list, even if they point to different reference images, the two motion vectors are averaged, with the reference image set to the one for p0Cand; if only one motion vector is available, that motion vector is used directly; if no motion vector is available, the list remains invalid. Furthermore, if the half-pixel interpolation filter indices for p0Cand and p1Cand are different, they are set to 0.

[0059] If the merge list is not full after adding the average number of merge candidates, a zero MVP will be inserted at the end until the maximum number of merge candidates is encountered.

[0060] Merging estimation areas Merge Estimation Region (MER) allows for the independent derivation of merge candidate lists for CUs within the same MER. Candidate blocks located within the same MER as the current CU are not included in the generation of the current CU's merge candidate list. Furthermore, the update process for predicting the candidate list based on historical motion vectors is only updated when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2ParMrgLevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left luminance sample position of the current CU in the image, and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder end and marked as log2_parallel_merge_level_minus2 in the sequence parameter set.

[0061] Various solutions have been disclosed to improve encoding / decoding performance or reduce the complexity of cross-component prediction.

[0062] Cross-component information is used to improve prediction accuracy for non-intra-frame blocks, such as an inter-frame block. In one example of improving prediction accuracy for the chroma component of an inter-frame block, luma information from the corresponding luma component and / or chroma information from a previously encoded chroma component are used.

[0063] The first approach is to improve the prediction of Cb and / or Cr by using information from Y for a codec unit (under single-tree splitting) that contains luminance (Y) and chrominance (Cb and / or Cr) components.

[0064] The second approach involves improving the prediction of Cr by using information from Cb, for an encoding / decoding unit (under a single-tree split) containing both luma (Y) and chroma (Cb and / or Cr) components, or for an encoding / decoding unit (under a chroma two-tree split) containing only chroma (Cb and / or Cr) components. For example, model parameters are derived by using neighboring reconstructed samples of Cb and Cr as input X (source term) and Y (target term). The derived model parameters and Cb reconstructed samples are then used to generate Cr predictions.

[0065] Next, several embodiments related to the first approach are proposed to utilize the inherited cross-component mode of the current chroma block. The methods include: a) establishing a candidate list for the current block, where the candidate list includes cross-component models; b) selecting one or more model information from the list; and / or c) using the model information (similar to the intra-chroma cross-component mode) to generate one or more hypothetical predictions for the current chroma component (Cb or Cr) by applying and / or modifying the selected model information to the reconstructed or predicted samples of the corresponding luma component. When the selected model information refers to a conventional cross-component linear model, the proposed method is called the inter-cross-component linear model (inter CCLM) mode. When the selected model information refers to a convolutional cross-component model derived through a regression-based method (e.g., CCCM), the proposed method is called the inter-cross-component convolutional model (inter CCCM) mode. Furthermore, in some embodiments, a self-derived (re-derived) cross-component mode is proposed that can be added to the candidate list in Section I. In some embodiments, the choice to use the proposed inheritance pattern, such as a model inherited from a previous block, and / or the proposed self-derived pattern, such as a model inferred from the current block, is determined according to explicit rules, implicit rules, or both. Section IV describes further details.

[0066] In one embodiment, the proposed embodiment can also be used in a second scheme by using a previously encoded chroma component (Cb) as the luminance component in the first scheme.

[0067] Model storage of the current block In another embodiment, when the current non-intra-frame block, such as an inter-frame block, uses model parameters from a self-derived cross-component mode, the used model parameters can be saved and / or referenced by subsequent codec blocks. For example, for an example where the self-derived cross-component mode is CCRM, all or any subset of the model parameters can be saved. If the subsequent codec block is intra-frame, the saved model parameters are allowed to be used. If the subsequent codec block is inter-frame or any other mode type (e.g., IBC), the saved model parameters are allowed to be used. In one sub-embodiment, if the subsequent block and the reference block containing the saved model parameters belong to different mode types (e.g., one intra-frame and one inter-frame), the buffer storing the model parameters can be different.

[0068] In another embodiment, when the current non-intra-frame block, such as an inter-frame block, uses an inherited cross-component mode, the model parameters used can be stored and / or referenced by subsequent codec blocks. For example, in an inherited CCCM, all or any subset of model parameters can be stored. If the subsequent codec block is intra-frame, the stored model parameters are allowed to be used. If the subsequent codec block is inter-frame or any other mode type (e.g., IBC), the stored model parameters are allowed to be used. In one sub-implementation, if the subsequent block and the reference block containing the stored model parameters belong to different mode types (e.g., one intra-frame and one inter-frame), the buffer storing the model parameters can be different.

[0069] I. Create a candidate list that includes cross-component models. In one embodiment, when a merged candidate model list (modelList) is created, one or more of the following candidate model information are included.

[0070] Spatial model information from spatially neighboring blocks (corresponding to "spatial MVP from spatially neighboring CUs" between frames) Temporal model information from co-located blocks (corresponding to the "temporal MVP from co-located CUs" between frames) Historical model information from the FIFO table (corresponding to "historical MVP from the FIFO table" between frames) Pairwise average model information (corresponding to the "pairwise average MVP" between frames) Default model information (corresponding to "zero MV" between frames) In one sub-implementation where the candidate type is "spatial model information from spatially neighboring blocks," a valid spatially neighboring block can come from one of its spatially adjacent and / or non-adjacent neighbors (or blocks from any subset of the current block's neighbor search area) and satisfy a predefined condition. For example, the predefined condition is that the neighbor is encoded by a cross-component mode (such as CCLM, MMLM, CCCM, GLM, a mode inheriting mode information from a similarly merged candidate list, MH CCLM (where multiple cross-component models or multiple cross-component predictions are used to generate MH CCLM blocks), and / or any cross-component mode whose syntax does not belong to a traditional (non-cross-component) intra-prediction mode) or a combination of cross-component modes (such as chroma fusion (or LM-assisted Angular / Planar mode, where existing prediction hypotheses are fused with additional cross-component prediction hypotheses to generate chroma fusion blocks), inter-frame CCLM, and / or any traditional mode whose syntax does not belong to a cross-component mode but uses cross-component information to generate predictions). When scanning spatially neighboring blocks, if a candidate is valid, it is added to the list.

[0071] In another sub-implementation of temporal model information from co-op blocks, co-op blocks are derived from blocks in a reference image or co-op image as inter-frame patterns. For example, when the current block is encoded by an inter-frame prediction pattern, the co-op block is derived using or referencing the motion information of the current block (including motion vectors and / or a reference image). If the current block is a sub-block motion pattern (e.g., an affine pattern), each sub-block in the current block has its own co-op temporal model information, and / or all or any subset of co-op temporal model information derived using or referencing the motion of different sub-blocks is added to the list. Another example is that the temporal model information may come from co-op blocks derived using or referencing the motion information of neighboring blocks of the current block. If the proposed method is applied to IBC blocks or any pattern that uses block vectors, the block vector information is used as motion vectors, where the block vector information is determined by signal and / or template matching within a predefined search range and / or any implicit or explicit predefined rules.

[0072] In another sub-implementation of history-based model information, a history-based table (FIFO table) is established to store model information from previously encoded blocks. This table can be reset at the beginning and / or end of CTUs (e.g., each CTU or CTU row), slices, pictures, tiles, and / or sequences. One or more history-based candidates can be added to the candidate list in either head-to-tail or tail-to-head order.

[0073] In another sub-implementation of pairwise averaging of model information, the candidate's model information is derived based on the model information of previous candidates in the list. For example, it can average and / or modify the model parameters of more than one candidate as the model parameters to be applied. Another example is that it can combine more than one prediction as the final prediction, where each of the more than one prediction is generated by applying a model from the list.

[0074] In another sub-implementation, if the list is not full after inserting all predefined candidates, default model information is added. Some examples of default CCLM model information are shown below: For example, the preset alpha (or called The scaling parameter (a, or scaling parameter) is {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, …}, while beta (or beta...) is... (b, or offset parameter) is based on the selected default alpha, average neighbor reconstructed luminance sample value, and / or average neighbor reconstructed chromaticity (Cb / Cr) sample value.

[0075] In another sub-implementation, details of the candidate list can be found in Section V.

[0076] In another sub-implementation, when a candidate is selected from a list and its model information is used as the current block, or when model information is inherited from a previously encoded block (when a candidate is added to the list), only a portion of the model information is inherited. For example, only alpha is inherited. Beta is obtained for the current block using the inherited alpha, average neighbor reconstructed luminance sample values, and / or average neighbor reconstructed chromaticity (Cb / Cr) sample values. For example, when inheriting MMLM model information, scaling parameters and / or classification thresholds are inherited. Offset parameters in each category are derived from the inherited classification threshold and average neighbor reconstructed luminance sample values, and / or average neighbor reconstructed chromaticity (Cb / Cr) sample values ​​in each category. If no neighbor reconstructed samples are available in a category, the offset parameters are inherited directly from the candidate. For example, when inheriting CCCM model information, all convolution parameters, offsets, and / or classification thresholds are inherited. For example, when inheriting GLM model information, if the GLM candidate is a 3-parameter GLM mode, all gradient mode exponents and model parameters are inherited; otherwise, if the GLM candidate is a 2-parameter GLM mode, the offset parameters are derived using the inherited scaling parameters, the average neighbor reconstructed luminance sample values, and / or the average neighbor reconstructed chromaticity (Cb / Cr) sample values. For example, when inheriting chromaticity fusion model information, the derived MMLM parameters are inherited and used as the current block's inherited MMLM candidate.

[0077] In another sub-implementation, when a candidate is selected from the list and its model information is used as the current block, all model information is inherited. For example, alpha and beta are both inherited.

[0078] In another embodiment, if the current block is a large block whose width, height, or area exceeds a predefined threshold, the current block is divided into multiple sub-blocks. For example, the segmentation rule predefines a minimum block size, and the current block is segmented until the width or height of the sub-blocks reaches the minimum block size. In another example, the segmentation rule follows a quadtree (4 sub-blocks) or a binary tree (2 sub-blocks) partitioning. In one sub-embodiment, each sub-block will have its own list. In another embodiment, an implicit rule is defined to select a model (from the list) for each sub-block. For example, the implicit rule is to use spatial model information for sub-blocks near the top or left boundary of the current block, and / or use temporal model information for sub-blocks far from the top or left boundary (e.g., sub-blocks located in the lower right portion of the current block). If the current block is divided into 4 sub-blocks by a quadtree, the top-left sub-block uses a spatial candidate model from B2, the top-right sub-block uses a spatial candidate model from B1 or B0, the bottom-left sub-block uses a spatial candidate model from A1 or A0, and / or the bottom-right sub-block uses a temporal candidate model. Subblocks without any significant brightness residuals and / or CBF are skipped.

[0079] In another embodiment, when building modelList, one or more self-derived cross-component candidates are included. In one sub-embodiment, an example of a self-derived cross-component candidate is CCRM. The cross-component prediction of the current block (containing the target prediction sample) is formed by combining one or more proposed source terms and the model (refer to the proposed weight settings). As shown in Equation (3), pred(i, j) is the target (predicted) sample in the current block, which can be obtained after our proposed mechanism, sourceTermSet0 includes one or more source terms from the luma component, sourceTermSet1 includes one or more source terms from the chroma component, and biasTermSet includes one or more bias terms.

[0080] Equation (3) is just one example; our proposed mechanism can use any subset or extension of sourceTermSet0, sourceTermSet1, and biasTermSet. Each sample in the current block, or any subset thereof, obtains its target (predicted) sample according to Equation (3): pred(i, j) = (sourceTermSet0(i, j) + sourceTermSet1(i, j) + … +biasTermSet) (3) The proposed weight setting is used, where (i, j) is the sample position in the current block.

[0081] In the following sections, the contents of sourceTermSet0 are described in Section I.1, the contents of sourceTermSet1 in Section I.2, the contents of biasTermSet in Section I.3, and the predictor derivation using the proposed source terms and weight settings is described in Section I.4. Several examples using our proposed mechanism are shown in Section I.4.

[0082] I.1. Contents of sourceTermSet0(i, j) SourceTermSet0(i, j) includes one or more terms denoted as sourceTerm00, sourceTerm01, ..., and / or sourceTerm0 n-1 The brightness source items. The value of n represents the number of points in the source item set. In another embodiment, the pattern of n points refers to the pattern defined around / containing the location (i L ,j L The pattern of any subset of the M x N window region. If the target sample is chroma (e.g., cb or cr), (i L ,j L ) is the same brightness position from (i, j).

[0083] For a source item in the source item set, the following examples are used to determine the generation of source content.

[0084] In one embodiment, the source content is based on a prediction sample generated from a prediction model and / or a reconstruction sample generated from the prediction model and reconstruction residuals.

[0085] In another sub-implementation, the source content is a filtered source or a source that has undergone any preprocessing. For example, the source content is a predicted / reconstructed sample filtered by a predefined model or filter.

[0086] In another sub-implementation, the source content is gradient information from the predicted samples and / or reconstructed samples.

[0087] In another sub-implementation, since the target sample is a chroma sample (e.g., Cb or Cr), the predicted and / or reconstructed samples are located within the co-located (luminance) block from the current (chroma) block. The predicted and / or reconstructed samples are considered as initial samples and used as source content for generating the target sample.

[0088] In another embodiment, the value of the source item is further adjusted by a predetermined offset (e.g., added to or subtracted from).

[0089] In another embodiment, the source item may also include location information.

[0090] I.2. Contents of sourceTermSet1(i, j) SourceTermSet1(i, j) includes one or more chromaticity (Cb or Cr) source terms, denoted as sourceTerm00, sourceTerm01, ... and / or sourceTerm0 m-1 The value of m represents the number of taps in the source itemset. In one embodiment, the source items can be linear and / or nonlinear, linear only, and / or nonlinear only. In another embodiment, the pattern of m taps refers to the number of taps defined around / containing positions (i... C j C The pattern of any subset of the window region M2 x N2. If the target sample is chromaticity (Cb or Cr), (i C j C ) is (i, j).

[0091] For source items in the source item set, the following examples are used to determine the generation of source content.

[0092] In one embodiment, the source content is based on a prediction sample generated from a prediction model and / or a reconstruction sample generated from the prediction model and reconstruction residual based on the prediction sample.

[0093] In another sub-implementation, the source content is a filtered source or a source that has undergone any preprocessing. For example, the source content is a predicted / reconstructed sample filtered by a predetermined model or filter.

[0094] In another sub-implementation, the source content is gradient information from the predicted samples and / or reconstructed samples.

[0095] In another sub-implementation, if the target sample is a chroma sample, the predicted and / or reconstructed samples are located within the current block. The predicted and / or reconstructed samples are considered as initial samples and used as source content for generating the target sample.

[0096] In another embodiment, the value of the source item is further adjusted by a predetermined offset (e.g., added to or subtracted from).

[0097] In another embodiment, the source term may also include location information. For example, if the target sample refers to chroma, then the horizontal position (i) of (i, j) is used for the source term, while the vertical position (j) of (i, j) is used for the source term.

[0098] I.3. Contents of biasTermSet The deviation term is a predetermined value. In one embodiment, the deviation term is the midValue of bitDepth specified in the standard. For example, the deviation term is set to... In another embodiment, the bias term is the same for every sample in the current block. That is, the bias term does not take position (i, j) into account.

[0099] I.4. Predictor Derivation for Sample (i, j) I.4.1. Suggested Weighting The proposed weighting is achieved by estimating the relationship (minimizing distortion) between "predicted and / or reconstructed samples on the reference region of the current (chroma) block" and "predicted and / or reconstructed samples on the reference region of the corresponding luma block" using a predetermined regression method to generate weights (referring to model parameters) based on the regression method. The weights derived from the source terms are then applied to obtain the target (predicted) samples in the current block. In one embodiment, the predetermined regression method may be the Linear Least Mean Squared Error (LMMSE) method for CCLM, or any method unified with the regression method used for CCLM. In another embodiment, the predetermined regression method may be the LDL decomposition method for CCCM, or any method unified with the regression method used for CCCM. In yet another embodiment, the predetermined regression method may be Gaussian elimination.

[0100] In one embodiment, the reference region of the current block is the spatial proximity region of the current block. The spatial proximity region of the current block (CurrentCU) 910 includes the upper reference region 912, the left reference region 914, the upper left reference region 916, and / or any subset thereof, such as... Figure 9 As shown.

[0101] The reference area for a corresponding luminance block is the spatially adjacent area of ​​that luminance block.

[0102] In another embodiment, the reference region of the current (chroma) block is the vector co-location region of the current block, and the reference region of the corresponding luma block can be the luma block co-located with the current chroma block, or the vector co-location region of the corresponding luma block. For an interactive codec unit containing luma and chroma blocks, the vector co-location region of the current block refers to the motion compensation result using the motion information (motion vector and / or reference image) of the current block, while the vector co-location region of the corresponding luma block refers to the motion compensation result using the motion information (motion vector and / or reference image) of the corresponding luma block. For IBC or intraTMP, the vector co-location region of the current block refers to the motion compensation result using the motion information (block vector and / or current image) of the current block, while the vector co-location region of the corresponding luma block refers to the motion compensation result using the motion information (block vector and / or current image) of the corresponding luma block.

[0103] In another embodiment, the two reference regions of the current block described above can be used together. For example, typically, when deriving model parameters, samples in the vector co-location region of the current block are used as input samples; however, for smaller blocks, when deriving model parameters, samples in the spatially adjacent reference region are used as additional input samples.

[0104] In another embodiment, more details on constructing the modelList can be found in Section 6.

[0105] II. Signals for model information control.

[0106] When the proposed inter-frame CCLM (or inter-frame CCCM) is not applied, the prediction for the current block comes from the original inter-frame prediction.

[0107] In another embodiment, whether inter-frame CCLM is applied depends on the signal.

[0108] In one sub-implementation, the signal refers to the flags at the encoded TU and / or TB and / or CU and / or CB levels. In another implementation, inter-frame CCLM (or inter-frame CCCM) is supported only if the size conditions of the current block are met.

[0109] In one sub-implementation, the size condition is that the block width, block height, or block area is greater than a predefined threshold. The predefined threshold can be a positive integer, such as 8, 16, 32, 64, 128, 256, ...

[0110] In another sub-implementation, the size condition is that the block width, block height, or block area is less than a predefined threshold. The predefined threshold can be a positive integer, such as 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096, etc.

[0111] In another embodiment, the original inter-frame prediction (generated by motion compensation) is used for luma, while the prediction for chroma components is generated by CCLM and / or any other cross-component model, such as a model from another LM mode.

[0112] In one sub-implementation, the current CU is considered to be an inter-frame CU, an intra-frame CU, or a novel prediction mode (neither intra-frame nor inter-frame).

[0113] In another embodiment, as a further proposed method related to Section V, one or more LM modes (or cross-component modes) predicted by one or more assumptions for generating LM-assisted angular / planar modes and / or inter-frame CCLM and / or MH CCLM are selected from a predefined merge candidate list (called modelList). A modelIdx is signaled to select a candidate from the candidate list (modelList), and the selected candidate is used for the current block. modelList contains one or more candidates, each referring to a model (or cross-component mode) information. If there is only one candidate in the list (the list size is only 1), modelIdx is not signaled, and / or modelIdx can be inferred to be 0 or a default value.

[0114] In one embodiment, when building the modelList, one or more predefined candidates are added. The predefined candidates can include any subset / extension of the following candidates: CCLM family: CCLM_LT, CCLM_L, CCLM_T MMLM family: MMLM_LT, MMLM_L, MMLM_T CCCM family: CCCM_LT, CCCM_L, CCCM_T The method described above can also be applied to IBC blocks or blocks of any IBC submode (e.g., IBC merging or IBCAMVP (or IBC Advanced MVP or IBC Inter-Frame) or any IBC mode under the IBC syntax). (In this invention, "inter-frame" can be replaced with IBC.) That is, for chroma components, block vector prediction can be combined with or replaced by cross-component prediction.

[0115] III. Generating Hypothesis Predictions III.1. Concept In one embodiment, a prediction- or reconstruction-based model is used to generate a hypothetical prediction of the current chromaticity component.

[0116] In a sub-implementation of a prediction-based linear model, the derived model parameters are applied to the predicted samples of the first component (Y) to obtain the predicted samples of the second or third component: .

[0117] The predicted samples of the first component are downsampled by downsampling filters (these filters may be fixed to a predefined filter or selected from some candidate filters).

[0118] In a sub-implementation of a reconstruction-based linear model, the derived model parameters are applied to the reconstructed samples of the first component (Y) to obtain predicted samples of the second or third component: .

[0119] The reconstructed samples of the first component are downsampled by downsampling filters (these filters may be fixed to a predefined filter or selected from some candidate filters).

[0120] Convolutional models based on prediction or reconstruction are similar to those proposed for linear models based on prediction or reconstruction. The main difference is that the model coefficient pattern follows CCCM (instead of CCLM), and luminance samples may or may not be downsampled first. If luminance samples are not downsampled, more taps (model coefficients) may be used to access the undownsampled luminance samples.

[0121] III.2. CCLM of Inter-Frame Blocks The CCLM of an inter-frame block can also be called an inter-frame CCLM, and "CCLM" can be extended to any LM mode (or any cross-component mode) or replaced by any LM mode (or any cross-component mode).

[0122] In one embodiment, for the chroma component, in addition to the original inter-frame prediction (generated by motion compensation, which may be a single prediction and / or bidirectional prediction, multiple hypothetical predictions of multiple motion candidates, which may refer to one or more merged candidates and / or one or more AMVP candidates, and / or any combination of the above, or may be only a single prediction), one or more hypothetical predictions (generated by CCLM and / or any other LM mode) are used to output the current prediction.

[0123] In one sub-implementation, the current prediction is a weighted sum of inter-frame prediction and CCLM prediction.

[0124] In another embodiment, inter-frame prediction can be generated by any of the inter-frame modes described above. For example, the inter-frame mode can be a regular merging mode. Another example is the CIIP mode. Yet another example is GPM or any GPM variant (e.g., GPM intra-references a prediction unit using intra-frame prediction).

[0125] In another embodiment, inter-frame CCLM is supported only when any (or more) of the predefined inter-frame modes are used as the current block, or when any (or more) of the enable flags of the predefined inter-frame modes are indicated as enabled. Supporting inter-frame CCLM means that the prediction of the current block can be selected between applying inter-frame CCLM and not applying inter-frame CCLM.

[0126] For example, if a CCLM mode is used to generate chroma prediction samples, and the luma prediction comes from inter-frame encoding / decoding tools, a flag is used to indicate whether the CCLM model used for chroma prediction is inherited from a CCLM model used in a previous coded block or from a predefined CCLM mode. If the CCLM model is inherited from a CCLM model used in a previous coded block, an index is used to indicate which model in the list was inherited or modified. Otherwise, a predefined CCLM mode is used to implicitly derive the CCLM model for the current chroma prediction.

[0127] IV. Choice of using the suggested inheritance pattern and / or self-derivative pattern In one embodiment, a flag can be sent to indicate / select whether to use a re-derived model. If the flag is 0, the cross-component model used for encoding and decoding neighbor merging candidate blocks is inherited. If the flag is 1, the re-derived approach is used.

[0128] In another embodiment, an implicit rule (without using additional flags) is used to determine whether to use the re-derived model.

[0129] In another embodiment, the proposed method, such as inheritance or self-derived methods, is used to implicitly select candidates for generating cross-component predictions by choosing the candidate with the lowest cost or model error (e.g., the first candidate in modelList). As another example, an index is sent to select one or more candidates from modelList. More details can be found in Section II.

[0130] V. Details of cross-component model information in the candidate list V.1. Inheriting CCM Information In one embodiment, inherited cross-component model (CCM) information can be stored along with inherited model parameters. CCM information can be inherited along with the inherited model parameters. Predictions for the current block can be generated based on the inherited CCM information and the inherited model parameters. CCM information may include, but is not limited to, prediction patterns (e.g., CCLM, MMLM, CCCM, 2-parameter GLM, 3-parameter GLM), model indices indicating which model shape is used in the convolutional model, classification thresholds for multiple models, information indicating the use of non-downsampled samples in the convolutional model, downsampled filter flags, downsampled filter indices when using multiple downsampled filters, the number of neighboring rows for the derived model, the template type for the derived model, post-filter flags, and model parameters.

[0131] In another embodiment, a CCLM model can be inherited. In addition to storing model parameters, a prediction pattern can also be stored in the CCM information to indicate that the inherited model is a CCLM model.

[0132] In another embodiment, a CCLM model with nonlinear terms can be inherited. In addition to storing model parameters, a prediction pattern can also be stored in the CCM information to indicate that the inherited model is a CCLM model with at least one nonlinear term.

[0133] In another embodiment, a CCCM model can be inherited. In addition to storing model parameters, a prediction mode indicating that the inherited model is a CCCM model can also be stored in the CCM information. Luminance and chromaticity offsets used to adjust the CCCM model inputs can also be stored in the CCM information.

[0134] In another embodiment, CCCM models with different convolutional filter shapes can be inherited. In addition to model parameters and prediction modes, a CCCM mode index can be stored in the CCM information to indicate which convolutional filter shape the inherited CCCM model uses. For example, a CCCM model with different convolutional filter shapes might only contain horizontal spatial terms. Another example is that a CCCM model with different convolutional filter shapes might only contain vertical spatial terms. Yet another example is that a CCCM model with different convolutional filter shapes might only contain diagonal spatial terms. Yet another example is that a CCCM model with different convolutional filter shapes might only contain anti-diagonal spatial terms. Yet another example is that a CCCM model with different convolutional filter shapes might contain X-shaped spatial terms.

[0135] In another embodiment, a CCCM model that uses unsampled samples can be inherited. In addition to storing model parameters, a prediction pattern can also be stored in the CCM information to indicate that the inherited model is a CCCM model that uses unsampled samples.

[0136] In another embodiment, a CCCM model with multiple downsampling filters can be inherited. In addition to storing model parameters, a prediction mode can be stored in the CCM information to indicate that the inherited model is a CCCM model with multiple downsampling filters, and a model index can also be stored in the CCM information to indicate which variant of the inherited CCCM model with multiple downsampling filters.

[0137] In one embodiment, a mixed convolutional cross-component model (CCCM model) composed of various terms (e.g., spatial, gradient, positional, nonlinear, and bias terms) can be inherited. The gradient term can be computed in either a downsampled or non-downsampled domain. The positional term can be computed relative to the top-left coordinates of the current block or the image. In addition to storing model parameters, a prediction pattern can be stored in the CCM information to indicate that the inherited model is a mixed CCCM model composed of various terms. If there are multiple types of mixed CCCM models, a model index can also be stored in the CCM information to indicate which type of mixed CCCM model is inherited. For example, gradient- and position-based CCCM (GL-CCCM) in JVET-AB0119 (Ramin G. Youvalari et al., “Non-EE2: Gradient- and position-based convolutional cross-component model (GL-CCCM) for intra-frame prediction,” Joint Film Exploration Group (JVET) ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 28th meeting, Mainz, DE, October 20-28, 2022, document: JVET-AB0119) is a hybrid CCCM model that includes a spatial term for the center position, two gradient terms for the horizontal and vertical directions, two positional terms X and Y relative to the horizontal and vertical positions, a nonlinear term, and a bias term. A prediction pattern can be stored in the CCM information to indicate that the inherited model is a GL-CCCM model.

[0138] In another embodiment, the GLM model can be inherited. In addition to storing model parameters, a prediction mode can be stored in the CCM information to indicate that the inherited model is a GLM model, and a downsampling filter index can also be stored in the CCM information to indicate which gradient downsampling filter was used in the inherited GLM model.

[0139] In another embodiment, a GLM model with a luminance term can be inherited. In addition to storing model parameters, a prediction mode can be stored in the CCM information to indicate that the inherited model is a GLM model with a luminance term, and a downsampling filter index can also be stored in the CCM information to indicate which gradient downsampling filter was used in the inherited GLM model with a luminance term.

[0140] In another embodiment, any type of cross-component multi-model can be inherited. In addition to storing model parameters and prediction modes, a multi-model on / off flag can be stored in the CCM information to indicate whether the inherited CCM model is a multi-model. If the multi-model on / off flag is true, the multi-model classification threshold is also stored in the CCM information.

[0141] In another embodiment, the CCM information may include information indicating how the inherited model is derived. For example, the CCM information may include the number of adjacent rows used to derive the cross-component model and / or the template type used to derive the model. For instance, a set of templates may be used to derive the CCCM model. This set of templates includes templates with different positions, sizes, and shapes. The CCM information may store the template index of which template the inherited CCCM model is derived from. For example, the inherited CCCM model may be derived based on a top-only template, a left-only template, or both left and top templates. As another example, the inherited CCCM model may be derived based on a 6-row template or a 2-row template.

[0142] In another embodiment, a post-filter flag can be stored in the CCM information. This information describes how the inherited model is used in blocks derived from that model. If the post-filter flag is enabled, this indicates that a filter is applied to the predictions of blocks derived from that inherited model.

[0143] V.2. Refinement of Inherited Model Parameters In one embodiment, the inherited model parameters can be further refined based on the inherited CCM information. The inherited CCM information may include how the inherited model is derived, such as the template type and / or the number of neighboring rows used to derive the model. The refined parameters are derived based on local information. The refinement process may follow the derivation method of the inherited model and use the same type of template and / or the same number of neighboring rows. For example, if the inherited model is CCLM and is derived based on a left-only template (e.g., the inherited model is CCLM_L), the offset parameters... The mean of the samples can be reconstructed from the templates adjacent to the left of the current block. For example, if the inherited model is CCLM and it is derived based only on the top template (the inherited model is CCLM_T), the offset parameter... The mean of the samples can be reconstructed from the template of the neighboring top of the current block. For another example, if the inherited model is CCCM and derived with a 2-row template, the offset value (e.g., in CCCM) can be used. (where c6 is the weight coefficient of the bias term) can be re-derived based on the 2-row template of the current block. For another example, if the inherited model is a multi-model (MMLM, CCCM with multi-model) and is derived based on only the left template, the classification threshold can be re-derived based on the reconstructed sample from the left template of the current block.

[0144] In another embodiment, the inherited model parameters are further refined using different types of templates and / or different numbers of rows, and the final model parameters are determined by the template cost. The template cost is calculated by applying candidate refined model parameters to neighboring templates to predict template samples and computing the difference between the predictions and reconstructed samples (SAD or SATD). For example, if the inherited model is CCLM, refined offset parameters are derived using the left, top, and top-left templates of the reconstructed samples of the current block, respectively. If the application The template cost is Choose the minimum value among them. As the final offset parameter.

[0145] In another embodiment, inherited model parameters can be further refined using predefined values. Template cost is used to determine whether to further refine the inherited model parameters. Template cost is calculated by applying candidate refined model parameters to neighboring templates to predict template samples and computing the difference between the predicted and reconstructed samples (SAD or SATD). For example, for the CCCM mode, for each inherited model parameter... This value is obtained through Refine and compare applications. and The template cost is used to determine which is the final model parameter value.

[0146] V.3. Inheritance Spatial Proximity Model Parameters In one embodiment, inherited model parameters can come from an immediately neighboring block. Models from blocks at predefined locations are added to the candidate list in a predefined order.

[0147] In one embodiment, the predefined positions and predefined order can be the same as the inter-frame merging mode of the spatial candidates.

[0148] In one embodiment, assuming the current block's position, width, and height are (x, y), W, and H, respectively, the predefined position can include the position directly above the current block, such as (x + W>>1, y-1) or (x + (W+1)>>1, y-1), if W is greater than or equal to a threshold TH. The predefined position can also include the position to the left of the current block, such as (x-1, y+H>>1) or (x-1, y+(H+1)>>1), if H is greater than or equal to the threshold TH. TH can be 2, 4, 8, 16, 32, or 64.

[0149] In one embodiment, the maximum number of models inherited from spatial neighbors can be added to the candidate list, and the maximum number is less than the number of predefined locations.

[0150] V.4. Inheritance Temporal Proximity Model Parameters In one embodiment, if the current slice / image is a non-intra-frame slice / image, the inherited model parameters can come from blocks in previously encoded / decoded slices / images.

[0151] In one embodiment, the current block is located at (x, y), and the block size is... Inherited model parameters can come from blocks at some predefined locations in previously encoded / decoded slices / images.

[0152] In one sub-implementation, the predefined location can be or ,in Two value sets and Defined as: , .

[0153] and All values ​​in the range are positive.

[0154] For example, It can be .

[0155] For example, It can be ,in and They are two fixed positive numbers.

[0156] To give another example, ,For example .

[0157] To give another example, ,For example and .

[0158] In one sub-implementation, a predefined location Within the corresponding region of the current encoding / decoding block, i.e. and The predefined location can be .

[0159] In one sub-implementation, a predefined location Outside the corresponding region of the current encoding / decoding block, i.e. and The predefined location can be .

[0160] In one embodiment, from closer The location model was first added to the final merge candidate list.

[0161] The previously encoded and decoded image obtains the inherited parameter model, which is then called the parsed image.

[0162] In one embodiment, the previously encoded / decoded image, i.e., the parsed image, is one of the images in the reference list.

[0163] In one embodiment, the corresponding image is marked in the image / slice header. A reference list and reference index are also marked in the image / slice header. For example, the corresponding image is selected as L0[0]. As another example, the corresponding image is selected as L1[0].

[0164] In one embodiment, such as Figure 10 As shown, the current block position is (x, y), and the block size is [size missing]. The inherited model parameters can come from blocks at positions (x', y'), (x', y' + h / 2), (x' + w / 2, y'), (x' + w / 2, y' + h / 2), (x' + w, y'), (x', y' + h), or (x' + w, y' + h) in previously encoded / decoded slices / images, where x' = x + Δx and y' = y + Δy.

[0165] In one sub-implementation, if the prediction mode for the current block is inter, Δx and Δy are set as the horizontal and vertical motion vectors of the current block.

[0166] In another sub-implementation, if the current block is an inter-double prediction, Δx and Δy are set to the horizontal and vertical motion vectors in the reference image list 0.

[0167] In another sub-implementation, if the current block is an inter-double prediction, Δx and Δy are set to the horizontal and vertical motion vectors in the reference image list 1.

[0168] V.5. Inheriting the proximity model of non-proximity spaces In one embodiment, the inherited model parameters can come from non-adjacent spatial neighbor blocks. Models from predetermined location blocks are added to the candidate list in a predetermined order.

[0169] In one sub-implementation, the predetermined positions and predetermined order are the same as those for non-adjacent spatial neighbor candidates used in the inter-frame merging mode.

[0170] In one sub-implementation, the predetermined position and predetermined order are as follows: Figure 11A andFigure 11B As shown. The positions of the numbered squares are predetermined positions. The numbers within each square indicate a predetermined order. Positions in pattern 1 (1110) are added to the list before those in pattern 2 (1120). The distance between each predetermined position is proportional to the width and height of the current block.

[0171] In one embodiment, there is a maximum limit on the number of models inherited from non-adjacent spatial neighbors that can be added to the candidate list, and this maximum number is less than the number of predetermined positions.

[0172] V.6 Inheriting Model Parameters from History Tables In one embodiment, the inherited model parameters can come from a cross-component model history table. The history table stores CCM information for valid previously encoded blocks. A valid previously encoded block refers to any block containing valid CCM information. Cross-component models in the history table can be added to a candidate list in a predetermined order. In one embodiment, the order in which historical candidates are added can be from the beginning to the end of the table. In another embodiment, the order in which historical candidates are added can be from the end to the beginning of the table.

[0173] In one embodiment, a cross-component model history table can be maintained to store previous cross-component models (i.e., CCM information), and the cross-component model history table can be reset at the start of the current image, current tile, current tile, every M CTU row, or every N CTU, where N and M can be any value greater than 0. In another embodiment, the cross-component model history table can be reset at the start of the current image, current tile, current tile, current CTU row, or the end of the current CTU.

[0174] In another embodiment, multiple history tables are used to store different types of cross-component models. For example, the first history table stores a single model, and the second history table stores multiple models. Another example is that the first history table stores gradient models, and the second history table stores non-gradient models. Yet another example is that the first history table stores simple linear models (e.g., y = ax + b), and the second history table stores complex models (e.g., CCCM).

[0175] In one embodiment, when adding historical candidates to a candidate list from multiple historical tables, the addition order can be from the beginning to the end of a table, and then the next historical table can be added in the same or reverse order.

[0176] V.7 Inherited from Fusion Mode Fusion mode refers to the mode of fusing two predictions to generate a final prediction. In chroma intra-frame fusion mode, a chroma intra-frame prediction generated without using cross-component prediction (CCP) codec tools (e.g., CCLM, MMLM, CCCM) is fused with another chroma intra-frame prediction generated using a cross-component prediction codec tool. For example, a non-CCLM encoded intra-frame prediction and a CCLM encoded intra-frame prediction are fused together to obtain the final intra-frame prediction.

[0177] In chroma intra-frame fusion mode, a non-CCLM encoded intra-frame prediction and a CCLM encoded intra-frame prediction are fused together to obtain the final intra-frame prediction. In one embodiment, when inheriting cross-component model parameters from blocks / positions encoded / decoded by the chroma intra-frame fusion mode, the model parameters used to obtain the CCLM encoded intra-frame prediction are inherited and further refined. In another embodiment, the fusion weights, the encoding / decoding mode of the non-CCLM encoded intra-frame prediction, and the model parameters used to obtain the CCLM encoded intra-frame prediction are inherited and further refined. In yet another embodiment, the encoding / decoding mode of the non-CCLM encoded intra-frame prediction is implicitly derived (e.g., derived as a DM or planar mode), while the fusion weights and the model parameters used to obtain the CCLM encoded intra-frame prediction are inherited and further refined. In another embodiment, if the intra-prediction of the block / position encoded by the chroma intra-fusion mode can be implicitly derived (e.g., the intra-prediction of the non-CCLM encoding / decoding is DM or planar mode), then the fusion weights and the model parameters used to obtain the intra-prediction of the CCLM encoding / decoding are inherited and further refined.

[0178] In one embodiment, when inheriting cross-component model parameters from blocks / positions encoded / decoded by chroma intra-fraction mode, the model parameters used to obtain intra-frame predictions for CCP encoding / decoding are inherited and further refined.

[0179] In one embodiment, in addition to inheriting and / or refining the CCP model parameters, the fusion weights and / or encoding / decoding modes of intra-frame prediction for non-CCP encoding / decoding are also inherited. That is, the chroma intra-frame fusion mode is inherited.

[0180] V.8 Limits buffer / storage resource requirements To limit the demand for buffer / storage resources, the available range of temporal candidates should be limited. The temporal candidates mentioned in this section refer to candidates that inherit model parameters from blocks in the previous encoded slice / image, as described in the "Inheriting Temporally Proximity Model Parameters" section. For example, assuming the current block position is (x, y), the position in the previous encoded slice / image that inherits the parameter model could be (x + Δx). i , y +Δy i), where i ranges from 1 to M, and M is a positive integer greater than 0. Δx i and Δy i It is a predefined displacement. To give another example, suppose the current block position is (x, y), the position in the previous encoded slice / image inherited from the parametric model could be (x + dx + Δx). i , y + dy + Δy i ), where i ranges from 1 to M, and M is a positive integer greater than 0. Δx i and Δy i These are predefined displacements. dx and dy are determined by the motion vectors of the current block's neighboring blocks. Details on how to determine these motion vectors are described in the "Inheriting Temporal Proximity Model Parameters" section. To give another example, assuming the current block is located at (x, y), the position in the previous encoded slice / image of the inherited parameter model could be (x + dx + Δx). i , y + dy + Δy i ), where i ranges from 1 to M, and M is a positive integer greater than 0. Δx i and Δy i It is a predefined displacement.

[0181] In one embodiment, if the prediction mode of the current block is inter, then dx and dy are set to the horizontal and vertical portions of the motion vector of the current block. If the horizontal or vertical portion of the motion vector is a fraction, then dx is set to the rounded value of the horizontal portion of the motion vector, and dy is set to the rounded value of the vertical portion of the motion vector. The rounding method used may be, but is not limited to, rounding towards negative infinity, rounding towards positive infinity, rounding towards zero, rounding to the nearest integer (e.g., rounding away from zero, rounding half up, rounding half down, etc.), or rounding to the nearest predefined precision (e.g., rounding to the nearest k-pixel or 1 / k-pixel precision position, where k can be 2, 4, 8, 16, or 32). If the prediction mode of the current block is IBC, then dx and dy are set to the horizontal and vertical block vectors of the current block. If the horizontal or vertical portion of the block vector is a fraction, then dx is set to the rounded value of the horizontal portion of the block vector, and dy is set to the rounded value of the vertical portion of the block vector. The rounding method used can be, but is not limited to, the following: rounding towards negative infinity, rounding towards positive infinity, rounding towards zero, rounding to the nearest integer (e.g., rounding away from zero, rounding half up, rounding half down, etc.), or rounding to the nearest predefined precision (e.g., rounding to the nearest k-pixel or 1 / k-pixel precision position, where k can be 2, 4, 8, 16, or 32).

[0182] In one embodiment, only the cross-component model (CCM) information of CTUs in the co-located image corresponding to the current codec CTU's position in the current image can be referenced by the time candidate. In another embodiment, only the CCM information of the CTUs in the co-located image corresponding to the current codec CTU's position, and / or the positions of the N CTUs to its left and / or the M CTUs to its right in the current image can be referenced by the time candidate, where N and M can be any integer greater than 0. In another embodiment, only the CCM information of the CTU rows in the co-located image corresponding to the current codec CTU's row position in the current image can be referenced by the time candidate. In yet another embodiment, only the current image containing the positions of the CTU rows in the co-located image corresponding to the current CTU's row, and / or the N rows of CTUs above and / or the M rows of CTUs below can be referenced, where N and M can be any integer greater than 0. Note that, as described in the sections “Inheriting CCM Information” and “Refining Inherited Model Parameters”, the CCM information mentioned in this disclosure includes, but is not limited to, prediction patterns (e.g., CCLM, MMLM, CCCM), GLM pattern indexes, model parameters, or classification thresholds.

[0183] exist Figure 12 In this context, a co-position CTU refers to a CTU within the same image, whose position corresponds to the position of the current codec CTU within the current image. Let the positions of the top-left and bottom-left corners of the co-position CTU be (x...). L ,y T ) and (x L ,y B Let the image width be w. Figure 12 The x and y ranges for each dashed line region in the diagram are defined as follows: Region 1: x L - N ≤ x <w, y T - N ≤ y <y T Region 2: x L - N ≤ x <x L , y T ≤ y ≤ y B Region 3: 0 ≤ x <x L - N, y B - N <y ≤ y B N is any positive integer, N>0.

[0184] In one embodiment, the time candidate can only reference CCM information in the co-located CTU, region 1, region 2, or region 3, such as... Figure 12 As shown. N is set to a predefined value. For example, N is set to the minimum block size allowed in the standard. This block can be CU / PU / TU. Another example is that N is set to 4. In one embodiment, the region where the time candidate can reference CCM information is the same as the region in the inter mode where the time motion vector can be referenced.

[0185] V.9 shares buffer resources with existing encoding and decoding tools. To further inherit the model, the buffer used to store CCM information (e.g., prediction mode, related submode flags, prediction mode or model parameters) and the buffer used to store inter-frame encoding / decoding information (e.g., motion vector buffers) are shared with cross-component information inheritance (or cross-component (CC) merging mode) for storing CCM information. Assume the minimum allowed block size is... The current CTU size is The current image size is The CTU-level buffer and image-level buffer are used to store the current CTU and the inter-frame encoding / decoding and CCM information for each image, respectively. A CTU-level buffer is created to store the final inter-frame encoding / decoding or CCM information; the size of this CTU-level buffer is... A picture-level buffer is created to store the final inter-frame encoding / decoding or CCM information of the current image. The size of this picture-level buffer is... ,in and After encoding and decoding the current block, the inter-frame encoding / decoding or CCM information of the current block is first... The unit is saved to the corresponding location in the CTU-level buffer, where the corresponding location is... The unit is located at the position covered by the current block. Afterwards, after encoding and decoding the current CTU, the inter-frame encoding / decoding or CCM information in the current CTU-level buffer is saved to the corresponding position in the image-level buffer, so that... The unit.

[0186] However, if the units of the CTU-level buffer and the image-level buffer are different (e.g., or This requires resampling of inter-frame encoding / decoding or CCM information in the CTU-level buffer to save it to the image-level buffer. Assuming... and Each CTU-level buffer Select a location in the grid to store inter-frame encoding / decoding or CCM information to the appropriate position in the image-level buffer. For example, ... Figure 13 As shown, if and A selected location within each 2x2 grid is used to store inter-frame encoding / decoding or CCM information to the corresponding location in the image-level buffer. In one embodiment, the selected location can be the top-left, bottom-left, top-right, or bottom-right of each 2x2 grid. Figure 13 As shown, the inter-frame encoding / decoding (CCM) information at the top-left position of the shaded area marker in each 2x2 grid is saved to the picture-level buffer. In another embodiment, when the CCM information in the CTU-level buffer is resampled for saving to the picture-level buffer, a conditional check can be performed. Prediction patterns within the grid. For example, if If more than a certain percentage of locations within the grid are intra-frame predictions (e.g., more than 50% or 75%), then the selected and saved data is CCM information. Otherwise (i.e., Most locations within the grid are for inter-frame prediction; the selected and saved data is inter-frame encoding / decoding information. When selecting candidates to save to the image-level buffer, the first allowed candidate can be selected following a predetermined scan order. For example, if the selected and saved data is CCM information, it can be selected according to a predetermined scan order. The first grid within the grid containing CCM information. Another example: if the selected and saved data is inter-frame encoding / decoding information, it can be selected according to a predetermined scan order. The first grid within the grid that contains inter-frame encoding / decoding information.

[0187] Because the buffer storing inter-frame codec information is shared with the CC merge mode, the CU prediction mode (e.g., intra-frame prediction or inter-frame prediction) can be checked to determine whether the information stored at a certain buffer location is inter-frame codec or CCM information. In one embodiment, if the CU prediction mode is intra-frame prediction, the stored information is CCM information. Otherwise (i.e., the CU prediction mode is not intra-frame prediction), the stored information is inter-frame codec information. In another embodiment, an invalid inter-frame prediction reference index or an invalid MV value (e.g., horizontal or vertical MV value) can be set to identify that the stored information is CCM information. Otherwise (i.e., a valid inter-frame prediction index), the stored information is inter-frame codec information. For example, in the VVC standard specification, an inter-frame prediction reference index greater than 2 is invalid, so the inter-frame prediction reference index can be set to a value greater than 2 to identify that the stored information is CCM information (e.g., the inter-frame prediction reference index is 3).

[0188] VI. Constructing the Candidate List In one embodiment, the candidate list is constructed by adding candidates in a predetermined order until a maximum number of candidates is reached. The added candidates may include, but are not limited to, all of the aforementioned candidates. For example, the predetermined order may be spatially adjacent candidates, temporal candidates, spatially non-adjacent candidates, historical candidates, and then default candidates.

[0189] In another embodiment, if all predetermined neighboring and historical candidates have been added but the maximum number of candidates has not been reached, some default candidates are added to the candidate list until the maximum number of candidates is reached.

[0190] In one embodiment, the preset candidate can be a Cross-Component Linear Model (CCLM). Scaling parameters From collection Where N is a positive integer. For example, the set could be {0, 1 / 8, -1 / 8, +2 / 8, -2 / 8, +3 / 8, -3 / 8, +4 / 8, -4 / 8}. Offset parameter It can be Alternatively, it can be derived from nearby luminance and chrominance samples. For example, if the average values ​​of nearby luminance and chrominance samples are lumaAvg and chromaAvg, In one sub-implementation, the order in which default candidates are included can depend on the scaling parameter. The absolute value and sign. For example, default candidates are added to the list in the following order: .

[0191] In another embodiment, a preset candidate can be an earlier candidate refined with an incremental scaling parameter. The earlier candidate is a CCLM model. If the scaling parameter of the earlier candidate is... The default scaling parameter for candidates is .For example, It can be 0, , where N is a positive integer. For example, It can be 0, Offset parameter Based on This is derived from the average of the luminance and chrominance samples of the current block's neighbors. In one sub-implementation, the earlier candidate is the first CCLM candidate added to the list. In another sub-implementation, the default candidate inclusion order may depend on refinement. The absolute value and sign. For example, default candidates are added to the list in the following order: 0, .

[0192] VII. Remove or modify similar neighboring model parameters When inheriting cross-component model parameters from other blocks, the similarity between the inherited model and existing models in the candidate list or model candidates derived from neighboring reconstruction samples of the current block (e.g., CCLM, MMLM, or CCCM models derived using neighboring reconstruction samples of the current block) can be further examined. If a candidate parameter's model is similar to an existing model, that model will not be included in the candidate list.

[0193] VIII. Reorder the candidates in the list Candidates in the list can be reordered to reduce the syntactic overhead of indexing candidates in signal selection.

[0194] In one embodiment, the reordering rule may rely on the encoding / decoding information or model error of neighboring blocks. For example, if the neighboring block above or to the left is encoded / decoded by MMLM, then MMLM candidates in the list can be moved to the head of the current list.

[0195] In one embodiment, the reordering rule is based on the model error of applying the candidate model to the neighboring templates of the current block, and then comparing the error with the reconstructed samples of the neighboring templates.

[0196] In one embodiment, the dequantized input (quantization index, qIdx) can be shifted using a derived offset. The derived offset can be predefined or selected from a predefined candidate set. After dequantizing the shifted input, a shifted dequantized result can be obtained, which can be further combined with the original unshifted dequantized result using weighted summation to form the final dequantized result. In one embodiment, the combined weighting depends on boundary matching. If the original dequantized result has a smaller boundary matching cost, its weight will be higher than that of the shifted dequantized result. In another embodiment, the derived offset is determined by boundary matching. For example, each candidate in the candidate set has a boundary matching cost, and the candidate with the lowest boundary matching cost is used as the shift offset. In another embodiment, a selection rule is used to determine which quantization index and / or which dequantized result to use. The selection rule depends on a predefined threshold. For example, if the quantization index is greater than or less than the threshold, it can be used to obtain the final result. Boundary matching can be used to determine the threshold.

[0197] In this invention, the term "block" may refer to TU / TB, CU / CB, PU / PB, or CTU / CTB.

[0198] The term "LM" in this invention can be considered as a CCLM / MMLM mode or any extension / variation of CCLM (e.g., the CCLM extension / variation proposed in this invention). One variant is MMLM, which uses a threshold to determine different models for different samples in the current chroma component. Another variant, for Cb (or Cr), derives model parameters from multiple co-located luma blocks. More possible variants are shown below. CCLM variants here mean that when the block indication reference uses one of these cross-component modes (e.g., CCLM_LT, MMLM_LT, CCLM_L, CCLM_T, MMLM_L, MMLM_T and / or an intra-prediction mode that is not one of the traditional DC, planar, and angular modes), several optional modes can be selected for the current block. An example of using Convolutional Cross-Component Mode (CCCM) as an optional mode is shown below. When this optional mode is applied to the current block, chroma predictions are generated using cross-component information including nonlinear terms with the model. Optional modes may follow the template selection of CCLM, therefore the CCCM family includes CCCM_LT, CCCM_L, and / or CCCM_T.

[0199] The method proposed in this invention (for CCLM) can be used for any other cross-component pattern.

[0200] Any combination of methods proposed in this invention can be applied.

[0201] Any of the aforementioned methods for storing cross-component model information can be implemented in the encoder and / or decoder. For example, any proposed method can be implemented in the inter-frame / intra-frame / prediction / IBC / quantization module of the encoder, and / or in the inter-frame / intra-frame / prediction / IBC / quantization module of the decoder. Alternatively, any proposed method for storing cross-component model information can be implemented as circuitry connected to the inter-frame / intra-frame / prediction / IBC / quantization module of the encoder and / or the inter-frame / intra-frame / prediction / IBC / quantization module of the decoder to provide the information required by the inter-frame / intra-frame / prediction / IBC / quantization module.

[0202] Cross-component prediction, where cross-component model information is stored as described above, can be implemented at the encoder or decoder level. For example, any proposed method for storing cross-component model information can be implemented in the decoder's inter-frame / intra-frame encoding / decoding module (e.g., Figure 1B Intra Pred. 150 / MC 152) or the encoder's inter-frame / intra-frame codec module (e.g. Figure 1AThis is implemented in Intra Pred. 110 / Inter Pred. 112. Any proposed method can also be implemented as circuitry connected to the inter-frame / intra-frame codec module of the decoder or encoder. However, the decoder or encoder may also use additional processing units to implement the required cross-component prediction processing. While the Intra Pred. / MC unit (e.g.) Figure 1A and Figure 1B Units 110 / 112 and 150 / 152 are shown as separate processing units, which may correspond to executable software or firmware code stored on a hard disk or flash memory for use by a CPU (Central Processing Unit) or a programmable device (such as a DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array)).

[0203] Figure 14 A flowchart illustrating an exemplary video codec system according to an embodiment of the present invention, which stores cross-component model information for use or reference by one or more subsequent codec blocks, is presented. According to this method, in step 1410, input data associated with the current block is received, including a first color patch and a second color patch, wherein the input data packet contains pixel data to be encoded at the encoder end or data associated with the current block to be decoded at the decoder end. In step 1420, a target CCP (Cross-Component Prediction) model for the current block is determined. In step 1430, the target CCP model is stored. In step 1440, the second color patch is encoded or decoded using the target prediction generated based on the target CCP model of the current block.

[0204] The flowchart shown is intended to illustrate the video encoding / decoding according to the present invention. Those skilled in the art can practice the invention by modifying each step, rearranging the steps, splitting the steps, or combining the steps without departing from the spirit of the invention. Specific syntax and semantics are used in the disclosure to illustrate examples of implementing the invention. Those skilled in the art can practice the invention by substituting equivalent syntax and semantics without departing from the spirit of the invention.

[0205] The foregoing description is intended to enable those skilled in the art to practice the invention according to specific applications and requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments. Therefore, the invention is not intended to be limited to the specific embodiments shown and described, but rather to be given the broadest scope consistent with the principles and novel features disclosed herein. In the foregoing detailed description, various specific details have been shown to provide a thorough understanding of the invention. However, the invention can be practiced by those skilled in the art without specific practice.

[0206] As described above, embodiments of the present invention can be implemented in various hardware, software code, or combinations thereof. For example, one embodiment of the invention may be one or more circuits integrated into a video compression chip, or program code integrated into video compression software, to perform the processes described herein. Another embodiment of the invention may be program code to be executed on a digital signal processor (DSP) to perform the processes described herein. The invention may also relate to multiple functions executed by a computer processor, digital signal processor, microprocessor, or field programmable gate array (FPGA). These processors can be configured to perform specific tasks according to the invention, defining the specific methods embodied in the invention by executing machine-readable software code or firmware code. The software code or firmware code can be developed in different programming languages ​​and different formats or styles. The software code can also be compiled for different target platforms. However, different code formats, styles, and languages ​​of the software code, as well as other configuration codes, to perform tasks according to the invention, do not depart from the spirit and scope of the invention.

[0207] This invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The examples described are to be regarded in all respects as illustrative only, and not restrictive. Therefore, the scope of the invention should be indicated by the appended claims rather than the foregoing description. All variations within the meaning and equivalence of the claims should be included within its scope.

Claims

1. A method of coding a color picture or video containing one or more cross- component model related modes using a coding tool, the method comprising: receiving input data related to a current block, including a first color block and a second color block, wherein the input data comprises pixel data to be encoded at an encoder end or data related to the current block to be decoded at a decoder end, and wherein the current block is coded in a non-intra mode; determining a target cross-component prediction (CCP) model for the current block; storing the target CCP model; and encoding or decoding the second color block using a target prediction generated for the current block according to the target CCP model.

2. The method of claim 1, wherein the target CCP model contains cross-component mode (CCM) information, Cross-Component Model (CCM) information, or both.

3. The method of claim 1, wherein the target CCP model corresponds to a self-derived cross-component model or an inherited cross-component model.

4. The method of claim 1, wherein the target CCP model contains CCM information and model parameters of an inherited cross-component model.

5. The method of claim 4, wherein the CCM information of the inherited cross- component model comprises a cross-component linear model (CCLM), a convolutional cross- component model (CCCM), a CCCM with different filters, or a combination thereof.

6. The method of claim 4, further comprising refining the CCM information according to previously stored CCM information.

7. The method of claim 1, wherein inherited model parameters related to CCM information are refined.

8. The method of claim 7, wherein the inherited model parameters are refined using different types of templates and / or different numbers of lines.

9. The method of claim 1, wherein the target CCP model corresponds to an inherited cross-component model from a chroma intra-merge mode.

10. The method of claim 9, wherein the chroma intra-merge mode is derived by merging an intra prediction of non-cross-component coding and an intra prediction of cross-component coding.

11. The method of claim 9, wherein when CCM information is inherited from a block or a position coded by the chroma intra-merge mode, model parameters used to obtain the intra prediction of cross-component coding are inherited and further refined.

12. The method of claim 1, wherein the stored target CCP model is used or referred to by one or more subsequent coded blocks.

13. A coding apparatus for coding a color picture or video containing one or more cross-component model related modes using a coding tool, the apparatus comprising one or more electronic circuits or processors configured to: ​ receiving input data associated with a current block, including a first color block and a second color block, wherein the input data reports pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side, and wherein the current block is coded in a non-intra mode; determining a target CCP model for the current block; storing the target CCP model; and encoding or decoding the second color block using a target prediction generated for the current block according to the target CCP model. ​