Methods and apparatus of block template for cross-component prediction modes

WO2026200869A1PCT designated stage Publication Date: 2026-10-01MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/085476
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-03-24
Publication Date
2026-10-01

Smart Images

  • Figure CN2026085476_01102026_PF_FP_ABST
    Figure CN2026085476_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for video coding using coding tools involving a template for evaluating cross-component models are disclosed. According to the method, a block template for a current block is determined based on MVs or BVs of the current block. A CCP (Cross-Component Prediction) candidate list comprising multiple candidates is derived, each CCP candidate being associated with a CCP model. A template cost for each CCP candidate is calculated for a respective CCP candidate. The template cost comprises a block template cost determined based on a distortion between a chroma component of the block template and a chroma prediction generated by applying the CCP model of the respective CCP candidate to a luma component of the block template. The CCP candidates are reordered into an ordered set based on the calculated template costs. A target CCP candidate is selected for coding the chroma component of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND APPARATUS OF BLOCK TEMPLATE FOR CROSS-COMPONENT PREDICTION MODESCROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 777,053, filed on March 25, 2025. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION

[0002] The present invention relates to video coding system that employs coding tools involving a neighbouring template for evaluating prediction candidates. In particular, the present invention relates to a new type of template to improve the coding performance. BACKGROUND AND RELATED ART

[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.

[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction 110, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.

[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.

[0006] The decoder, as shown in Fig. 1B, can use some of the functional blocks as the encoder. For example, the decoder can reuse Inverse Quantization 124 and Inverse Transform 126; however, Transform 118 and Quantization 120 are not needed at the decoder. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.

[0007] In VVC, the Sequence Parameter Set (SPS) and the Picture Parameter Set (PPS) contain high-level syntax elements that apply to entire coded video sequences and pictures, respectively. The Picture Header (PH) and Slice Header (SH) contain high-level syntax elements that apply to a current coded picture and a current coded slice, respectively.

[0008] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs. The individual CTUs in a slice are processed in raster-scan order. A bi-predictive (B) slice may be decoded using intra prediction or inter prediction with at most two motion vectors and reference indices to predict the sample values of each block. A predictive (P) slice is decoded using intra prediction or inter prediction with at most one motion vector and reference index to predict the sample values of each block. An intra (I) slice is decoded using intra prediction only.

[0009] I. 1 Cross-Component Prediction (CCP) Merge (a.k.a., Non-Local CCP) Mode

[0010] For chroma coding, a CCP merge candidate list is constructed from the spatial adjacent, temporal, spatial non-adjacent, history-based m or shifted temporal candidates. After including these candidates, default models are further included to fill the remaining empty positions in the merge list. Pruning operation is applied. After constructing the list, the CCP models in the list are reordered according to the SAD costs, which are obtained using the neighbouring template of the current block. More details are described below.

[0011] Spatial adjacent and non-adjacent candidates

[0012] The positions and inclusion order of the spatial adjacent and non-adjacent candidates are the same as those defined in ECM for regular inter merge prediction candidates.

[0013] Temporal and shifted temporal candidates

[0014] Temporal candidates are selected from the collocated picture. The position and inclusion order of the temporal candidates are the same as those defined in ECM for regular inter merge prediction candidates. The shifted temporal candidates are also selected from the collocated picture. The position of temporal candidates is shifted by a selected motion vector which is derived from motion vectors of neighbouring blocks.

[0015] History-based candidates

[0016] A history-based table is maintained to include the recently used CCP models, and the table is reset at the beginning of each CTU row. If the current list is not full after including spatial adjacent and non-adjacent candidates, the CCP models in the history-based table are added into the list.

[0017] Default candidates

[0018] CCLM candidates with default scaling parameters are considered, only when the list is not full after including the spatial adjacent, spatial non-adjacent, or history-based candidates. If the current list has no candidates with the single model CCLM mode, the default scaling parameters are {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, 4 / 8, -4 / 8, 5 / 8, -5 / 8, 6 / 8} . Otherwise, the default scaling parameters are {0, the scaling parameter of the first CCLM candidate + {1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, 4 / 8, -4 / 8, 5 / 8, -5 / 8, 6 / 8} } .

[0019] A flag is signalled to indicate whether the CCP merge mode is applied or not. If CCP merge mode is applied, an index is signalled to indicate which candidate model is used by the current block.

[0020] I. 2 CCP Merge for Chroma Inter Blocks

[0021] The cross-component prediction merge mode, described in Section “Cross-Component Prediction (CCP) merge (a.k.a., non-local CCP) mode” , is extended to chroma inter coding. The CCP models including CCLM, MMLM, CCCM, GLM, chroma fusion, CCP merge modes, and inter CCCM are stored and inherited for the following coding chroma intra and inter blocks. Similar to the CCP merge for chroma intra blocks, a flag is signalled to indicate whether a chroma inter block is coded using this mode. If the CCP merge mode is used, a CCP merge list is constructed in a similar way as that for chroma intra blocks except that additional shifted temporal candidate and on-the-fly derived candidates are included in the CCP merge list.

[0022] The additional shifted temporal candidates are derived from the collocated picture. The position of these candidates are the same as those defined in ECM for regular inter merge prediction candidates with a shift obtained from the motion vector of the current block. The on-the-fly derived candidates are only used for low delay pictures and are obtained using the neighbouring reconstructed samples of the current block. At most one on-the-fly derived candidates including single / multi-model CCCM and single / multi-model CCLM is added to the CCP merge list. After the CCP merge list is constructed, the candidate with the lowest template cost is selected for the chroma inter block. The chroma inter block is then predicted in the same way as that of inter CCCM. That is, the motion compensation predicted samples are blended with the cross-component predicted samples to form the final prediction.

[0023] I. 3 Decoder Derived CCP Mode

[0024] In this method, a candidate list of cross-component prediction (CCP) modes is constructed, and to select the best candidate from the list a template cost is calculated to compare the reconstructed samples and the prediction values generated by the evaluated CCP mode. The template is indicated by grey areas in Fig. 2.

[0025] The CCP mode list is constructed from the already existed in ECM modes by single model CCLM, single model CCCM, multi-model CCCM, single model GLCCCM, single model CCCM applied with LBCCP, and multi-model CCCM applied with LBCCP.

[0026] In the second aspect of the method, various decoder-derived CCP fusion candidates are added. A fusion candidate is the combination of two CCP modes selected from the existing CCP mode lists reordered by template costs. Mode flag and a fusion flag are signalled to indicate the mode usage.

[0027] I. 4 Intra Template Matching

[0028] Intra template matching prediction (IntraTMP) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the most similar template to the current template in a reconstructed part of the current frame and uses the corresponding block as a prediction block. The encoder then signals the usage of this mode, and the same prediction operation is performed at the decoder side.

[0029] The prediction signal is generated by matching the L-shaped, Top-only or Left-Only causal neighbour of the current block with another block in a predefined search area.

[0030] I. 5 Intra Block Copy (IBC)

[0031] Intra block copy (IBC) is a tool adopted in HEVC extensions on SCC. It is well known that it significantly improves the coding efficiency of screen content materials. Since IBC mode is implemented as a block level coding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, a block vector is used to indicate the displacement from the current block to a reference block, which is already reconstructed inside the current picture. An IBC-coded CU is treated as the third prediction mode other than intra or inter prediction modes.

[0032] At CU level, IBC mode is signalled with a flag and it can be signalled as IBC AMVP mode or IBC skip / merge mode as follows:

[0033] IBC skip / merge mode: a merge candidate index is used to indicate which of the block vectors in the list from neighbouring candidate IBC coded blocks is used to predict the current block. The merge list consists of spatial, HMVP, and pairwise candidates.

[0034] IBC AMVP mode: block vector difference is coded in the same way as a motion vector difference. The block vector prediction method uses two candidates as predictors, one from left neighbour and one from above neighbour (if IBC coded) . When either neighbour is not available, a default block vector will be used as a predictor. A flag is signalled to indicate the block vector predictor index.

[0035] I. 6 Decoder Side Intra Mode Derivation (DIMD)

[0036] When DIMD is applied, a number of intra modes are derived from the reconstructed neighbour samples by computing a histogram of gradients (HoG) , which is built by applying edge operators to a reference area. The edge operators are used to derive the direction and magnitude (amplitude) of gradients within the area covered by an edge operator. In particular, up to seven intra modes are derived if the block size area is larger or equal than 128 samples, and up to five intra modes are derived otherwise. Each derived intra mode is used to form an intra predictor. These predictors are combined with the non-directional predictor (planar or block vector based predictor) with the weights derived from the histogram of gradients as described in JVET-O0449 and JVET-AJ0267. The decision between for the non-directional modes is taken according to the template cost. Specifically, the block vectors of all adjacent and non-adjacent merge candidates (coded in IntraTMP or IBC) are compared to planar prediction on the reconstructed template. The template cost (SATD) is used to select the best predictor among them.

[0037] I. 7 Fusion for Template-Based Intra Mode Derivation (TIMD)

[0038] For each intra prediction mode in MPMs, as well as the wide-angle modes if the above-right and / or bottom-left reference samples are available, SATD between the prediction and reconstruction samples of the template is calculated. First two intra prediction modes with the minimum SATD and one non-angular intra prediction mode (i.e. DC or Planar) with the lowest SATD cost are selected as the TIMD modes. These three TIMD modes are fused with the weights after applying PDPC process, and such weighted intra prediction is used to code the current CU. Position dependent intra prediction combination (PDPC) is included in the derivation of the TIMD modes.

[0039] I. 8 Spatial Geometric Partitioning Mode (SGPM)

[0040] SGPM is an intra mode that resembles the inter coding tool of GPM, where the two prediction parts are generated from intra predicted process. In this mode, a candidate list is built with each entry containing one partition split and two intra prediction modes. 26 partition modes and 9 of intra prediction modes are used to form the combinations. The length of the candidate list is set equal to 16. The selected candidate index is signalled.

[0041] For each partition mode, an IPM list is derived for each part using the same intra-inter GPM list derivation as GPM. The list is further augmented with block-vector based prediction candidates obtained from the adjacent and non-adjacent merge candidates coded in IntraTMP or IBC mode. The final list contains up to 9 predictors: 3 regular intra modes and up to 6 block vectors based predictors.

[0042] I. 9 JVET-AK0062 EE2-1.14: Block Vector Guided Extrapolation Intra Prediction (EIP)

[0043] The block-vector guided EIP (BV-EIP or BVG-EIP) method uses a block vector to determine the reference area for calculating the EIP filter parameters. Then the current coding unit utilizes the calculated EIP filter parameters and neighbouring reconstructed samples to create the prediction. Fig. 3 illustrates the reference area and the current block along with the template of the current block used in the BV-EIP method. The block vector used in this method is derived from the rough searching process of the intraTMP in the current ECM.

[0044] I. 10 JVET-AK0060 EE2-1.7: Block vector guided DIMD

[0045] Proposing to use the reference samples pointed at by the block vectors to build HoG and then derive intra modes. Up to 5 block vectors are obtained by the same searching process as in intraTMP. With each block vector, a reference area in the current picture is determined and its samples are used to derive the intra modes and corresponding amplitudes. Overlapping reference areas are clustered. All results are counted in constructing the HoG. Two intra modes with the highest amplitudes are selected from the HoG and the prediction of the current block is the blending of those 2 predictors and a non-directional predictor. The non-directional predictor is the blending of up to 5 BV-based predictors obtained using the block vectors.

[0046] In the present invention, methods and apparatus to use a new type of template are disclosed to improve the performance, where the new type of template is referred as a block template that is derived using motion vectors or block vectors associated with the current block. BRIEF SUMMARY OF THE INVENTION

[0047] A method and apparatus for video coding using coding tools involving a template for evaluating cross-component models are disclosed. According to the method, input data associated with a current block is received, the current block comprises a luma component and at least one chroma component. A block template corresponding to the current block is determined based on one or more MVs (Motion Vectors) or one or more BVs (Block Vectors) associated with the current block. A CCP (Cross-Component Prediction) candidate list comprising a plurality of CCP candidates is derived, each CCP candidate being associated with a CCP model. A template cost for each of the plurality of CCP candidates is calculated, wherein the template cost for a respective CCP candidate comprises a block template cost determined based on a distortion or differences between a chroma component of the block template and a chroma prediction generated by applying the CCP model of the respective CCP candidate to a luma component of the block template. The plurality of CCP candidates is reordered into an ordered set based on the calculated template costs. At least one target CCP candidate is selected from the ordered set. The chroma component of the current block is encoded or decoded using the at least one target CCP candidate.

[0048] In one embodiment, the luma component and the chroma component of the block template are determined using reference blocks identified by said one or more MVs or said one or more BVs. In one embodiment, the luma component and the chroma component of the block template correspond to reconstruction samples generated based on motion compensation using said one or more MVs of the current block or based on prediction generated according to said one or more BVs of the current block.

[0049] In one embodiment, the block template and the current block have a same size.

[0050] In one embodiment, for the current block coded in Intra Template Matching Prediction (IntraTMP) , Intra Block Copy (IBC) , Decoder Side Intra Mode Derivation (DIMD) , Spatial Geometric Partitioning Mode (SGPM) , Template-based Intra Mode Derivation (TIMD) , Block-Vector Guided Extrapolation Intra Prediction (BVG-EIP) , or BVG-DIMD, the luma component and the chroma component of the block template are determined using reference blocks identified by said one or more BVs.

[0051] In one embodiment, the block template is determined based on said one or more MVs or said one or more BVs from one or more neighbouring blocks of the current block.

[0052] In one embodiment, the template cost for the respective CCP candidate further comprises a neighbouring-area template cost calculated using a neighbouring-area template of the current block. In one embodiment, the plurality of CCP candidates are reordered based on: the block template cost, conditioned upon the current block being coded in inter mode, IBC mode, IntraTMP mode, coding mode utilizing block vectors; and the neighbouring-area template cost, conditioned upon the current block being coded in other modes.

[0053] In one embodiment, a final template cost is calculated as a weighted combination of the block template cost and the neighbouring-area template cost. In one embodiment, weighting for calculating the weighted combination of the block template cost and the neighbouring-area template cost is normalized based on at least one of: a size of the neighbouring-area template, a width of the current block, a height of the current block, an area of the current block, a value of the block template cost, or a value of the neighbouring-area template cost. In one embodiment, in response to a difference between a normalized block template cost and a normalized neighbouring-area template cost exceeding a threshold value, weighting factors are adjusted to increase relative contribution from one of the normalized block template cost and the normalized neighbouring-area template cost having a lower value. In one embodiment, a trust value is derived based on at least one of: the block template cost, the neighbouring-area template cost, the normalized block template cost, or the normalized neighbouring-area template cost; and wherein, in response to a difference between the block template cost and the neighbouring-area template cost or between the normalized block template cost and the normalized neighbouring-area template cost exceeding a threshold, the trust value is adjusted to increase the final template cost or to move the associated CCP candidate backward in the ordered set.

[0054] In one embodiment, weighting for calculating the weighted combination of the block template cost and the neighbouring-area template cost is adjusted based on a coding mode for at least one of the neighbouring-area template or the luma component of the block template.

[0055] In one embodiment, the CCP model of the respective CCP candidate is derived from samples in both the neighbouring-area template and the block template.

[0056] In one embodiment, a size-based condition is evaluated by comparing at least one of a size, a width, or a height of the current block with a predefined threshold; and wherein, in response to the size-based condition being satisfied, a final template cost is computed exclusively based on the neighbouring-area template cost.

[0057] In one embodiment, a template selection syntax element is signalled in a bitstream to specify a template configuration for determining a final template cost, the template configuration being selected from a group comprising: the block template cost, the neighbouring-area template cost, and a combination thereof.

[0058] In one embodiment, the block template cost is computed using a subset of samples selected from the block template, the subset of samples comprising fewer than all samples within the block template. In one embodiment, use of the subset of samples is performed in response to a size-based condition being satisfied, the size-based condition being evaluated by comparing at least one of a size, a width, or a height of the current block with a predefined threshold.

[0059] In one embodiment, the template cost for the respective CCP candidate is determined by generating a chroma prediction signal based exclusively on luma samples within a template.BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Fig. 1A illustrates an exemplary adaptive Inter / Intra video coding system incorporating loop processing.

[0061] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.

[0062] Fig. 2 illustrates an example of a template area adjacent to the current chroma CU.

[0063] Fig. 3 illustrates an example of reference area for Block-Vector Guided Extrapolation Intra Prediction (BV-EIP) .

[0064] Fig. 4 illustrates an example of neighbouring templates used to derive cross-component model.

[0065] Fig. 5A illustrates an example, where a reference block being located by the motion vector is used to generate the luma and chroma components of the block template through motion compensation for a uni-directional coded block.

[0066] Fig. 5B illustrates an example, where a reference block being located by the motion vectors is used to generate the luma and chroma components of the block template through motion compensation for a bi-directional coded block.

[0067] Fig. 6 illustrates an example, where when the current block is coded in a mode utilizing block vectors, one or more block vectors used to determine the block template can be the block vectors of the current block.

[0068] Fig. 7 illustrates an example, where, when computing the template cost of a CCP model, only the luma samples inside the template can be used to generate the chroma prediction of the template.

[0069] Fig. 8 illustrates a flowchart of an exemplary video coding system that uses a new type of template to improve the performance according to an embodiment of the present invention, where the new type of template is referred as a block template that is derived using motion vectors or block vectors associated with the current block.DETAILED DESCRIPTION OF THE INVENTION

[0070] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.

[0071] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.

[0072] PROPOSED METHOD

[0073] Cross-component prediction (CCP) tools are widely adopted in video coding standards. In CCP tools, the redundancy between cross-components is exploited by generating the prediction of a second colour component based on a first colour component. In some CCP tools, one or more lists comprising CCP models are constructed and the potential prediction errors of CCP models in the list are evaluated. The potential prediction errors of CCP models may be used to reorder the CCP models in the list to reduce the signalling overhead. The potential prediction errors of CCP models may be used to select one or more CCP models from the CCP models in the list. The CCP models with the smallest potential prediction error may be selected. For example, for CCP merge mode described in Section I. 1 “Cross-Component Prediction (CCP) merge (a.k.a., non-local CCP) mode” , the CCP models in the merge candidate list are reordered to reduce the signalling overhead. For another example, for a CCP merge mode for chroma inter blocks described in Section I. 2, and Decoder Derived CCP mode described in Section I. 3, the potential prediction errors of CCP models are evaluated to select a CCP model from the candidate list.

[0074] The potential prediction error of a CCP model can be evaluated by using a template matching method. A template can be an area in previous coded region / pictures or an area containing samples derived based on previous coded region / pictures. The chroma prediction of the template is generated based on the corresponding luma component of the template and the CCP model. The template cost is the distortion / difference between the chroma prediction (generated based on the CCP model and the corresponding luma component) and the chroma component of the template. The luma and chroma components of the template can be the luma reconstruction samples and the chroma reconstruction samples respectively. The distortion metric can be SAD (sum of absolute difference) or SATD (sum of absolute transformed difference) .

[0075] As shown in Fig. 4, the template can be a neighbouring area above the current block, a neighbouring area left to the current block, or the combination thereof. The size of above neighbouring template is wa×ha, and the size of left neighbouring template is wb×hb. For example, in CCP merge mode and CCP merge mode for chroma inter blocks, wa is the width of the current block, hb is the height of the current block, ha and wb are 1. For example, ha and wb can be a pre-defined positive number such as 1, 2, 3, 4, 6, 8, etc. For another example, ha and wb can be adaptive determined. ha and wb can be determined based on the size / width / height of the current block.

[0076] In this invention, methods to determine the template for cross-component prediction tools are proposed. In one embodiment, when one or more motion vectors can be identified at the current block, instead of using neighbouring area as template, the template can be a block determined based on the one or more motion vectors. In another embodiment, when one or more block vectors can be identified at the current block, instead of using neighbouring area as template, the template can be a block determined based on the one or more block vectors. The block size can be the same as the current block. The block template cost of a CCP model can be computed as the distortion / difference between the chroma component of the block template and the chroma prediction generated based on the luma component of the block template and the CCP model.

[0077] For CCP tools that use template cost to evaluate the potential prediction errors of CCP models, as described in previous paragraphs, the block template cost can be used instead of neighbouring-area template cost to evaluate the potential prediction errors of CCP models. For example, for CCP merge mode for chroma inter blocks described in Section I. 2, the CCP models with the smallest block template cost can be selected from the merge candidate list. For example, for Decoder Derived CCP mode described in Section I. 3, the CCP models with the smallest block template cost can be selected from the candidate list. For another example, for CCP merge mode described in Section I. 1, the CCP models in in the merge candidate list can be reordered based on the block template cost to reduce the signalling overhead.

[0078] In one embodiment, when the current block is coded in an inter mode, the one or more motion vectors used to determine the block template can be the motion vectors of the current block, as depicted in Fig. 5A and Fig. 5B. The reference blocks located by the motion vectors can be used to determine the luma and chroma components of the block template. The luma and chroma components of the block template are computed by motion compensation based on the motion vectors of the current block. When the current block utilizes uni-direction prediction as in Fig. 5A, the reference block located by the motion vector is used to generate the luma and chroma components of the block template through motion compensation. When the current block utilizes bi-directional prediction (i.e., the current block has more than one motion vector) , the luma and chroma components of the block template can be generated by bi-prediction motion compensation as shown in Fig. 5B. For example, for CCP merge mode for chroma inter blocks described in Section I. 2, since the current block is coded in inter mode, the one or more motion vectors of the current block can be used to determine the block template.

[0079] In one sub-embodiment, the luma components of the block template is the reconstruction samples generated based on motion compensation based on the motion vectors of the current block.

[0080] In one embodiment, when the current block is coded in a mode utilizing block vectors, the one or more block vectors used to determine the block template can be the block vectors of the current block, as depicted in Fig. 6. For example, the coding mode of the current block can be IntraTMP, IBC. For another example, the coding mode of the current block may be DIMD, SGPM, TIMD, BVG-EIP, BVG-DIMD. The reference blocks located by the block vectors can be used to determine the luma and chroma components of the block template. For example, if the current block is coded in IBC mode, the luma and chroma components of the block template are computed by motion compensation based on the block vectors of the current block. When the current block utilizes uni-direction prediction, the reference block located by the block vector is used to generate the luma and chroma components of the block template through motion compensation. When the current block utilizes bi-directional prediction (i.e., the current block has more than one block vectors) , the luma and chroma components of the block template can be generated by bi-prediction motion compensation. For another example, if the current block is coded in IntraTMP mode, the luma and chroma component of the block template can be the IntraTMP prediction, which is the combination of multiple BV-predictors. For example, for CCP merge mode described in Section I. 1 and Decoder Derived CCP mode described in Section I. 3, when the current block is coded in IntraTMP mode or other intra coding modes utilizing block vectors, the one or more block vectors of the current block can be used to determine the block template.

[0081] In one sub-embodiment, the luma and chroma components of the block template are the same as the reference block located by the block vector.

[0082] In one sub-embodiment, the luma components of the block template is the reconstruction samples generated based on the prediction generated according to the block vectors of the current block. For example, if the current block is coded in IntraTMP mode, the reconstruction samples are generated based on the combination of multiple BV-predictors. For another example, if the current block is coded in IBC mode, the reconstruction samples are generated based on BV-based motion compensation.

[0083] In one embodiment, the one or more motion vectors used to determine the block template can be the motion vectors of a neighbouring block of the current block. The luma and chroma components of the block template are computed by motion compensation based on the motion vectors of the neighbouring block.

[0084] In one embodiment, the one or more block vectors used to determine the block template can be the block vectors of a neighbouring block of the current block. The luma and chroma components of the block template are computed by motion compensation based on the block vectors of the neighbouring block.

[0085] In one embodiment, the block template cost is used to evaluate the potential prediction error only when the current block is coded in inter mode, IBC mode, IntraTMP mode, coding mode utilizing block vectors (e.g. DIMD, SGPM, TIMD, BVG-EIP, BVG-DIMD) , or the combination thereof. Otherwise, the neighbouring-area template is used.

[0086] In one embodiment, when selecting CCP models from candidate CCP models in the list, if there are M candidate CCP models in the list, N (N is smaller than M) candidate CCP models with lowest N template costs based on neighbouring-area template are selected first. Then block template costs are calculated for these N candidate CCP models. The one or more candidate CCP models with the smallest block template costs is selected as the final candidates.

[0087] In one embodiment, when evaluating the potential prediction errors of CCP models in the list, if there are M CCP models in the list, neighbouring-area template cost is computed for each M CCP models. For N (N is smaller than M) CCP models with lowest N neighbouring-area template costs, block template cost is computed for each of the N CCP models. For example, the M CCP models can be reordered first based on the neighbouring-area template cost, then the N CCP models are reordered again based on the block template cost.

[0088] In one embodiment, the final template cost used to evaluate the potential prediction error is computed based on the block template cost and the neighbouring-area template cost. The final template cost can be a weighted combination of block template cost and the neighbouring-area template cost.

[0089] In one sub-embodiment, when the size / width / height of the current block is greater than / greater than or equal to / smaller than / smaller than or equal to a threshold, the final template cost is computed only based on the neighbouring-area template cost.

[0090] In one embodiment, a flag can be signalled to indicate the type of template used to compute template cost. For example, if the flag is true, the block template cost is used. Otherwise, the neighbouring-area template cost is used. For another example, if the flag is true, the neighbouring-area template cost is used. Otherwise, the block template cost is used. The flag can be signalled at block level, CU level, PU level, TU level, CTU row level, area level, slice level, picture level, sequence level, or a high-level syntax.

[0091] In one sub-embodiment, the block template cannot be used when the size / width / height of the current block is greater than / greater than or equal to / smaller than / smaller than or equal to a threshold. If the size / width / height of the current block is greater than / greater than or equal to / smaller than / smaller than or equal to a threshold, the flag is not signalled and the neighbouring-area template is used to compute template cost.

[0092] In one embodiment, the block template cost is computed based on only partial samples in the block template. For example, block template cost is computed based on only sub-sampled positions of the template. For another example, the block template cost is computed based only partial region of the block template.

[0093] In one embodiment, the block template cost is computed based on only partial samples in the block template if the size / width / height of the current block is greater than / greater than or equal to / smaller than / smaller than or equal to a threshold.

[0094] In another embodiment, when computing the template cost of a CCP model, only the luma samples inside the template can be used to generate the chroma prediction of the template. For example, a 5-tap convolutional filter (in a cross shape) is used in CCCM, assume the current template size is M×N, after applying the 5-tap convolutional filter to the luma samples (maybe downsampled according to the colour format) , the generated chroma prediction size is (M-2) × (N-2) . For example, as shown in the following figure, given an 8×8 template, the 5-tap convolutional filter is applied to the luma reconstruction samples from the centre positions (1, 1) to (6, 6) , and the generated chroma prediction size is (8-2) × (8-2) = 6×6, which is the region marked in dash lines in Fig. 7.

[0095] In one embodiment, when computing the template cost of a CCP model, if luma samples outside of the template are needed to generate the chroma prediction of the template. The luma samples outside of the template are generated by using padding. For example, for a 5-tap convolutional filter (in a cross shape) is used in CCCM, one line of luma samples above / left to / right to / below the template are generated by padding.

[0096] In one embodiment, the final template cost used to evaluate the potential prediction error is computed based on the block template cost and the neighbouring-area template cost. The final template cost is a weighted combination of the block template cost and the neighbouring-area template cost according to a weighting setting with normalization. The normalization depends on the neighbouring-area template size, block width, block height, block area, the block template cost, the neighbouring-area template cost, or any combination thereof. For example, the neighbouring-area template cost is calculated based on the neighbouring-area including above neighbouring-area (block width x neighbouring-area template size, or wa×ha) and left neighbouring-area (neighbouring-area template size x block height, or wb×hb) , and the block template cost is calculated based on the block template (block width x block height) . In one way, the normalization is applied to increase the neighbouring-area template cost and then the increased neighbouring-area template cost is combined with the block template cost. For example, the neighbouring-area template cost is multiplied by a factor proportional to (block width x block height)  /  (wa×ha + wb×hb) . In another way, the normalization is applied to decrease the block template cost and then the decreased block template cost is combined with the neighbouring-area template cost. For example, the block template cost is multiplied by a factor proportional to (wa×ha + wb×hb)  /  (block width x block height) . In another example, the weight used to do weighted combination or normalization can be signalled at block level, CU level, PU level, TU level, CTU row level, area level, slice level, picture level, sequence level, or a high-level syntax.

[0097] In another example, if the normalized block template cost and the normalized neighbouring-area template cost are very different, it implies that the content of block template and neighbouring template are very different. The weight of final cost can favour either the block template cost or the neighbouring-area template cost. In another example, the weight of the block template cost and the neighbouring-area template cost can be derived from the information of the block template cost and the neighbouring-area template cost. For example, if the ratio of the normalized block template cost and the normalized neighbouring-area template cost is between a first range, the weight is equal to Value1; if the ratio is between a second range, the weight is equal to Value2.

[0098] In another embodiment, the cost / normalized cost of the block template and the neighbouring-area template can be used to derive a trust value or a weight for the final cost. The modified final cost (e.g. multiplied by the trust value or the weight) is used to do the reordering. For example, if the normalized cost of the block template and the neighbouring-area template are very different, it implies that the content of block template and neighbouring template are very different. The cost of this candidate is not trustable. The final cost of this candidate should be increased, or the priority of this candidate should be moved backward in the list.

[0099] In one embodiment, the final template cost used to evaluate the potential prediction error is computed based on the block template cost and the neighbouring-area template cost. The final template cost is a weighted combination of the block template cost and the neighbouring-area template cost according to a weighting setting based on coding modes for the neighbouring-area template and / or the luma component of the block template. For example, when the neighbouring-area template is associated with intra coding modes, the weight for the neighbouring-area template cost is higher than the weight for the block template cost. When the luma component of the block template is associated with intra coding modes, the weight for the block template cost is higher than the weight for the neighbouring-area template cost. In another example, the weight used to do weighted combination can be signalled at block level, CU level, PU level, TU level, CTU row level, area level, slice level, picture level, sequence level, or a high-level syntax.

[0100] In another embodiment, the CCP model can be derived from the samples in both neighbouring-area template and block template. If the costs of two templates are very different, the CCP model can either favour the model derived from neighbouring-area template or the model derived from block template.

[0101] The term “block” in this invention can refer to TU / TB, CU / CB, PU / PB, pre-defined region, or CTU / CTB.

[0102] Any combination of the proposed methods in this invention can be applied.

[0103] Any of the foregoing proposed methods of using a block template to evaluate cross-component candidates can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter / intra / IBC / prediction / transform module of an encoder, and / or an inter / intra / IBC / prediction / transform module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / IBC / prediction / transform module of the encoder and / or the inter / intra / IBC / prediction / transform module of the decoder, so as to provide the information needed by the inter / intra / IBC / prediction / transform module.

[0104] With reference to the exemplary encoder and decoder in Fig. 1A and Fig. 1B, any of the proposed methods can be implemented in an Intra / Inter coding module (e.g. Intra Pred. 150 / MC 152 in Fig. 1B) in a decoder or an Intra / Inter coding module is an encoder (e.g. Intra Pred. 110 / Inter Pred. 112 in Fig. 1A) . Any of the proposed methods can also be implemented as a circuit coupled to the intra / inter coding module at the decoder or the encoder. The proposed methods may also correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .

[0105] Fig. 8 illustrates a flowchart of an exemplary video coding system that uses a new type of template to improve the performance according to an embodiment of the present invention, where the new type of template is referred as a block template that is derived using motion vectors or block vectors associated with the current block. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, input data associated with a current block is received in step 810, the current block comprises a luma component and at least one chroma component. A block template corresponding to the current block is determined based on one or more MVs (Motion Vectors) or one or more BVs (Block Vectors) associated with the current block in step 820. A CCP (Cross-Component Prediction) candidate list comprising a plurality of CCP candidates is derived in step 830, each CCP candidate being associated with a CCP model. A template cost for each of the plurality of CCP candidates is calculated in step 840, wherein the template cost for a respective CCP candidate comprises a block template cost determined based on a distortion or differences between a chroma component of the block template and a chroma prediction generated by applying the CCP model of the respective CCP candidate to a luma component of the block template. The plurality of CCP candidates is reordered into an ordered set based on the calculated template costs in step 850. At least one target CCP candidate is selected from the ordered set in step 860. The chroma component of the current block is encoded or decoded using the at least one target CCP candidate in step 870.

[0106] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0107] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.

[0108] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.

[0109] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1.A method of coding colour pictures, the method comprising:receiving input data associated with a current block, the current block comprises a luma component and at least one chroma component;determining a block template corresponding to the current block based on one or more MVs (Motion Vectors) or one or more BVs (Block Vectors) associated with the current block;deriving a CCP (Cross-Component Prediction) candidate list comprising a plurality of CCP candidates, each CCP candidate being associated with a CCP model;calculating a template cost for each of the plurality of CCP candidates, wherein the template cost for a respective CCP candidate comprises a block template cost determined based on a distortion or differences between a chroma component of the block template and a chroma prediction generated by applying the CCP model of the respective CCP candidate to the luma component of the block template;reordering the plurality of CCP candidates into an ordered set based on the calculated template costs;selecting at least one target CCP candidate from the ordered set; andencoding or decoding the chroma component of the current block using the at least one target CCP candidate.2.The method of Claim 1, wherein the luma component and the chroma component of the block template are determined using reference blocks identified by said one or more MVs or said one or more BVs.3.The method of Claim 2, wherein the luma component and the chroma component of the block template correspond to reconstruction samples generated based on motion compensation using said one or more MVs of the current block or based on prediction generated according to said one or more BVs of the current block.4.The method of Claim 1, wherein the block template and the current block have a same size.5.The method of Claim 1, wherein, for the current block coded in Intra Template Matching Prediction (IntraTMP) , Intra Block Copy (IBC) , Decoder Side Intra Mode Derivation (DIMD) , Spatial Geometric Partitioning Mode (SGPM) , Template-based Intra Mode Derivation (TIMD) , Block-Vector Guided Extrapolation Intra Prediction (BVG-EIP) , or BVG-DIMD, the luma component and the chroma component of the block template are determined using reference blocks identified by said one or more BVs.6.The method of Claim 1, wherein the block template is determined based on said one or more MVs or said one or more BVs from one or more neighbouring blocks of the current block.7.The method of Claim 1, wherein the template cost for the respective CCP candidate further comprises a neighbouring-area template cost calculated using a neighbouring-area template of the current block.8.The method of Claim 7, wherein the plurality of CCP candidates are reordered based on: the block template cost, conditioned upon the current block being coded in inter mode, IBC mode, IntraTMP mode, coding mode utilizing block vectors; and the neighbouring-area template cost, conditioned upon the current block being coded in other modes.9.The method of Claim 7, wherein a final template cost is calculated as a weighted combination of the block template cost and the neighbouring-area template cost.10.The method of Claim 9, wherein weighting for calculating the weighted combination of the block template cost and the neighbouring-area template cost is normalized based on at least one of: a size of the neighbouring-area template, a width of the current block, a height of the current block, an area of the current block, a value of the block template cost, or a value of the neighbouring-area template cost.11.The method of Claim 10, wherein, in response to a difference between a normalized block template cost and a normalized neighbouring-area template cost exceeding a threshold value, weighting factors are adjusted to increase relative contribution from one of the normalized block template cost and the normalized neighbouring-area template cost having a lower value.12.The method of Claim 11, wherein a trust value is derived based on at least two of: the block template cost, the neighbouring-area template cost, the normalized block template cost, or the normalized neighbouring-area template cost; and wherein, in response to a difference between the block template cost and the neighbouring-area template cost or between the normalized block template cost and the normalized neighbouring-area template cost exceeding a threshold, the trust value is adjusted to increase the final template cost or to move the associated CCP candidate backward in the ordered set.13.The method of Claim 9, wherein weighting for calculating the weighted combination of the block template cost and the neighbouring-area template cost is adjusted based on a coding mode for at least one of the neighbouring-area template or the luma component of the block template.14.The method of Claim 7, wherein the CCP model of the respective CCP candidate is derived from samples in both the neighbouring-area template and the block template.15.The method of Claim 7, wherein a size-based condition is evaluated by comparing at least one of a size, a width, or a height of the current block with a predefined threshold; and wherein, in response to the size-based condition being satisfied, a final template cost is computed exclusively based on the neighbouring-area template cost.16.The method of Claim 7, wherein a template selection syntax element is signalled in a bitstream to specify a template configuration for determining a final template cost, the template configuration being selected from a group comprising: the block template cost, the neighbouring-area template cost, and a combination thereof.17.The method of Claim 1, wherein the block template cost is computed using a subset of samples selected from the block template, the subset of samples comprising fewer than all samples within the block template.18.The method of Claim 17, wherein use of the subset of samples is performed in response to a size-based condition being satisfied, the size-based condition being evaluated by comparing at least one of a size, a width, or a height of the current block with a predefined threshold.19.The method of Claim 1, wherein the template cost for the respective CCP candidate is determined by generating a chroma prediction signal based exclusively on luma samples within a template.20.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block, the current block comprises a luma component and at least one chroma component;determine a block template corresponding to the current block based on one or more MVs (Motion Vectors) or one or more BVs (Block Vectors) associated with the current block;derive a CCP (Cross-Component Prediction) candidate list comprising a plurality of CCP candidates, each CCP candidate being associated with a CCP model;calculate a template cost for each of the plurality of CCP candidates, wherein the template cost for a respective CCP candidate comprises a block template cost determined based on a distortion or differences between a chroma component of the block template and a chroma prediction generated by applying the CCP model of the respective CCP candidate to the luma component of the block template;reorder the plurality of CCP candidates into an ordered set based on the calculated template costs;select at least one target CCP candidate from the ordered set; andencode or decode the chroma component of the current block using the at least one target CCP candidate.