Methods and apparatus of block template for unified intra merge candidate list for inter and IBC prediction

WO2026201126A1PCT designated stage Publication Date: 2026-10-01MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/086503
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure CN2026086503_01102026_PF_FP_ABST
    Figure CN2026086503_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus of deriving a unified merge candidate list are disclosed. According to this method, a block template for the current block is determined. A merge candidate list comprising at least one joint luma-chroma intra prediction information entry is determined, and each entry comprises both luma intra prediction information and chroma intra prediction information. Template costs associated with entries in the merge candidate list are evaluated, wherein the template costs comprise block template costs and each of the block template costs is evaluated between the block template and a predicted block template derived using corresponding joint luma-chroma intra-prediction information. The merge candidate list is reordered based on the template costs to generate a reordered merge candidate list. At least one target joint luma-chroma intra prediction information entry is selected from the reordered merge candidate list.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND APPARATUS OF BLOCK TEMPLATE FOR UNIFIED INTRA MERGE CANDIDATE LIST FOR INTER AND IBC PREDICTIONCROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 779,374, filed on March 28, 2025. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION

[0002] The present invention relates to merge mode for video coding systems. In particular, the present invention relates to using cost measurement based on a block template for reordering and / or selecting candidates from a merge candidate list comprising one or more intra prediction candidates that jointly inherit the luma and chroma intra prediction information from neighbouring spatial or temporal positions / blocks to the current block. BACKGROUND AND RELATED ART

[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.

[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter-Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.

[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.

[0006] The decoder, as shown in Fig. 1B, can use some of the functional blocks as the encoder. For example, the decoder can reuse Inverse Quantization 124 and Inverse Transform 126; however, Transform 118 and Quantization 120 are not needed at the decoder. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.

[0007] According to VVC, an input picture is partitioned into non-overlapped square block regions referred as CTUs (Coding Tree Units) , similar to HEVC. Each CTU can be partitioned into one or multiple smaller size coding units (CUs) . The resulting CU partitions can be in square or rectangular shapes. Also, VVC divides a CTU into prediction units (PUs) as a unit to apply prediction process, such as Inter prediction, Intra prediction, etc.

[0008] I. 1 Intra Block Copy (IBC)

[0009] Since IBC mode is implemented as a block level coding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, a block vector (BV) is used to indicate the displacement from the current block to a reference block, which is already reconstructed inside the current picture. An IBC-coded CU is treated as the third prediction mode other than intra or inter prediction modes.

[0010] At CU level, IBC mode is signalled with a flag and it can be signalled as IBC AMVP mode or IBC skip / merge mode as follows:

[0011] I. 2 Intra Mode Coding with 67 Intra Prediction Modes

[0012] To capture the arbitrary edge directions presented in natural video, the number of directional intra modes in VVC is extended from 33, as used in HEVC, to 65. Several conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for the non-square blocks.

[0013] I. 3 Decoder Side Intra Mode Derivation (DIMD)

[0014] When DIMD is applied, multiple intra modes are derived from the reconstructed neighbour samples (template) . Those predictors from the derived intra modes are combined with the non-directional mode (e.g. planar mode predictor or BV-based predictor) with the weights derived from the gradients.

[0015] A texture gradient analysis is performed at both encoder and decoder. This process starts with an empty Histogram of Gradient (HoG) with 65 entries, corresponding to the 65 angular modes. Amplitudes of these entries are determined during the texture gradient analysis.

[0016] I. 4 Template-Based Intra Mode Derivation (TIMD)

[0017] TIMD mode implicitly derives the intra prediction mode of a CU by a neighbouring template at both encoder and decoder, instead of being signal exact intra prediction mode bits to the decoder. The prediction samples of the template are generated using the reference samples of the template for each candidate mode. A cost is calculated as the SATD between the prediction and the reconstruction samples of the template. First two intra prediction modes with the minimum SATD are selected as the TIMD modes. These two TIMD modes are fused with the weights to generate prediction for the current CU.

[0018] I. 5 Extrapolation Filter-Based Intra Prediction (EIP) Mode

[0019] In EIP mode, the samples in a CU are predicted from the top-left position to the bottom-right position by applying an extrapolation filter to neighbouring reconstructed samples or predicted samples. The EIP mode uses a 15-tap filter for prediction as below: where pred (x, y) is the predicted value at position (x, y) in the CU, ci is the filter coefficient, and the represents the reconstructed samples or predicted samples.

[0020] The EIP filter can be derived from the neighbouring reconstructed samples or be inherited from the previous EIP coded blocks. There are three EIP filter shapes and three types of reconstructed area.

[0021] I. 6 Template-Based Multiple Reference Line Intra Prediction

[0022] Template-based multiple reference line intra prediction (TMRL) mode combines reference line and prediction mode together and uses a template matching method to construct a list of candidate combinations. An index to the candidate combination list is signalled.

[0023] The extended reference line starts from reference line 1. Reference line 0 is used for template matching. The SAD costs (TMRL costs) over the template area (Fig. 2) are calculated between the predictions (generated by 50 combinations) and the reconstructions. The 20 combinations with the least SAD cost are selected in an ascending order to form the TMRL candidate list.

[0024] I. 7 Spatial Geometric Partitioning Mode (SGPM)

[0025] SGPM is an intra mode that resembles the inter coding tool of GPM, where the two prediction parts are generated from intra predicted process. In this mode, a candidate list is built with each entry containing one partition split and two intra prediction modes. 26 partition modes and 9 of intra prediction modes are used to form the combinations. The selected candidate index is signalled. For each partition mode, an IPM list is derived for each part using the same intra-inter GPM list derivation as GPM. The list is further augmented with block-vector based prediction candidates obtained from the adjacent and non-adjacent merge candidates coded in IntraTMP or IBC mode.

[0026] I. 8 Intra Template Matching

[0027] Intra template matching prediction (IntraTMP) is a intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the most similar template to the current template in a reconstructed part of the current frame and uses the corresponding block as a prediction block. The encoder then signals the usage of this mode, and the same prediction operation is performed at the decoder side.

[0028] The prediction signal is generated by matching the L-shaped, Top-only or Left-Only causal neighbour of the current block with another block in a predefined search area.

[0029] I. 9 Cross-Component Linear Model Prediction

[0030] To reduce the cross-component redundancy, a cross-component linear model (CCLM) prediction mode is used in the VVC, for which the chroma samples are predicted based on the reconstructed luma samples of the same CU by using a linear model as follows:  predC (i, j) =α·recL′ (i, j) + β      (1) where predC (i, j) represents the predicted chroma samples in a CU and recL′ (i, j) represents the downsampled reconstructed luma samples of the same CU.

[0031] The CCLM parameters (α and β) are derived with at most four neighbouring chroma samples and their corresponding down-sampled luma samples.

[0032] This parameter computation is performed as part of the decoding process, and is not just as an encoder search operation. As a result, no syntax is used to convey the α and β values to the decoder.

[0033] I. 9.1 Multiple model CCLM

[0034] Multiple model CCLM mode (MMLM) uses two models for predicting the chroma samples from the luma samples for the whole CU. In MMLM, neighbouring luma samples and neighbouring chroma samples of the current block are classified into two groups, each group is used as a training set to derive a linear model (i.e., a particular α and β are derived for a particular group) . Furthermore, the samples of the current luma block are also classified based on the same rule for the classification of neighbouring luma samples. Three MMLM model modes (MMLM_LT, MMLM_T, and MMLM_L) are allowed for choosing the neighbouring samples from left-side and above-side, above-side only, and left-side only, respectively.

[0035] Fig. 3 shows an example of classifying the neighbouring samples into two groups. Threshold is calculated as the average value of the neighbouring reconstructed luma samples. A neighbouring sample with Rec′L [x, y] <= Threshold is classified into group 1; while a neighbouring sample with Rec′L [x, y] > Threshold is classified into group 2.

[0036] I. 10 Convolutional Cross-Component Model (CCCM)

[0037] In CCCM, a convolutional model is applied to improve the chroma prediction performance. The convolutional model has 7-tap filter consisting of a 5-tap plus sign shape spatial component, a nonlinear term and a bias term. The input to the spatial 5-tap component of the filter consists of a centre (C) luma sample which is collocated with the chroma sample to be predicted and its above / north (N) , below / south (S) , left / west (W) and right / east (E) neighbours

[0038] The nonlinear term (denoted as P) is represented as power of two of the centre luma sample C and scaled to the sample value range of the content: P = (C*C + midVal ) >> bitDepth.

[0039] That is, for 10-bit content it is calculated as: P = (C*C + 512 ) >> 10.

[0040] The bias term (denoted as B) represents a scalar offset between the input and output (similarly to the offset term in CCLM) and is set to middle chroma value (512 for 10-bit content) .

[0041] Output of the filter is calculated as a convolution between the filter coefficients ci and the input values and clipped to the range of valid chroma samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B.

[0042] The filter coefficients ci are calculated by minimising MSE between predicted and reconstructed chroma samples in the reference area. The reference area which consists of 6 lines of chroma samples above and left of the PU. Reference area extends one PU width to the right and one PU height below the PU boundaries.

[0043] I. 10.1 CCCM variants

[0044] CCCM mode may feed non-downsampled luma samples, two gradient luma samples, two location information into a filter. The luma samples may be downsampled by multiple downsampling filters. The filter parameters may be derived based on reference blocks located by BV of the current block. More details can be found in “Algorithm description of Enhanced Compression Model” (for example, JVET-AH2025)

[0045] I. 11 Inter Prediction

[0046] Inter prediction techniques are briefly reviewed as follow. The details can be from JVET-T2002.

[0047] I. 11.1 Spatial candidate derivation

[0048] Spatial merge candidates are selected among candidates located in the positions depicted in Fig. 4 for a current block 410.

[0049] I. 11.2 Temporal candidate derivation

[0050] In the derivation of this temporal merge candidate for a current CU 510, a scaled motion vector is derived based on the co-located CU 520 belonging to the collocated reference picture as shown in Fig. 5. The reference picture list and the reference index to be used for the derivation of the co-located CU is explicitly signalled in the slice header. The scaled motion vector 530 for the temporal merge candidate is obtained as illustrated by the dashed line in Fig. 5, which is scaled from the motion vector 540 of the co-located CU using the POC (Picture Order Count) distances, tb and td, where tb is defined to be the POC difference between the reference picture of the current picture and the current picture and td is defined to be the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of temporal merge candidate is set equal to zero.

[0051] The position for the temporal candidate is selected between candidates C0 and C1, as depicted in Fig. 6.

[0052] I. 11.3 Non-adjacent spatial candidate

[0053] The non-adjacent spatial merge candidates as in JVET-L0399 are inserted after the TMVP in the regular merge candidate list. The pattern of spatial merge candidates is shown in Fig. 7. The distances between non-adjacent spatial candidates and current coding block are based on the width and height of current coding block.

[0054] In the present invention, methods and apparatus of using cost measurement based on a block template for reordering and / or selecting candidates from a merge candidate list comprising one or more intra prediction candidates are disclosed, where the intra prediction candidates jointly inherit the luma and chroma intra prediction information from neighbouring spatial or temporal positions / blocks to the current block. BRIEF SUMMARY OF THE INVENTION

[0055] A method and apparatus of deriving a unified merge candidate list are disclosed. According to this method, input data associated with a current block is received, wherein the input data comprises a luma component and at least one chroma component. A block template for the current block is determined. A merge candidate list comprising at least one joint luma-chroma intra prediction information entry is determined, and each entry comprises both luma intra prediction information and chroma intra prediction information. Template costs associated with entries in the merge candidate list are evaluated, wherein the template costs comprise block template costs and each of the block template costs is evaluated between the block template and a predicted block template derived using corresponding joint luma-chroma intra-prediction information. The merge candidate list is reordered based on the template costs to generate a reordered merge candidate list. At least one target joint luma-chroma intra prediction information entry is selected from the reordered merge candidate list. The current block is encoded or decoded using the at least one target joint luma-chroma intra prediction information entry.

[0056] In one embodiment, each of the block template costs comprises a luma block template cost calculated based on first distortion between the luma component of the block template and a luma intra prediction generated based on the luma intra prediction information, and a chroma block template cost calculated based on second distortion between the chroma component of the block template and a chroma intra prediction generated based on the chroma intra prediction information. In one embodiment, each of block template costs is calculated as a weighted sum of the luma block template cost and the chroma block template cost. In one embodiment, the first distortion, and the second distortion are evaluated using at least one of SAD (sum of absolute differences) or SATD (sum of absolute transformed differences) .

[0057] In one embodiment, the chroma intra prediction of the block template is generated based on the luma component of the block template and the chroma intra prediction information associated with one corresponding entry. In one embodiment, the luma component of the block template and the chroma component of the block template consist of reconstructed samples. In one embodiment, the chroma intra prediction information associated with one corresponding entry indicates a cross-component prediction mode, the chroma intra prediction of the block template is generated based on the chroma intra prediction information associated with said one corresponding entry and collocated luma samples for the chroma intra prediction of the block template. In one embodiment, the collocated luma samples are determined based on the luma component of the block template or based on a combination of the luma component of the block template and the luma intra prediction generated based on the luma intra prediction information.

[0058] In one embodiment, the block template is determined based on one or more motion vectors of the current block or a neighbouring block of the current block, or based on one or more block vectors of the current block or the neighbouring block of the current block. In one embodiment, when said one or more motion vectors or said one or more block vectors correspond to a single motion vector or a single block vector associated with a uni-directional prediction, a reference block located by the single motion vector or the single block vector is used to generate the block template through motion compensation. In one embodiment, when said one or more motion vectors or said one or more block vectors correspond to two motion vectors or two block vectors associated with a bi-directional prediction, two reference blocks located by the two motion vectors or the two block vectors are used to generate the block template through motion compensation.

[0059] In one embodiment, evaluating the template costs associated with entries in the merge candidate list comprises evaluating neighbouring-area template costs associated with the entries in the merge candidate list, and wherein each of the neighbouring-area template costs is evaluated between a neighbouring-area template and a predicted neighbouring-area template derived using corresponding joint luma-chroma intra-prediction information. In one embodiment, final template costs based on at least one of the block template costs or the neighbouring-area template costs are used to reorder the merge candidate list. In one embodiment, each of the final template costs is calculated as a weighted combination of one of the block template costs and one of the neighbouring-area template costs. In one embodiment, the method further comprises determining a normalization factor based on at least one of a size of the neighbouring-area template, a block template width, a block template height, a block template area, the block template costs, or the neighbouring-area template costs, and wherein the block template costs and the neighbouring-area template costs are adjusted according to the normalization factor. In one embodiment, the method further comprises determining a cost disparity by calculating a difference between a normalized block template cost and a normalized neighbouring-area template cost after applying the normalization factor, wherein, in response to the cost disparity exceeding a disparity threshold, weighting used for a target final template cost favours either a target block template cost or a target neighbouring-area template cost by using a lower weighting value. In one embodiment, the method further comprises determining a weight value for a target final template cost based on a target block template cost and a target neighbouring-area template cost, or based on a normalized target block template cost and a normalized target neighbouring-area template cost, and wherein the target final template cost, generated by applying the weight value to at least one of the target block template cost, the normalized target block template cost, the target neighbouring-area template cost, or the normalized target neighbouring-area template cost, is used for reordering the merge candidate list. In one embodiment, the target final template cost is increased if a cost difference between the target block template cost and the target neighbouring-area template cost, or between on the normalized target block template cost and the normalized target neighbouring-area template cost exceed a threshold. In one embodiment, the method further comprises determining a weight value for a target final template cost according to a coding mode of at least one of the neighbouring-area template or the luma component of the block template, and wherein the weight value is used to generate the target final template cost.

[0060] In one embodiment, the method further comprises determining a condition by comparing at least one of a size, a width or a height of the current block with a threshold, and wherein in response to the condition being satisfied, each of the final template costs is calculated using exclusively one of the block template costs or one of the neighbouring-area template costs.

[0061] In one embodiment, each of the block template costs is evaluated using exclusively a subset of available samples within the block template and a corresponding subset of samples within the predicted block template. In one embodiment, the method further comprises determining a condition by comparing at least one of a size, a width or a height of the current block with a threshold, and wherein in response to the condition being satisfied, evaluating each of the block template costs is performed using exclusively the subset of available samples within the block template and the corresponding subset of samples within the predicted block template.

[0062] In one embodiment, when the chroma intra prediction information associated with one corresponding entry indicates a cross-component prediction mode, a chroma intra prediction of the block template is generated based on luma samples inside the block template.BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing.

[0064] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.

[0065] Fig. 2 illustrates an example of template area with multiple reference lines.

[0066] Fig. 3 shows an example of classifying the neighbouring samples into two groups according to multiple mode CCLM.

[0067] Fig. 4 illustrates the neighbouring blocks used for deriving spatial merge candidates for VVC.

[0068] Fig. 5 illustrates an example of temporal candidate derivation, where a scaled motion vector is derived according to POC (Picture Order Count) distances.

[0069] Fig. 6 illustrates the position for the temporal candidate selected between candidates C0 and C1.

[0070] Fig. 7 illustrates an exemplary pattern of the non-adjacent spatial merge candidates.

[0071] Fig. 8A illustrates an example of the neighbouring templates comprising an above neighbouring template and a left neighbouring template, where the neighbouring templates are immediately adjacent to the current block.

[0072] Fig. 8B illustrates an example of the neighbouring templates comprising an above neighbouring template and a left neighbouring template, where the neighbouring templates are na and nb lines away from the top side and left side of the current block respectively.

[0073] Fig. 9A illustrates an example that a reference block located by the motion vector of the current block is used to generate the luma and chroma components of the block template through uni-directional motion compensation.

[0074] Fig. 9B illustrates an example that reference blocks located by the motion vectors of the current block are used to generate the luma and chroma components of the block template through bi-directional motion compensation.

[0075] Fig. 10 illustrates an example that a reference block located by the block vector of the current block is used to generate the luma and chroma components of the block template through uni-directional motion compensation.

[0076] Fig. 11 illustrates an example where only the luma samples inside the template can be used to generate the chroma prediction of the template when computing the template cost of a CCP model.

[0077] Fig. 12 illustrates an example where the size of the block template is larger than the size of the current block.

[0078] Fig. 13 illustrates a flowchart of an exemplary video coding system that uses cost measurement based on a block template for reordering and / or selecting candidates from a merge candidate list comprising one or more intra prediction candidates according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION

[0079] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.

[0080] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.

[0081] Proposed Method

[0082] For a non-intra block, such as inter or IBC block, the prediction accuracy of both luma and chroma components of the non-intra block can be improved by jointly inheriting luma and chroma intra prediction information from previous coded blocks. The luma and chroma intra prediction of the current block can be generated based on the inherited luma and chroma intra prediction information. The final prediction of the current block can be a combination of the non-intra prediction and the intra prediction.

[0083] In one embodiment, let jointIntraInfo represents a set of luma and chroma intra prediction information. One or more merge candidate lists comprising candidates with associated jointIntraInfo are constructed. At least one of the candidates in the merge candidate list is selected, and the selected jointIntraInfo is inherited by the current block. jointIntraInfo is associated with a luma coding mode and a chroma coding mode. The luma and chroma intra prediction of the current block is generated based on the luma and chroma intra prediction information of the inherited jointIntraInfo following the associated luma and chroma coding mode respectively.

[0084] In one embodiment, the final prediction can be a weighted combination of the non-intra prediction and the intra prediction generated based on the inherited jointIntraInfo. For example, for a given position (x, y) in the current block, suppose the non-intra prediction is Pnon-intra (x, y) and the intra prediction is PjointIntraInfo (x, y) , the final prediction at (x, y) can be Pfinal (x, y) =w0×Pnon-intra (x, y) +w1×PjointIntraInfo (x, y) . In another embodiment, the combination of multiple predictions can depend on many factors. For example, the combination can depend on color components (e.g., the combination is only applied to the luma component, or the combination is only applied to the chroma components) .

[0085] In one embodiment, if more than one candidate is selected from the merge candidate list, more than one intra prediction generated based on inherited jointIntraInfo can be combined with the non-intra prediction as the final prediction,  where N is the number of selected candidates.

[0086] In one embodiment, the current block is partitioned into two or more prediction subblocks, and the prediction of at least one of prediction subblocks is generated based on the inherited jointIntraInfo. The resulting partitions are not necessary to be rectangular. The prediction of one of the subblocks of the current block is generated by an inherited jointIntraInfo, and other subblocks are predicted by inter, IBC or other intra modes.

[0087] In one embodiment, a jointIntraInfo is allowed to be added into the merge candidate list if the associated luma intra coding mode of the jointIntraInfo is in a pre-defined set of allowed luma intra coding modes and / or the associated chroma intra coding mode of the jointIntraInfo is in a pre-defined set of allowed chroma intra coding modes. For the jointIntraInfo in the candidate list, the jointIntraInfo associated with each prediction candidate is also referred as a jointIntraInfo entry in this disclosure.

[0088] In one embodiment, the pre-defined set of allowed luma intra coding modes includes, but not limited to, DIMD, TIMD, TMRL, EIP, SGPM, IntraTMP, ISP (intra sub-partition) , MIP (matrix-based intra prediction) , PDP (matrix based position dependent intra prediction replacing conventional intra modes) , intra directional and non-directional prediction modes, any luma intra coding modes, or a combination thereof. More details regarding the above mentioned mode can be found in “Algorithm description of Enhanced Compression Model” (for example, JVET-AH2025) .

[0089] In one embodiment, the pre-defined set of allowed luma intra coding modes includes DIMD, TIMD, TMRL, EIP, or a combination thereof.

[0090] In one embodiment, the pre-defined set of allowed luma intra coding modes contains only one luma intra coding mode. For example, the luma intra coding mode may be DIMD, TIMD, TMRL, EIP, SGPM, or IntraTMP.

[0091] In one embodiment, the pre-defined set of allowed luma intra coding modes contains more than one luma intra coding mode.

[0092] In one embodiment, the pre-defined set of allowed chroma intra coding modes includes cross-component prediction modes (e.g., CCLM, MMLM, CCCM, variants of CCCM and CCLM) , chroma DIMD mode, or a combination thereof.

[0093] II. 1 jointIntraInfo Setting

[0094] jointIntraInfo can contain any coding mode information, any sample information, any block information, any model information, and / or any information associated with prediction generation.

[0095] In one embodiment, the luma intra prediction information of a jointIntraInfo may include, but not limited to, associated luma intra coding mode, sub-mode flags (e.g., a flag is signalled depending on if a prediction coding mode flag is true) , intra prediction directions, intra prediction blending weights, intra prediction blending modes, intra interpolation filter shape, intra interpolation filter parameters, multiple reference line indices, transform kernel types (e.g., DCT2, DST7, DCT8) , transform set, transform index (e.g., MTS or LFNST index) , transform flags, or a combination thereof.

[0096] In one embodiment, the chroma intra prediction information includes, but not limited to, associated chroma intra coding mode (e.g., CCLM, MMLM, CCCM, variants of CCCM or CCLM, chroma DIMD) , GLM pattern index, model parameters, prediction fusion modes, prediction fusion weights, classification threshold, down-sampling filters, using location terms or not, using gradient terms or not, using non-downsampled samples or not, intra prediction directions, intra prediction blending weights, intra prediction blending modes, intra interpolation filter shape, intra interpolation filter parameters, multiple reference line indices, transform kernel types (e.g., DCT2, DST7, DCT8) , transform set, transform index (e.g., MTS or LFNST index) , transform flags, or a combination thereof.

[0097] II. 1.1 Luma Intra Prediction Information

[0098] In one embodiment, if the associated luma intra coding mode of the jointIntraInfo is DIMD, the luma intra prediction information may include (a) the N intra directional prediction modes (with the highest N histogram bars) suggested by the histogram values, (b) the intra non-directional prediction modes (e.g., the planar mode or a BV-predictor) , (c) one or more histogram (bar) values for the available DIMD intra prediction modes (such as DC, planar, and / or directional prediction modes) , (d) DIMD weighting information and / or fusion or not, (e) reference line information and / or wide-angle conditions, or a combination thereof.

[0099] In one embodiment, if the associated luma intra coding mode of the jointIntraInfo is TIMD, the luma intra prediction information may include (a) the N intra prediction modes (with the smallest N TIMD costs) suggested by the TIMD costs, (b) the intra non-directional prediction modes (e.g., the planar mode or a BV-predictor) , (c) one or more TIMD cost values for the available TIMD intra prediction modes (such as DC, planar, and / or directional prediction modes) , (d) the metric of computing TIMD costs (e.g., SAD or SATD) , (e) TIMD weighting information and / or whether to use fusion or not, (f) reference line information and / or wide-angle conditions, or a combination thereof.

[0100] In one embodiment, if the associated luma intra coding mode of the jointIntraInfo is TMRL, the luma intra prediction information may include (a) the intra prediction mode (such as DC, planar, and / or directional prediction modes) jointly with the reference line information suggested by the template costs, (b) weighting information and / or fusion or not, (c) reference line information and / or wide-angle conditions, or a combination thereof.

[0101] In one embodiment, if the associated luma intra coding mode of the jointIntraInfo is EIP, the luma intra prediction information may include (a) the filter shape, (b) all or parts of the filter coefficients, (c) the template used to derive the filter coefficients, (d) multi-model flag, (e) the number of filter taps, or a combination thereof.

[0102] II. 1.2 Chroma Intra Prediction Information

[0103] In one embodiment, if the associated chroma intra coding mode of the jointIntraInfo is cross-component prediction modes, the chroma intra prediction information may include (a) the coding mode (e.g., CCLM, MMLM, CCCM, GLM, GL-CCCM, MDF-CCCM, NSCCCM, BVGCCCM, InterCCCM, and / or variants of these modes. ) , (b) the model parameters, (c) the classifying threshold if the coding mode is a multi-model cross-component mode, (d) the low pass filter flag, or a combination thereof.

[0104] In one embodiment, if the associated chroma intra coding mode of the jointIntraInfo is DIMD chroma mode, the chroma intra prediction information may include (a) one or more histogram (bar) values for the available intra prediction modes (such as DC, planar, and / or directional prediction modes) , (b) the N intra prediction modes (with the highest histogram bars) suggested by the histogram values, (c) DIMD weighting information and / or fusion or not, (d) reference line information and / or wide-angle conditions, or a combination thereof. DIMD chroma mode predicts the current chroma block using intra directional prediction modes and neighbouring reference lines. The derivation of intra directional prediction modes follows DIMD.

[0105] II. 2 Constructing a Merge Candidate List

[0106] In one embodiment, jointIntraInfo from previous coded blocks is inserted into the merge candidate list according to a pre-defined inclusion order.

[0107] In one embodiment, a maximum number of allowed candidates is imposed on the merge candidate list. The jointIntraInfo from previous coded blocks is added into the merge candidate list according to a pre-defined inclusion order until the maximum number of allowed candidates is reached.

[0108] In one embodiment, one or more candidates of spatial adjacent candidates and / or non-adjacent candidates, history candidates, temporal candidates, default candidates, combined candidates (e.g., combination of earlier candidates) , or any subset of the above-mentioned candidates containing jointIntraInfo are added into the merge candidate list according to a pre-defined inclusion order.

[0109] In one embodiment, the inclusion order can be spatial adjacent candidates → temporal candidates → spatial non-adjacent candidates → history candidate → default candidates.

[0110] II. 2.1 Spatial adjacent candidates and non-adjacent candidates

[0111] The spatial adjacent candidates are from the adjacent neighbouring blocks of the current block. The pre-defined positions and inclusion order of the spatial adjacent candidates can be the same as those of inter merge mode. The pre-defined positions of the spatial adjacent candidates can be any subset of the adjacent neighbouring blocks of the current block. For example, for adding the spatial adjacent candidates into the merge list, as in Fig. 4, the inclusion order can be A1 → B1 → A0 → B0 → B2 or B1 → A1 → B0 → A0 → B2. The non-adjacent candidates are from a search range around (but not adjacent to) the current block. The search range can be the same as the search range of non-adjacent candidates for inter merge mode. The non-adjacent candidates can be from pre-defined positions and are added into the merge list in a pre-defined inclusion order. For example, the pre-defined positions and the inclusion order are the same as those of the non-adjacent candidates of inter merge mode.

[0112] II. 2.2 History candidates

[0113] The history candidates are from a history-based buffer array. In the history-based buffer array, the jointIntraInfo of each valid previous coded block is stored, where the valid previous coded block refers to any block containing jointIntraInfo.

[0114] II. 2.3 Temporal candidates

[0115] The temporal candidates are jointIntraInfo from positions in one or more previous coded picture. The temporal candidates are obtainable when the current slice / picture is a non-intra slice / picture.

[0116] In one embodiment, the temporal candidates can be from the block at some pre-defined positions (x′, y′) of the previous coded slices / picture.

[0117] In one sub-embodiment, the positions are inside and / or outside of the corresponding area of the current encoding block.

[0118] In one sub-embodiment, the pre-defined positions can be determined based on the position, width and height of the current block.

[0119] In one sub-embodiment, the pre-defined positions can be determined based on the position, and some pre-defined fixed x-y distances.

[0120] In one embodiment, the pre-defined positions and the inclusion order are the same as those of regular inter merge mode.

[0121] In one embodiment, the previous coded pictures are among the pictures in the reference lists.

[0122] In one embodiment, the previous coded pictures are the same pictures as the collocated picture of the regular inter merge mode.

[0123] II. 3 Signalling

[0124] In one embodiment, an on / off flag is signalled to indicate whether the current block jointly inherits the luma and chroma intra prediction information from previous coded blocks to improve the non-intra prediction or not. The flag can be signalled per CU / CB, per PU, per TU / TB, or per colour component. A high-level syntax can be signalled in SPS, PPS, PH or SH to indicate whether the proposed method is allowed for the current sequence, picture, or slice.

[0125] In one embodiment, if the current block jointly inherits the luma and chroma intra prediction information from previous coded blocks to improve the non-intra prediction, the one or more candidates selected from the merge candidate list can be explicitly signalled or implicitly determined. For example, the one or more candidates with the smallest potential prediction error (i.e., cost) are implicitly selected. For example, explicit indexes are signalled to indicate one or more candidates in the reordered merge candidate list as the selected candidates. The index can be signalled using a truncated unary code, an Exp-Golomb code, or a fix length code. The cost calculation and / or list reordering may depend on methods described in Section II.4

[0126] II. 4 Reordering of the Merge Candidate List

[0127] The potential prediction errors (i.e., the costs) for candidates with associated jointIntraInfo entries in the merge candidate list can be evaluated. The potential prediction errors associated with a jointIntraInfo entry can be used to reorder the jointIntraInfo entry in the merge candidate list to reduce the signalling overhead. The potential prediction errors associated with the jointIntraInfo entries may be used to select one or more jointIntraInfo entries from the merge candidate list. The jointIntraInfo entries with the smallest potential prediction error may be selected.

[0128] In one embodiment, the potential prediction error associated with a jointIntraInfo entry can be evaluated by using a template matching method. A template can be an area in previous coded region / pictures or an area containing samples derived based on previous coded region / pictures. The prediction of the template is generated based on each jointIntraInfo entry. The template cost is the distortion / difference between the prediction of the template (generated based on jointIntraInfo) and the template.

[0129] In one embodiment, the luma template cost is the distortion / difference between the luma intra prediction (generated based on luma intra prediction information of jointIntraInfo) and the luma component of the template. The chroma template cost is the distortion / difference between the chroma intra prediction (generated based on chroma intra prediction information of jointIntraInfo) and the chroma component of the template. The template cost can be a weighted sum of the luma template cost and the chroma template cost. That is, template cost = a*luma template cost + b*chroma template cost. For example, the weighting setting may be equal weighting. For example, the weighting setting may be (a, b) = (4, 4) . For example, the weighting setting may be unequal weighting. For example, the weighting setting may be (a, b) = (6, 2) .

[0130] In one embodiment, the luma and chroma components of the template can be the luma reconstruction samples and the chroma reconstruction samples respectively.

[0131] In one embodiment, the distortion metric can be SAD (sum of absolute difference) or SATD (sum of absolute transformed difference) .

[0132] In one embodiment, as shown in Fig 8A, the template can be a neighbouring area 820 above the current block 810, a neighbouring area 830 left to the current block 810, or a combination thereof. The size of above neighbouring template is wa×ha, and the size of left neighbouring template is wb×hb. For example, wa can be the width of the current block, hb can be the height of the current block. For example, ha and wb can be a pre-defined positive number such as 1, 2, 3, 4, 6, 8, …etc. For another example, ha and wb can be adaptive determined. ha and wb can be determined based on the size / width / height of the current block. As shown in Fig. 8B, the luma and / or chroma neighbouring templates can be the lines close to the current block but not immediately adjacent to the current block. na and nb are positive integers. For example, na and nb equal to one.

[0133] In one embodiment, the luma neighbouring template can be different from the chroma neighbouring template. For example, the luma neighbouring templates are the lines immediately adjacent to the current block, and the chroma neighbouring templates are the lines close to the current block but not immediately adjacent to the current block. For example, the luma neighbouring templates are as depicted in Fig. 8A, with ha and wb equal to 1. The chroma neighbouring templates are as depicted in Fig. 8B, with ha, wb, na and nb equal to 1.

[0134] In one embodiment, the chroma neighbouring templates are determined based on the associated chroma coding mode. If the associate chroma coding mode is DIMD chroma mode, the chroma neighbouring templates are the lines immediately adjacent to the current block, if the associated chroma coding mode is a cross-component prediction mode, the chroma neighbouring templates are the lines close to the current block but not immediately adjacent to the current block.

[0135] In one embodiment, when one or more motion vectors can be identified at the current block, instead of using neighbouring area as template, the template can be a block determined based on the one or more motion vectors. In another embodiment, when one or more block vectors can be identified at the current block, instead of using neighbouring area as template, the template can be a block determined based on the one or more block vectors. The block size can be the same as the current block. The block template cost of a jointIntraInfo can be computed as the distortion / difference between the block template and the prediction generated based on the jointIntraInfo. The potential prediction error of a jointIntraInfo can be evaluated based on the block template cost instead of neighbouring-area template cost.

[0136] In one embodiment, when the current block is coded in an inter mode, the one or more motion vectors used to determine the block template can be the motion vectors of the current block, as depicted in Fig. 9A and Fig. 9B. The reference blocks located by the motion vectors can be used to determine the luma and chroma components of the block template. The luma and chroma components of the block template are computed by motion compensation based on the motion vectors of the current block. When the current block utilises uni-direction prediction as in Fig. 9A, the reference block located by the motion vector is used to generate the luma and chroma components of the block template through motion compensation. When the current block utilises bi-directional prediction (i.e., the current block has more than one motion vector) , the luma and chroma components of the block template can be generated by bi-prediction motion compensation as shown in Fig. 9B.

[0137] In one embodiment, when the current block is coded in IBC mode, the one or more block vectors used to determine the block template can be the block vectors of the current block, as depicted in Fig. 10. The reference blocks located by the block vectors can be used to determine the luma and chroma components of the block template by motion compensation. The luma and chroma components of the block template are computed by motion compensation based on the block vectors of the current block. When the current block utilises uni-direction prediction, the reference block located by the block vector is used to generate the luma and chroma components of the block template through motion compensation. When the current block utilizes bi-directional prediction (i.e., the current block has more than one block vector) , the luma and chroma components of the block template can be generated by bi-prediction motion compensation.

[0138] In one embodiment, the chroma prediction of the block template is generated based on the luma component of the template and the chroma intra prediction information of the jointIntraInfo.

[0139] In one embodiment, the luma and chroma components of the block template are the reconstruction samples.

[0140] In one embodiment, when the associated chroma intra coding mode of the jointIntraInfo is a cross-component prediction mode, the chroma intra prediction of the block template is generated based on chroma intra prediction information of the jointIntraInfo and the collocated luma samples. For example, the collocated luma samples can be the luma component of the block template. For another example, the collocated luma sample can be the combination of the luma component of the block template (may be generated based on motion compensation) and the luma intra prediction generated based on the luma intra prediction information of the jointIntraInfo.

[0141] In one embodiment, the one or more motion vectors used to determine the block template can be the motion vectors of a neighbouring block of the current block. The luma and chroma components of the block template are computed by motion compensation based on the motion vectors of the neighbouring block.

[0142] In one embodiment, the one or more block vectors used to determine the block template can be the block vectors of a neighbouring block of the current block. The luma and chroma components of the block template are computed by motion compensation based on the block vectors of the neighbouring block.

[0143] In one embodiment, when evaluating the potential prediction errors associated with the jointIntraInfo entries in the merge candidate list, if there are M jointIntraInfo entries in the list, neighbouring-area template cost is computed for each of the M jointIntraInfo entries. For N (N is smaller than M) jointIntraInfo entries with lowest N neighbouring-area template costs, block template cost is computed for each of the N jointIntraInfo. For example, the M jointIntraInfo entries can be reordered first based on the neighbouring-area template cost, then the N jointIntraInfo entries are reordered again based on the block template cost. For another example, one or more jointIntraInfo entries associated with the smallest block template costs are selected from N candidate jointIntraInfo entries as the final candidates.

[0144] In one embodiment, the final template cost used to evaluate the potential prediction error is computed based on the block template cost and the neighbouring-area template cost. The final template cost can be a weighted combination of block template cost and the neighbouring-area template cost.

[0145] In one sub-embodiment, when the size, width, or height of the current block is greater than, greater than or equal to, smaller than, or smaller than or equal to a threshold, the final template cost is computed based on the neighbouring-area template cost only.

[0146] In one sub-embodiment, when the size, width, or height of the current block is greater than, greater than or equal to, smaller than, or smaller than or equal to a threshold, the final template cost is computed based on the block template cost only.

[0147] In one embodiment, a flag can be signalled to indicate the type of template used to compute the template cost. For example, if the flag is true, the block template cost is used. Otherwise, the neighbouring-area template cost is used. For another example, if the flag is true, the neighbouring-area template cost is used. Otherwise, the block template cost is used. The flag can be signalled at a block level, CU level, PU level, TU level, CTU row level, area level, slice level, picture level, sequence level, or a high-level syntax.

[0148] In one sub-embodiment, the block template cannot be used when the size / width / height of the current block is greater than, greater than or equal to, smaller than, or smaller than or equal to a threshold. If the size, width, or height of the current block is greater than, greater than or equal to, smaller than, smaller than or equal to a threshold, the flag is not signalled and the neighbouring-area template is used to compute template cost.

[0149] In one sub-embodiment, the neighbouring-area template cannot be used when the size, width, or height of the current block is greater than, greater than or equal to, smaller than, or smaller than or equal to a threshold. If the size, width, or height of the current block is greater than, greater than or equal to, smaller than, or smaller than or equal to a threshold, the flag is not signalled and the block template is used to compute template cost.

[0150] In one embodiment, the block template cost is computed based on partial samples in the block template only. For example, the block template cost is computed based on sub-sampled positions of the template only. For another example, the block template cost is computed based on partial region of the block template only.

[0151] In one embodiment, the block template cost is computed only based on partial samples in the block template if the size, width, or height of the current block is greater than, greater than or equal to, smaller than, or smaller than or equal to a threshold.

[0152] In another embodiment, if the chroma coding mode associated with the jointIntraInfo is a cross-component prediction mode, only the luma samples inside the template can be used to generate the chroma prediction of the template. For example, a 5-tap convolutional filter (in a cross shape) is used in CCCM, assume the current template size is M×N, after applying the 5-tap convolutional filter to the luma samples (maybe downsampled according to the colour format) , the generated chroma prediction size is (M-2) × (N-2) . For example, as shown in Fig. 11, given an 8×8 template, the 5-tap convolutional filter is applied to the luma component samples (or the combination of the luma component of the template and the luma intra prediction) from the centre positions (1, 1) to (6, 6) , and the generated chroma prediction size is (8-2) × (8-2) = 6×6, which is the region marked in dash lines in Fig. 11.

[0153] In one embodiment, if the chroma coding mode associated with the jointIntraInfo is a cross-component prediction mode and if luma samples outside of the template is needed to generate the chroma prediction of the template, the luma samples outside of the template are generated by using padding. For example, for a 5-tap convolutional filter (in a cross shape) is used in CCCM, one line of luma samples above, left to, right to, or below the template are generated by padding.

[0154] In one embodiment, the size of the block template is larger than the size of the current block. The size of the block template may consider the needed reference samples for generating luma and chroma intra prediction. As depicted in Fig. 12, the block template includes the collocated block, located by the block vector, and the outer dotted region. The block template cost is computed based on the collocated block region only. To generate luma intra prediction for the collocated block region, above and left reference lines may be needed. For example, if the associated luma coding mode is DIMD or TIMD, one above and one left reference line are needed. Hence the block template is enlarged to include the above and left reference lines. To generate chroma intra prediction for the collocated block region, above, left, right, or below reference lines may be needed. For example, if the associated chroma coding mode is CCCM, one above, left, right, and below reference lines are need. Hence, the block template is enlarged to include the above, left, right and below reference lines.

[0155] In one embodiment, the final template cost used to evaluate the potential prediction error is computed based on the block template cost and the neighbouring-area template cost. The final template cost is a weighted combination of the block template cost and the neighbouring-area template cost according to a weighting setting with normalization. The normalization depends on the neighbouring-area template size, block width, block height, block area, the block template cost, the neighbouring-area template cost, or any combination thereof. For example, the neighbouring-area template cost is calculated based on the neighbouring-area including above neighbouring-area (block width x neighbouring-area template size, or wa×ha) and left neighbouring-area (neighbouring-area template size x block height, or wb×hb) , and the block template cost is calculated based on the block template (block width x block height) . In one way, the normalization is applied to increase the neighbouring-area template cost and then the increased neighbouring-area template cost is combined with the block template cost. For example, the neighbouring-area template cost is multiplied by a factor proportional to (block width x block height)  /  (wa×ha + wb×hb) . In another way, the normalization is applied to decrease the block template cost and then the decreased block template cost is combined with the neighbouring-area template cost. For example, the block template cost is multiplied by a factor proportional to (wa×ha + wb×hb)  /  (block width x block height) . In another example, the weight used to do weighted combination or normalization can be signalled at block level, CU level, PU level, TU level, CTU row level, area level, slice level, picture level, sequence level, or a high-level syntax.

[0156] In another example, if the normalized block template cost and the normalized neighbouring-area template cost are very different, it implies that the content of block template and neighbouring template are very different. The weight of final cost can favour either the block template cost or the neighbouring-area template cost. For example, the weight of final cost favours the one with smaller template cost. In another example, the weight of the block template cost and the neighbouring-area template cost can be derived from the information of the block template cost and the neighbouring-area template cost. For example, if the ratio of the normalized block template cost and the normalized neighbouring-area template cost is between a first range, the weight is equal to Value1; if the ratio if between a second range, the weight is equal to Value2.

[0157] In another embodiment, the cost / normalized cost of the block template and the neighbouring-area template can be used to derive a trust value or a weight for the final cost. The modified final cost (e.g., multiplied by the trust value or the weight) are used to do the reordering. For example, if the normalized cost of the block template and the neighbouring-area template are very different, it implies that the content of block template and neighbouring template are very different. The cost of this candidate is not trustable. The final cost of this candidate should be increased, or the priority of this candidate should be moved backward in the list.

[0158] In one embodiment, the final template cost used to evaluate the potential prediction error is computed based on the block template cost and the neighbouring-area template cost. The final template cost is a weighted combination of the block template cost and the neighbouring-area template cost according to a weighting setting based on coding modes for the neighbouring-area template and / or the luma component of the block template. For example, when the neighbouring-area template is associated with intra coding modes, the weight for the neighbouring-area template cost is higher than the weight for the block template cost. When the luma component of the block template is associated with intra coding modes, the weight for the block template cost is higher than the weight for the neighbouring-area template cost. In another example, the weight used to do weighted combination can be signalled at a block level, CU level, PU level, TU level, CTU row level, area level, slice level, picture level, sequence level, or a high-level syntax.

[0159] Any of the foregoing proposed methods can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter / intra / prediction module of an encoder, and / or an inter / intra / prediction module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module.

[0160] With reference to the exemplary encoder and decoder in Fig. 1A and Fig. 1B, the proposed methods of merge list derivation with candidate diversity can be implemented in an Intra / Inter coding module (e.g. Intra Pred. 150 / MC 152 in Fig. 1B) in a decoder or an Intra / Inter coding module in an encoder (e.g. Intra Pred. 110 / Inter Pred. in Fig. 1A) . However, the decoder or encoder may also use additional processing unit to implement the required cross-component prediction processing. While the Intra / Inter Pred. units are shown as individual processing units, they may correspond to executable software or firmware codes for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .

[0161] Fig. 13 illustrates a flowchart of an exemplary video coding system that uses cost measurement based on a block template for reordering and / or selecting candidates from a merge candidate list comprising one or more intra prediction candidates according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to this method, input data associated with a current block is received in step 1310, wherein the input data comprises a luma component and at least one chroma component. A block template for the current block is determined in step 1320. A merge candidate list comprising at least one joint luma-chroma intra prediction information entry is determined in step 1330, and each entry comprises both luma intra prediction information and chroma intra prediction information. Template costs associated with entries in the merge candidate list are evaluated in step 1340, wherein the template costs comprise block template costs and each of the block template costs is evaluated between the block template and a predicted block template derived using corresponding joint luma-chroma intra-prediction information. The merge candidate list is reordered based on the template costs to generate a reordered merge candidate list in step 1350. At least one target joint luma-chroma intra prediction information entry is selected from the reordered merge candidate list in step 1360. The current block is encoded or decoded using the at least one target joint luma-chroma intra prediction information entry in step 1370.

[0162] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arrange the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0163] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.

[0164] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.

[0165] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1.A method of coding colour pictures, the method comprising:receiving input data associated with a current block, wherein the input data comprises a luma component and at least one chroma component;determining a block template for the current block;determining a merge candidate list comprising at least one joint luma-chroma intra prediction information entry, and each entry comprises both luma intra prediction information and chroma intra prediction information;evaluating template costs associated with entries in the merge candidate list, wherein the template costs comprise block template costs and each of the block template costs is evaluated between the block template and a predicted block template derived using corresponding joint luma-chroma intra-prediction information;reordering the merge candidate list based on the template costs to generate a reordered merge candidate list;selecting at least one target joint luma-chroma intra prediction information entry from the reordered merge candidate list; andencoding or decoding the current block using the at least one target joint luma-chroma intra prediction information entry.2.The method of Claim 1, wherein each of the block template costs comprises a luma block template cost calculated based on first distortion between the luma component of the block template and a luma intra prediction generated based on the luma intra prediction information, and a chroma block template cost calculated based on second distortion between the chroma component of the block template and a chroma intra prediction generated based on the chroma intra prediction information.3.The method of Claim 2, wherein each of block template costs is calculated as a weighted sum of the luma block template cost and the chroma block template cost.4.The method of Claim 2, wherein the first distortion, and the second distortion are evaluated using at least one of SAD (sum of absolute differences) or SATD (sum of absolute transformed differences) .5.The method of Claim 2, wherein the chroma intra prediction of the block template is generated based on the luma component of the block template and the chroma intra prediction information associated with one corresponding entry.6.The method of Claim 2, wherein the luma component of the block template and the chroma component of the block template consist of reconstructed samples.7.The method of Claim 2, wherein when the chroma intra prediction information associated with one corresponding entry indicates a cross-component prediction mode, the chroma intra prediction of the block template is generated based on the chroma intra prediction information associated with said one corresponding entry and collocated luma samples for the chroma intra prediction of the block template.8.The method of Claim 7, wherein the collocated luma samples are determined based on the luma component of the block template or based on a combination of the luma component of the block template and the luma intra prediction generated based on the luma intra prediction information.9.The method of Claim 1, wherein the block template is determined based on one or more motion vectors of the current block or a neighbouring block of the current block, or based on one or more block vectors of the current block or the neighbouring block of the current block.10.The method of Claim 9, wherein when said one or more motion vectors or said one or more block vectors correspond to a single motion vector or a single block vector associated with a uni-directional prediction, a reference block located by the single motion vector or the single block vector is used to generate the block template through motion compensation.11.The method of Claim 9, wherein when said one or more motion vectors or said one or more block vectors correspond to two motion vectors or two block vectors associated with a bi-directional prediction, two reference blocks located by the two motion vectors or the two block vectors are used to generate the block template through motion compensation.12.The method of Claim 1, wherein evaluating the template costs associated with entries in the merge candidate list comprises evaluating neighbouring-area template costs associated with entries in the merge candidate list, and wherein each of the neighbouring-area template costs is evaluated between a neighbouring-area template and a predicted neighbouring-area template derived using corresponding joint luma-chroma intra-prediction information.13.The method of Claim 12, wherein final template costs based on at least one of the block template costs or the neighbouring-area template costs are used to reorder the merge candidate list.14.The method of Claim 13, wherein each of the final template costs is calculated as a weighted combination of one of the block template costs and one of the neighbouring-area template costs.15.The method of Claim 14, further comprising determining a normalization factor based on at least one of a size of the neighbouring-area template, a block template width, a block template height, a block template area, the block template costs, or the neighbouring-area template costs, and wherein the block template costs and the neighbouring-area template costs are adjusted according to the normalization factor.16.The method of Claim 15, further comprising determining a cost disparity by calculating a difference between a normalized block template cost and a normalized neighbouring-area template cost after applying the normalization factor, wherein, in response to the cost disparity exceeding a disparity threshold, weighting used for a target final template cost favours either a target block template cost or a target neighbouring-area template cost by using a lower weighting value.17.The method of Claim 15, further comprising determining a weight value for a target final template cost based on a target block template cost and a target neighbouring-area template cost, or based on a normalized target block template cost and a normalized target neighbouring-area template cost, and wherein the target final template cost, generated by applying the weight value to at least one of the target block template cost, the normalized target block template cost, the target neighbouring-area template cost, or the normalized target neighbouring-area template cost, is used for reordering the merge candidate list.18.The method of Claim 17, wherein the target final template cost is increased if a cost difference between the target block template cost and the target neighbouring-area template cost, or between on the normalized target block template cost and the normalized target neighbouring-area template cost exceed a threshold.19.The method of Claim 15, further comprising determining a weight value for a target final template cost according to a coding mode of at least one of the neighbouring-area template or the luma component of the block template, and wherein the weight value is used to generate the target final template cost.20.The method of Claim 13, further comprising determining a condition by comparing at least one of a size, a width or a height of the current block with a threshold, and wherein in response to the condition being satisfied, each of the final template costs is calculated using exclusively one of the block template costs or one of the neighbouring-area template costs.21.The method of Claim 1, wherein each of the block template costs is evaluated using exclusively a subset of available samples within the block template and a corresponding subset of samples within the predicted block template.22.The method of Claim 21, further comprising determining a condition by comparing at least one of a size, a width or a height of the current block with a threshold, and wherein in response to the condition being satisfied, evaluating each of the block template costs is performed using exclusively the subset of available samples within the block template and the corresponding subset of samples within the predicted block template.23.The method of Claim 1, wherein when the chroma intra prediction information associated with one corresponding entry indicates a cross-component prediction mode, a chroma intra prediction of the block template is generated based on luma samples inside the block template.24.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block, wherein the input data comprises a luma component and at least one chroma component;determine a block template for the current block;determine a merge candidate list comprising at least one joint luma-chroma intra prediction information entry, and each entry comprises both luma intra prediction information and chroma intra prediction information;evaluate template costs associated with entries in the merge candidate list, wherein the template costs comprise block template costs and each of the block template costs is evaluated between the block template and a predicted block template derived using corresponding joint luma-chroma intra-prediction information;reorder the merge candidate list based on the template costs to generate a reordered merge candidate list;select at least one target joint luma-chroma intra prediction information entry from the reordered merge candidate list; andencode or decode the current block using the at least one target joint luma-chroma intra prediction information entry.