Method and apparatus of regression-based blending for improving intra prediction fusion in video coding system
The regression-based method for chroma intra fusion in video coding systems addresses inefficiencies by utilizing position-dependent information to enhance chroma prediction, resulting in improved coding efficiency and quality.
Patent Information
- Application Number
- PCT/CN2024/140205
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2024-12-18
- Publication Date
- 2025-07-17
AI Technical Summary
Existing video coding systems face challenges in efficiently predicting chroma components due to insufficient utilization of position-dependent information in chroma intra fusion modes, leading to suboptimal coding performance.
A regression-based method is employed to derive fusion weights for chroma predictors using position-dependent information, allowing for improved blending of non-linear model predictors with reconstructed samples, thereby enhancing chroma intra prediction accuracy.
This approach improves coding efficiency by refining chroma prediction, leading to better compression performance and quality in video coding systems.
Smart Images

Figure CN2024140205_17072025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS OF REGRESSION-BASED BLENDING FOR IMPROVING INTRA PREDICTION FUSION IN VIDEO CODING SYSTEMCROSS REFERENCE TO RELATED APPLICATIONSThe present invention is a Non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 618,932, filed on January 9, 2024. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTIONThe present invention relates to video coding using chroma intra fusion mode. In particular, the present invention relates to schemes to derive blending weights using a regression-based method with location information for chroma intra fusion mode.BACKGROUND AND RELATED ARTVersatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter-Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, are provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.The decoder, as shown in Fig. 1B, can use similar or portion of the same functional blocks as the encoder except for Transform 118 and Quantization 120 since the decoder only needs Inverse Quantization 124 and Inverse Transform 126. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.According to VVC, an input picture is partitioned into non-overlapped square block regions referred as CTUs (Coding Tree Units) , similar to HEVC. Each CTU can be partitioned into one or multiple smaller size coding units (CUs) . The resulting CU partitions can be in square or rectangular shapes. Also, VVC divides a CTU into prediction units (PUs) as a unit to apply prediction process, such as Inter prediction, Intra prediction, etc.Partitioning of the CTUs Using a Tree StructureIn HEVC, a CTU is split into CUs by using a quaternary-tree (QT) structure denoted as coding tree to adapt to various local characteristics. The decision whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the leaf CU level. Each leaf CU can be further split into one, two or four PUs according to the PU splitting type. Inside one PU, the same prediction process is applied and the relevant information is transmitted to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU splitting type, a leaf CU can be partitioned into transform units (TUs) according to another quaternary-tree structure similar to the coding tree for the CU. One of key feature of the HEVC structure is that it has the multiple partition conceptions including CU, PU, and TU.In VVC, a quadtree with nested multi-type tree using binary and ternary splits segmentation structure replaces the concepts of multiple partition unit types, i.e. it removes the separation of the CU, PU and TU concepts except as needed for CUs that have a size too large for the maximum transform length, and supports more flexibility for CU partition shapes. In the coding tree structure, a CU can have either a square or rectangular shape. A coding tree unit (CTU) is first partitioned by a quaternary tree (a.k.a. quadtree) structure. Then the quaternary tree leaf nodes can be further partitioned by a multi-type tree structure. As shown in Fig. 2, there are four splitting types in multi-type tree structure, vertical binary splitting (SPLIT_BT_VER 210) , horizontal binary splitting (SPLIT_BT_HOR 220) , vertical ternary splitting (SPLIT_TT_VER 230) , and horizontal ternary splitting (SPLIT_TT_HOR 240) . The multi-type tree leaf nodes are called coding units (CUs) , and unless the CU is too large for the maximum transform length, this segmentation is used for prediction and transform processing without any further partitioning. This means that, in most cases, the CU, PU and TU have the same block size in the quadtree with nested multi-type tree coding block structure. The exception occurs when maximum supported transform length is smaller than the width or height of the colour component of the CU.Fig. 3 shows a CTU divided into multiple CUs with a quadtree and nested multi-type tree coding block structure, where the bold block edges represent quadtree partitioning and the remaining edges represent multi-type tree partitioning. The quadtree with nested multi-type tree partition provides a content-adaptive coding tree structure comprised of CUs. In VVC, the maximum supported luma transform size is 64×64 and the maximum supported chroma transform size is 32×32. When the width or height of the CB is larger the maximum transform width or height, the CB is automatically split in the horizontal and / or vertical direction to meet the transform size restriction in that direction.Intra Mode Coding with 67 Intra Prediction ModesTo capture the arbitrary edge directions presented in natural video, the number of directional intra modes in VVC is extended from 33, as used in HEVC, to 65. The new directional modes not in HEVC are depicted as dotted arrows in Fig. 4, and the planar and DC modes remain the same. These denser directional intra prediction modes apply for all block sizes and for both luma and chroma intra predictions.In VVC, several conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for the non-square blocks.In HEVC, every intra-coded block has a square shape and the length of each of its side is a power of 2. Thus, no division operations are required to generate an intra-predictor using DC mode. In VVC, blocks can have a rectangular shape that necessitates the use of a division operation per block in the general case. To avoid division operations for DC prediction, only the longer side is used to compute the average for non-square blocks.Cross-Component Linear Model (CCLM) PredictionTo reduce the cross-component redundancy, a cross-component linear model (CCLM) prediction mode is used in the VVC, for which the chroma samples are predicted based on the reconstructed luma samples of the same CU by using a linear model as follows:predC (i, j) =α·recL′ (i, j) + β (1)where predC (i, j) represents the predicted chroma samples in a CU and recL′ (i, j) represents the downsampled reconstructed luma samples of the same CU.The CCLM parameters (α and β) are derived with at most four neighbouring chroma samples and their corresponding down-sampled luma samples. Suppose the current chroma block dimensions are W×H, then W’a nd H’a re set as–W’ = W, H’ = H when LM_LA mode is applied;–W’ =W + H when LM_Amode is applied;–H’ = H + W when LM_L mode is applied.The above neighbouring positions are denoted as S [0, -1 ] …S [W’ -1, -1 ] and the left neighbouring positions are denoted as S [-1, 0 ] …S [-1, H’ -1 ] . Then the four samples are selected as-S [W’ / 4, -1 ] , S [3 *W’ / 4, -1 ] , S [-1, H’ / 4 ] , S [-1, 3 *H’ / 4 ] when LM mode is applied and both above and left neighbouring samples are available;-S [W’ / 8, -1 ] , S [3 *W’ / 8, -1 ] , S [5 *W’ / 8, -1 ] , S [7 *W’ / 8, -1 ] when LM-A mode is applied or only the above neighbouring samples are available;-S [-1, H’ / 8 ] , S [-1, 3 *H’ / 8 ] , S [-1, 5 *H’ / 8 ] , S [-1, 7 *H’ / 8 ] when LM-L mode is applied or only the left neighbouring samples are available.The four neighbouring luma samples at the selected positions are down-sampled and compared four times to find two larger values: x0A and x1A, and two smaller values: x0B and x1B. Their corresponding chroma sample values are denoted as y0A, y1A, y0B and y1B. Then xA, xB, yA and yB are derived as:Xa= (x0A + x1A +1) >>1;Xb= (x0B+ x1B+1) >>1;Ya= (y0A + y1A +1) >>1;Yb= (y0B+ y1B+1) >>1. (2)Finally, the linear model parameters α and β are obtained according to the following equations.β=Yb-α·Xb (4)Fig. 5 shows an example of the location of the left and above samples and the sample of the current block involved in the LM_LA mode. Fig. 5 shows the relative sample locations of N × N chroma block 510, the corresponding 2N × 2N luma block 520 and their neighbouring samples (shown as filled circles) .The division operation to calculate parameter α is implemented with a look-up table. To reduce the memory required for storing the table, the diff value (difference between maximum and minimum values) and the parameter α are expressed by an exponential notation. For example, diff is approximated with a 4-bit significant part and an exponent. Consequently, the table for 1 / diff is reduced into 16 elements for 16 values of the significand as follows:DivTable [] = {0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0 } (5)This would have a benefit of both reducing the complexity of the calculation as well as the memory size required for storing the needed tables.Besides the above template and left template can be used to calculate the linear model coefficients together, they also can be used alternatively in the other 2 LM modes, called LM_A, and LM_L modes.In LM_Amode, only the above template is used to calculate the linear model coefficients. To get more samples, the above template is extended to (W+H) samples. In LM_L mode, only left template are used to calculate the linear model coefficients. To get more samples, the left template is extended to (H+W) samples.In LM_LA mode, left and above templates are used to calculate the linear model coefficients.To match the chroma sample locations for 4: 2: 0 video sequences, two types of down-sampling filter are applied to luma samples to achieve 2 to 1 down-sampling ratio in both horizontal and vertical directions. The selection of down-sampling filter is specified by a SPS level flag. The two down-sampling filters are as follows, which are corresponding to “type-0” and “type-2” content, respectively.RecL′ (i, j) = [recL (2i-1, 2j-1) +2·recL (2i-1, 2j-1) +recL (2i+1, 2j-1) +recL (2i-1, 2j) +2·recL (2i, 2j) +recL (2i+1, 2j) +4] >>3 (6) recL′ (i, j) =recL (2i, 2j-1) +recL (2i-1, 2j) +4·recL (2i, 2j) +recL (2i+1, 2j) +recL (2i, 2j+1) +4] >>3 (7)Note that only one luma line (general line buffer in intra prediction) is used to make the down-sampled luma samples when the upper reference line is at the CTU boundary.This parameter computation is performed as part of the decoding process, and is not just as an encoder search operation. As a result, no syntax is used to convey the α and β values to the decoder.For chroma intra mode coding, a total of 8 intra modes are allowed for chroma intra mode coding. Those modes include five traditional intra modes and three cross-component linear model modes (LM_LA, LM_A, and LM_L) .Multiple Model CCLM (MMLM)In the JEM (J. Chen, E. Alshina, G. J. Sullivan, J. -R. Ohm, and J. Boyce, Algorithm Description of Joint Exploration Test Model 7, document JVET-G1001, ITU-T / ISO / IEC Joint Video Exploration Team (JVET) , Jul. 2017) , multiple model CCLM mode (MMLM) is proposed for using two models for predicting the chroma samples from the luma samples for the whole CU. In MMLM, neighbouring luma samples and neighbouring chroma samples of the current block are classified into two groups, each group is used as a training set to derive a linear model (i.e., a particular α and β are derived for a particular group) . Furthermore, the samples of the current luma block are also classified based on the same rule for the classification of neighbouring luma samples. Three MMLM model modes (MMLM_LA, MMLM_T, and MMLM_L) are allowed for choosing the neighbouring samples from left-side and above-side, above-side only, and left-side only, respectively.Fig. 6 shows an example of classifying the neighbouring samples into two groups. Threshold is calculated as the average value of the neighbouring reconstructed luma samples. A neighbouring sample with Rec′L [x, y] <= Threshold is classified into group 1; while a neighbouring sample with Rec′L [x, y] >Threshold is classified into group 2.Convolutional Cross-Component Model (CCCM)In CCCM, a convolutional model is applied to improve the chroma prediction performance. The convolutional model has 7-tap filter consist of a 5-tap plus sign shape spatial component, a nonlinear term and a bias term. The input to the spatial 5-tap component of the filter consists of a centre (C) luma sample which is collocated with the chroma sample to be predicted and its above / north (N) , below / south (S) , left / west (W) and right / east (E) neighbours as illustrated in Fig. 7.The nonlinear term (denoted as P) is represented as power of two of the centre luma sample C and scaled to the sample value range of the content:P = (C*C + midVal ) >> bitDepth.That is, for 10-bit content it is calculated as:P = (C*C + 512 ) >> 10The bias term (denoted as B) represents a scalar offset between the input and output (similarly to the offset term in CCLM) and is set to middle chroma value (512 for 10-bit content) .Output of the filter is calculated as a convolution between the filter coefficients ci and the input values and clipped to the range of valid chroma samples:predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6BThe filter coefficients ci are calculated by minimising MSE between predicted and reconstructed chroma samples in the reference area. Fig. 8 illustrates the reference area which consists of 6 lines of chroma samples above and left of the PU. Reference area extends one PU width to the right and one PU height below the PU boundaries. Area is adjusted to include only available samples. The extensions to the area shown in blue are needed to support the “side samples” of the plus shaped spatial filter and are padded when in unavailable areas.The MSE minimization is performed by calculating autocorrelation matrix for the luma input and a cross-correlation vector between the luma input and chroma output. Autocorrelation matrix is LDL decomposed and the final filter coefficients are calculated using back-substitution. The process follows roughly the calculation of the ALF filter coefficients in ECM, however LDL decomposition was chosen instead of Cholesky decomposition to avoid using square root operations.Position Dependent Intra Prediction Combination (PDPC)In VVC, the results of intra prediction of DC, planar and several angular modes are further modified by a position dependent intra prediction combination (PDPC) method. PDPC is an intra prediction method which invokes a combination of the boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signalling: planar, DC, intra angles less than or equal to horizontal, and intra angles greater than or equal to vertical and less than or equal to 80. If the current block is BDPCM mode or MRL index is larger than 0, PDPC is not applied.The prediction sample pred (x′, y′) is predicted using an intra prediction mode (e.g. DC, planar, angular) and a linear combination of reference samples according to the Equation 8 as follows:pred (x′, y′) =Clip(0, (1<<BitDeph) -1, (wL×R-1, y′+wT×Rx′, -1+ (64-wL-wT) ×pred (x′, y′) +32)>>6) (8)where Rx, -1, R-1, y represent the reference samples located at the top and left boundaries of current sample (x,y) , respectively.If PDPC is applied to DC, planar, horizontal, and vertical intra modes, additional boundary filters are not needed, as required in the case of HEVC DC mode boundary filter or horizontal / vertical mode edge filters. PDPC process for DC and Planar modes is identical. For angular modes, if the current angular mode is HOR_IDX or VER_IDX, left or top reference samples is not used, respectively. The PDPC weights and scaling factors are dependent on prediction modes and the block sizes. PDPC is applied to the block with both width and height greater than or equal to 4.Figs. 9A-D illustrate the definition of reference samples (Rx, -1 and R-1, y) for PDPC applied over various prediction modes, where Fig. 9A corresponds to the diagonal top-right mode, Fig. 9B corresponds to the diagonal bottom-left mode, Fig. 9C corresponds to the adjacent diagonal top-right mode and Fig. 9D corresponds to the adjacent diagonal bottom-left mode. The prediction sample pred (x′, y′) is located at (x′, y′) within the prediction block. As an example, the coordinate x of the reference sample Rx, -1 is given by: x=x′+y′+1, and the coordinate y of the reference sample R-1, y is similarly given by: y=x′+y′+1 for the diagonal modes. For the other angular mode, the reference samples Rx, -1and R-1, ycould be located in fractional sample position. In this case, the sample value of the nearest integer sample location is used.Intra Prediction FusionThis intra prediction method derives predicted samples as a weighted combination of multiple predictors generated from different reference lines. In this process multiple intra predictors are generated and then fused by weighted averaging. The process of deriving the predictors to be used in the fusion process is described as follows:1) For angular intra prediction modes including the single mode case of TIMD and DIMD, the proposed method derives intra prediction by weighted combination of intra predictions obtained from multiple reference lines represented as pfusion=w0pline+w1pline+1, where pline is the intra prediction from the default reference line and pline+1 is the prediction from the line above the default reference line. The weights are set as w0=3 / 4 and w1=1 / 4.2) For TIMD mode with blending, pline is used for the first mode (w0=1, w1=0) and pline+1 is used for the second mode (w0=0, w1=1) .3) For DIMD mode with blending, the number of predictors selected for a weighted average is increased from 3 to 6.Intra prediction fusion method is applied to luma blocks when angular intra mode has non-integer slope (required reference samples interpolation) and the block size is greater than 16, it is used with MRL and not applied for ISP coded blocks. In one method, PDPC is applied for the intra prediction mode using the nearest line to the current block reference line.Intra Template MatchingIntra template matching prediction (IntraTMP) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the most similar template to the current template in a reconstructed part of the current frame and uses the corresponding block as a prediction block. The encoder then signals the usage of this mode, and the same prediction operation is performed at the decoder side.The prediction signal is generated by matching the L-shaped, Top-only or Left-Only causal neighbour of the current block with another block in a predefined search area in Fig. 10. There are 6 predefined search areas, i.e., R1 to R6 in Fig. 10 which contain the reconstructed samples from the top and left CTUs as well as part of the reconstructed samples within the current CTU that are located above, left, bottom-left and top-right to the current block.Sum of absolute differences (SAD) is used as a cost function.A given search order of the 6 regions is utilized, i.e., R4, R5, R6, R1, R2, and R3. Within each region, the decoder constructs a candidate list of up to “19” template matching block vectors, where the candidates (i.e., template matching block vectors) are ranked in ascending order according to the template cost (SAD) . The following modes are supported:a. Single predictor: A single predictor is selected from the candidate list.b. Fusion of multiple predictors: multiple predictors are blended to derive the final prediction block. The blending weights are either computed from the template matching cost of each predictor, or with Wiener-filter based weight derivation method.c. Sub-pel precision: When single predictor is used, sub-pel precision can be used with 1 / 2-pel precision, 1 / 4-pel precision and 3 / 4-pel precision, each with 8 possible directions.d. Linear filter model: A linear filter can be learned between the reference template and current template and the linear model can be applied to the reference block. This mode can be used for single predictor when sub-pel precision is not used.The dimensions of all regions (SearchRange_w, SearchRange_h) are set proportional to the block dimension (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is:SearchRange_w = min (64, a*BlkW) ,SearchRange_h = min (64, a*BlkH) ,where ‘a’ is a constant that controls the gain / complexity trade-off. In practice, ‘a’ is equal to 5.To speed-up the template matching process, the search range of all search regions is subsampled by a factor of 3. After finding the best match, a refinement process is performed. The refinement is done via a second template matching search around the best match with a reduced range.The Intra template matching tool is enabled for CUs with size less than or equal to 64 in width and height. This maximum CU size for Intra template matching is configurable.The Intra template matching prediction mode is signalled at CU level through a dedicated flag when DIMD is not used for the current CU.Matrix weighted Intra Prediction (MIP)Matrix weighted intra prediction (MIP) method is a newly added intra prediction technique into VVC. For predicting the samples of a rectangular block of width W and height H, matrix weighted intra prediction (MIP) takes one line of H reconstructed neighbouring boundary samples left of the block and one line of W reconstructed neighbouring boundary samples above the block as input. If the reconstructed samples are unavailable, they are generated as it is done in the conventional intra prediction. The generation of the prediction signal is based on the following three steps, which are averaging, matrix vector multiplication and linear interpolation as shown in Fig. 11. One line of H reconstructed neighbouring boundary samples 1112 left of the block and one line of W reconstructed neighbouring boundary samples 1110 above the block are shown as dot-filled small squares. After the averaging process, the boundary samples are down-sampled to top boundary line 1114 and left boundary line 1124. The down-sampled samples are provided to the matric-vector multiplication unit 1120 to generate the down-sampled prediction block 1130. An interpolation process is then applied to generate the prediction block 1140.Decoder side Intra Mode Derivation (DIMD)When DIMD is applied, up to five intra modes are derived from the reconstructed neighbour samples, and those five predictors are combined with the planar mode predictor with the weights derived from the histogram of gradients as described in JVET-O0449. The division operations in weight derivation are performed utilizing the same lookup table (LUT) based integerization scheme used by the CCLM. For example, the division operation in the orientation calculation:Orient=Gy / Gxis computed by the following LUT-based scheme:x = Floor (Log2 (Gx ) ) norm Diff = ( (Gx<< 4 ) >> x ) &15x += (3 + (normDiff ! = 0 ) ? 1 : 0 ) Orient = (Gy* (DivSigTable [normDiff ] | 8 ) + (1<< (x-1 ) ) ) >> xwhereDivSigTable
[0016] = {0, 7, 6, 5 , 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0 } .For a block of size W×H, the weight for each of the five derived modes is modified if the one the above or left histogram magnitudes is twice larger than the other one. In this case, the weights are location dependent and computed as follows.If the above histogram is twice the left, then:If the left histogram is twice the above, then:where wDimdi is the unmodified uniform weight of the DIMD selected as in JVET-O0449, Δi is pre-defined and set to 10.Derived intra modes are included into the primary list of intra most probable modes (MPM) , so the DIMD process is performed before the MPM list is constructed. The primary derived intra mode of a DIMD block is stored with a block and is used for MPM list construction of the neighbouring blocks.Finally, note the region of neighbouring reconstructed samples used for computing the histogram of gradients is modified compared to JVET-O0449 method, depending on reconstructed samples availability. The region of decoded reference samples of current WxH luma CB is extended towards the above-right side if available, up to W additional columns. It is extended towards the bottom-left side if available, up to H additional rows.Fusion for Template-based Intra Mode Derivation (TIMD)For each intra prediction mode in MPMs, as well as the wide-angle modes if the above-right and / or bottom-left reference samples are available, SATD between the prediction and reconstruction samples of the template is calculated. First two intra prediction modes with the minimum SATD are selected as the TIMD modes. These two TIMD modes are fused with the weights after applying PDPC process, and such weighted intra prediction is used to code the current CU. Position dependent intra prediction combination (PDPC) is included in the derivation of the TIMD modes.The costs of the two selected modes are compared with a threshold, in the test the cost factor of 2 is applied as follows:costMode2 < 2*costMode1.If this condition is true, the fusion is applied, otherwise the only mode1 is used.Weights of the modes are computed from their SATD costs as follows:weight1 = costMode2 / (costMode1+ costMode2) weight2 = 1 -weight1The division operations are conducted using the same lookup table (LUT) based integerization scheme used by the CCLM.Regression-based GPM BlendingRegression-based GPM blending mode is designed as an additional GPM implicit mode, where the two integer blending matrices (W0 and W1) are derived from the template (1 line above, 1 column left) . The blending matrices are modelled as an affine linear function of the sample positions (x, y) in the current CU:W0(x, y) = a. x + b. y + c and W1 (x, y) = 1 -W0 (x, y) .The parameters (a, b, c) are derived from the reference template using the same solver (MSE minimization) as the one used for CCCM. A list of pair of candidates is built from the regular GPM candidates and re-ordered with the template cost.The GPM implicit mode is signalled by a CU-level flag (gpm_implicit_flag) . If gpm_implicit_flag is true, a merge-idx is coded to signal the pair of GPM candidates to be used. If gpm_implicit_flag is false, the regular GPM syntax elements are signalled.In the present invention, methods and apparatus to improve the coding performance of chroma intra fusion mode for video coding systems using regression-based method with position-dependent information are disclosed.BRIEF SUMMARY OF THE INVENTIONA method and apparatus for video coding are disclosed. According to this method, input data associated with a current block comprising a current luma block and a current chroma block is received, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side. A non-LM (Linear Model) predictor is generated for the current chroma block. One or more chroma predictors for the current chroma block are determined. One or more fusion weights for the non-LM predictor and said one or more chroma predictors are derived by using a regression-based process with training data comprising reconstructed samples and position-dependent information. A final chroma predictor is derived by blending the non-LM predictor and said one or more chroma predictors using said one or more fusion weights. The current chroma block is encoded or decoded by using coding information comprising the final chroma predictor.In one embodiment, said one or more chroma predictors are generated according to a single model or multiple models. In one embodiment, said one or more chroma predictors are generated according to a single model or multiple models inherited from luma component, and chroma component reuses luma models with or without refinement.In one embodiment, the current block comprises a Cb block and a Cr block, and wherein the Cb block and the Cr block use a joint single model or joint multiple models. In one embodiment, the current block comprises a Cb block and a Cr block, and wherein neighbouring Cb and Cr components are mapped into a transformed domain for the regression-based process to calculate cost. In one embodiment, the regression-based process uses multiple reconstruction sample lines.In one embodiment, it is determined by explicit signalling or implicitly determining whether to apply position dependent chroma fusion mode to process the current chroma block. The position dependent chroma fusion mode comprises said using the regression-based process with the training data comprising the neighbouring reconstructed samples and the position-dependent information. In one embodiment, the position dependent chroma fusion mode is treated as an additional mode to existing chroma fusion modes.In one embodiment, it is determined by explicit signalling or implicitly determining whether to blend the non-LM predictor and one or more chroma predictors to derive the final chroma predictor. The blending is performed by using one or more fusion weighs.In one embodiment, it is determined by explicit signalling or implicitly determining whether to use the coding information comprising the final chroma predictor to encode or decode the current chroma block.In one embodiment, the training data comprises down-samples luma reconstructed samples or non-downsampled luma reconstructed samples.In one embodiment, fusion weights or blending lines for one or more chroma predictors are adaptively changes according to regression-derived model resulted from the regression-based process.BRIEF DESCRIPTION OF THE DRAWINGSFig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing.Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.Fig. 2 illustrates examples of a multi-type tree structure corresponding to vertical binary splitting (SPLIT_BT_VER) , horizontal binary splitting (SPLIT_BT_HOR) , vertical ternary splitting (SPLIT_TT_VER) , and horizontal ternary splitting (SPLIT_TT_HOR) .Fig. 3 shows an example of a CTU divided into multiple CUs with a quadtree and nested multi-type tree coding block structure, where the bold block edges represent quadtree partitioning and the remaining edges represent multi-type tree partitioning.Fig. 4 shows the 67 intra prediction modes as adopted by the VVC video coding standard.Fig. 5 shows an example of the location of the left and above samples and the sample of the current block involved in the LM_LA mode.Fig. 6 shows an example of classifying the neighbouring samples into two groups according to multiple mode CCLM.Fig. 7 illustrates an example of spatial part of the convolutional filter.Fig. 8 illustrates an example of reference area with paddings used to derive the filter coefficients.Figs. 9A-D illustrate examples of the definition of reference samples for PDPC applied over various prediction modes, where Fig. 9A corresponds to the diagonal top-right mode, Fig. 9B corresponds to the diagonal bottom-left mode, Fig. 9C corresponds to the adjacent diagonal top-right mode and Fig. 9D corresponds to the adjacent diagonal bottom-left mode.Fig. 10 illustrates the concept of IntraTMP where the L-shape template is used to search the areas R1 to R6 to get a displacement vector (i.e., block vector) .Fig. 11 illustrates an example of processing flow for Matrix weighted intra prediction (MIP) .Fig. 12 illustrates an example of target sample and neighbouring sample usage during training sample collection.Fig. 13 illustrates an example of spatial weighting WS and distance weighting WD in the template region.Fig. 14 illustrates an example of adaptive blending lines or adaptive blending weightings in position-dependent chroma fusion mode.Fig. 15 illustrates an example of certain region blending in position-dependent chroma fusion mode.Fig. 16 illustrates an example of spatial coordinates of four corners inside the current block.Fig. 17 illustrates a flowchart of an exemplary video coding system that derives chroma intra fusion mode using regression-based method with position-dependent information according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTIONIt will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.Several methods are proposed to improve the luma / chroma intra coding prediction accuracy or coding performance. In one embodiment, whether to allow or apply the following proposed methods can depend on SPS / PPS / SH / PH syntax, or CTU or CU / PU / TU level syntax or semantic. In another embodiment, whether to allow or apply the following proposed methods can also depend on implicit conditions. For example, the following proposed methods can be applied based on block width, block height, or block area. In another embodiment, the following proposed improvements of regression-based blending can be conditionally applied to partial positions / sub-block of the current block. In another embodiment, the term “block” in this invention can refer to TU / TB, CU / CB, PU / PB, pre-defined region, or CTU / CTB. Any combination of the following proposed methods in this invention can be applied.Fusion of Chroma Intra Prediction Modes with Regression-Based BlendingIn existing chroma intra prediction modes fusion, three fusion options are available to blend non-LM chroma predictor and MMLM chroma predictor / down-sampled luma reconstructed samples. In regression-based GPM blending, position information is utilized to assist the refinement in the predictor blending process. A new chroma intra prediction mode fusion method, namely, position-dependent chroma fusion mode, is disclosed, which exploits the additional position information to assist model training and model application.Joint or Separate Training for Luma and Chroma ComponentsIn the proposed method, chroma components may use a joint single or joint multiple models when the current block performs chroma intra prediction mode with regression-based blending or position-dependent chroma fusion mode.In one embodiment, single model or multiple models can be derived or inherited from luma component and chroma components reuse the luma derived models with or without refinement.Example (1) : Extrapolation filter-based intra prediction (EIP) mode derives single model or multiple models for luma CBs. For chroma components, they can reuse the EIP derived single model or multiple models from luma CBs. To better fit chroma components, EIP luma derived models can be refined, such as, combining different model parameters together, removing model parameters or adjusting model parameters.Example (2) : Chroma components may inherit partial luma derived models, and according to Cb and Cr information, some model parameters may be modified.In another embodiment, chroma components Cb and Cr derive a joint single mode or joint multiple models.Example (1) : Down-sampled or non-down-sampled luma reconstructed samples are used to derive the model, and neighbouring chroma reconstructed samples are used as the training target in the regression process. When calculating the training cost, both neighbouring Cb and Cr reconstructed samples are used, and the objective function is to minimize the cost by considering Cb and Cr jointly. The jointly derived single model or multiple models are used in the current block’s prediction stage.Example (2) : Down-sampled or non-down-sampled luma reconstructed samples are used to derive the model, and neighbouring chroma reconstructed samples are used as the training target in regression process. When calculating training cost, both neighbouring Cb and Cr reconstructed samples are firstly converted into another domain using some transforms. Then, the objective function is to minimize the cost by considering transformed Cb and Cr jointly. The jointly derived single model or multiple models are used in the current block’s prediction stage. At the current block, Cb and Cr components may be converted into another domain using some transforms or omit this process.Reconstruction Lines and Samples Collection and Samples Usage in Model TrainingIn existing regression-based GPM blending, one reconstruction sample line is used for the model training and template-matching, respectively. In the proposed method, multiple reconstruction sample lines and different reconstruction sample collection methods during the training are disclosed.In one embodiment, multiple reconstruction sample lines are used in the regression-based method.In another embodiment, multiple reconstruction sample lines are used to train one singe model or multiple models.In another embodiment, subsampled reconstruction samples are collected in the model training, and subsampled reconstruction samples are collected according to some pre-defined or classification rules. For example:(1) Reconstruction line: …╳○○○╳○○○╳…In example (1) , ○ and ╳ are uniformly distributed reconstruction samples. Reconstruction samples marked as ○ are used in the model training and reconstruction samples marked as ╳ are used in cost calculation.(2) Reconstruction line: …╳○○○○○○○○○╳…In example (2) , ○ and ╳ are non-uniformly distributed reconstruction samples. The distribution of ○ and ╳ can be determined according to some pre-defined positions or classification results. The reconstruction samples marked as ○ are used in the model training and reconstruction samples marked as ╳are used in the cost calculation.(3) Reconstruction line: …╳○○△△○○○╳…In example (3) , ○, △ and ╳ are non-uniformly distributed reconstruction samples. The distribution of △, ○ and ╳ can be determined according to some pre-defined positions or classification results. The reconstruction sample ○ is used in one model training, the reconstruction sample marked as △ are used in another model training and reconstruction samples marked as ╳ are used in the cost calculation.In another embodiment, when training one target sample, its neighbouring samples can be considered or omitted in the model training, as shown in Fig. 12. Sample at location A is the target sample, and samples B, C, D and E can be considered in the model training.In another embodiment, to derive one predictor sample, one sample or multiple samples can be used.In another embodiment, in the model training, spatial weighting WS or distance weighting WD is considered, as shown in the Fig. 13. The spatial weighting WS and the distance weighting WD is the weighting distribution along the horizontal direction and vertical direction in the template region, respectively. For example, inside the template region, the spatial weighting is decreased from far to near with respect to the current block boundary, and the distance weight is increased from left-side to right-side of the above template region or from top-side to bottom-side of the left template region.In another embodiment, during the model training, if only the left-template or the top-template is available, either spatial weighting WS or distance weighting WD is considered.In another embodiment, during the model training, if only left-template or the top-template is available, both spatial weighting WS and distance weighting WD are considered.In another embodiment, during the model training, if only the left-template or the top-template available, neither spatial weighting WS nor distance weighting WD is considered.In another embodiment, the right-top reconstruction sample and the left-bottom reconstruction sample in the reconstruction sample line are padded during model training. For example, in Fig. 12, the reconstruction sample A, B, C and D are padded during the model training.In another embodiment, the predictors utilized in model training can be implicitly determined. For example, non-LM chroma predictor, LM chroma predictor, or down-sampled luma reconstructed samples can be used. In another example, multiple models can be trained using different combinations of the mentioned predictors and the costs of different trained models are calculated. The model with the smallest cost is selected, and the corresponding predictors in position-dependent chroma fusion mode is used.Regression-based Adaptive Weightings and LinesIn the proposed method, based on the regression-derived model, adaptive multiple blending lines and blending weightings are generated and applied to the chroma intra prediction mode with regression-based blending. Four corner positions can be used to test the weighting strength change inside the current block. To constrain the weighting strength, weighting clipping can be considered.In one embodiment, blending weightings or blending lines of two predictors in position-dependent chroma fusion mode are adaptively changed according to the regression-derived model, as shown in Fig. 14. For example, the weightings for predictor A increases from left part to right part and weightings for predictor B increases from right part to left part.In another embodiment, the position-dependent chroma fusion mode is only performed at certain region inside predictor, similar to partial blending, as shown in Fig. 15. For example, only the left region of predictor performs the position-dependent chroma fusion to refine the predictor boundary.In another embodiment, similar to planar mode with the regression-based blending, two temporary predictors (i.e., horizontal planar prediction and vertical planar prediction) can be firstly derived using two model weightings and the final predictor is generated using the average of the two temporary predictors.In another embodiment, four corner positions (0, 0) , (width, 0) , (0, height) and (width, height) are used to check the generated weightings for position-dependent chroma fusion mode.Example (1) : If the weighting difference between any two corner positions is smaller than or equal to a threshold, the regression-based model is not applied to the current block’s position-dependent chroma fusion mode.Example (2) : If weighting of any two corner positions is smaller than a threshold, the regression-based model is not applied to current block’s position-dependent chroma fusion mode.In another embodiment, the generated weighting for position-dependent chroma fusion mode is clipped to be in a range. For example, the clipping range is within [1 / 32, 31 / 32] for weightings of non-LM chroma predictors, weightings of down-sampled luma reconstructed samples, or weightings of LM chroma predictors.Syntax DesignThe new position-dependent chroma fusion mode is proposed and can be explicitly or implicitly signalled.In one embodiment, the position-dependent chroma fusion mode is an additional mode before or after existing single linear model (single LM) mode, or multiple linear model (multiple LM) mode in chroma intra prediction modes. An example of the signalling syntax is showed in the following table.Table 1: Syntax example of position-dependent chroma fusion modeIn another embodiment, the position information is further considered in single LM or multiple LM cases. Single LM with position information or multiple LM with position information can be regarded as another kind of position-dependent chroma intra prediction mode. An example of syntax design is listed in the following table.Table 2: Syntax example of single LM with position information and multiple LM with position informationIn another embodiment, the position information can be considered in single LM or multiple LM cases by implicit method. For instance, single LM with position information or without position information can be implicitly determined. For example, four corner positions can be used to test generated model weightings, and if any of the two corner position weightings difference is smaller than a threshold, position information is not involved in single LM scenario. Similar concept may be extended to multiple LM scenario. In this example, there is no additional syntax overhead and an example of syntax design is listed in the following table.Table 3: Syntax example of existing modes with and without position informationIntra Template Matching (IntraTMP) with Regression-Based BlendingIn the latest ECM, IntraTMP has two fusion modes. They both use five candidates to generate the final prediction. One of them uses a regression-based method to implicitly derive the fusion weights, and the other uses the template matching cost (e.g., SAD or SATD) of each candidate to implicitly derive the fusion weights. The following is an example of the regression model used in current IntraTMP fusion mode. P0 to P4 are five to-be-fused candidates, and c0 to c4 are their coefficients:pred=c0P0+c1P1+c2P2+c3P3+c4P4.To improve the performance of IntraTMP fusion modes, some techniques from regression-based GPM blending can be used in IntraTMP fusion modes.In one embodiment, the weight of each candidate in the IntraTMP fusion regression model can be related to the horizontal location information x and / or vertical location information y as shown in the following equation, and the optimal parameters are derived using the reconstructed samples in the neighbouring template region:pred= (a0x+b0y+c0) P0+ (a1x+b1y+c1) P1+... (a4x+b4y+c4) P4.In one embodiment, the new IntraTMP fusion mode with location information can be a sub-mode of IntraTMP fusion mode. If the regression-based IntraTMP fusion mode is selected, one additional flag is signalled to indicate whether the original regression model without location information is used or the new regression model with location information is used.In another embodiment, if the regression-based IntraTMP fusion mode is selected, a template matching method can be used to implicitly determine which model is used for the current block. First, two regression models are derived, and then two regression models can be applied to the neighbouring template region and calculate the template matching cost between the reconstructed samples and the predicted samples. The predictor of the regression model with smaller template matching cost will be selected as the final predictor.In one embodiment, the number of template lines used to derive the regression-based IntraTMP fusion model and the number of template lines used to calculate the IntraTMP template matching cost can be different.SGPM with Regression-Based BlendingThe main difference between SGPM mode and GPM mode is that both predictors in SGPM are intra-coded. The techniques from regression-based GPM blending can be easily used in SGPM mode to improve the coding efficiency.In one embodiment, a regression-based model with location information can be used to determine the SGPM blending weight. For example, w0 is the blending weight of the first predictor which can be derived from the following regression model using horizontal location information and / or vertical location information, and w1 is the blending weight of the second predictor which can be derived by the equation w1=1-w0:w0=ax+by+c.In another embodiment, a regression-based model with location information can be used to determine the SGPM blending weight. The blending weight w0 of the first predictor can be derived from the following regression model using horizontal location information and / or vertical location information, and the blending weight w1 of the second predictor can be derived from another regression model. Two regression models can be solved together:w0=a0x+b0y+c0,w1=a1x+b1y+c1.In one embodiment, if the current block is coded by SGPM mode, a regression-based SGPM flag can be signalled to indicate whether the proposed regression-based model with location information or the original SGPM mode is used for determining the blending weight.In one embodiment, regression-based GPM with the weight clipping method can be used in the proposed regression-based SGPM mode. After obtaining the blending weights from the regression model, a clipping operation can be further applied to the output of the regression model to clip the values within a pre-defined range. For example, if the blending weight has 5 bits for the fractional part, the clipping range can be 0 to 32, or 1 to 31.In one embodiment, regression-based GPM mode with single-line template region can be used to train the SGPM regression model.In another embodiment, multiple-line template region can be used to train the SGPM regression model.Any of the foregoing proposed methods of chroma intra fusion mode using a regression-based method with location information can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter / intra / prediction module of an encoder, and / or an inter / intra / prediction module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module.With reference to the encoder and decoder in Fig. 1A and Fig. 1B, any of the proposed methods of chroma intra fusion mode using a regression-based method with location information can be implemented in an inter / intra / prediction / transform module (e.g. Intra Pred. 110 in Fig. 1A) of an encoder, and / or an inter / intra / prediction / transform module (e.g. Intra Pred. 150 in Fig. 1B) of a decoder.Fig. 17 illustrates a flowchart of an exemplary video coding system that derives chroma intra fusion mode using regression-based method with position-dependent information according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side and / or decoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to this method, input data associated with a current block comprising a current luma block and a current chroma block is received in step 1710, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side. A non-LM (Linear Model) predictor is generated for the current chroma block in step 1720. One or more chroma predictors for the current chroma block are determined in step 1730. One or more fusion weights for the non-LM predictor and said one or more chroma predictors are derived by using a regression-based process with training data comprising reconstructed samples and position-dependent information in step 1740. A final chroma predictor is derived by blending the non-LM predictor and said one or more chroma predictors using said one or more fusion weights in step 1750. The current chroma block is encoded or decoded by using coding information comprising the final chroma predictor in step 1760.The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1.A method of coding colour pictures or video, the method comprising:receiving input data associated with a current block comprising a current luma block and a current chroma block, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side; andderiving a position dependent chroma fusion mode for the current block, wherein said processing the current chroma block using the position dependent chroma fusion mode comprises:generating a non-LM (Linear Model) predictor for the current chroma block;determining one or more chroma predictors for the current chroma block; andderiving one or more fusion weights for the non-LM predictor and said one or more chroma predictors by using a regression-based process with training data comprising reconstructed samples and position-dependent information;deriving a final chroma predictor by blending the non-LM predictor and said one or more chroma predictors using said one or more fusion weights; andencoding or decoding the current chroma block by using coding information comprising the final chroma predictor.2.The method of Claim 1, wherein said one or more chroma predictors are generated according to a single model or multiple models.3.The method of Claim 1, wherein said one or more chroma predictors are generated according to a single model or multiple models inherited from luma component, and chroma component reuses luma models with or without refinement.4.The method of Claim 1, wherein the current block comprises a Cb block and a Cr block, and wherein the Cb block and the Cr block use a joint single model or joint multiple models.5.The method of Claim 1, wherein the current block comprises a Cb block and a Cr block, and wherein neighbouring Cb and Cr components are mapped into a transformed domain for the regression-based process to calculate cost.6.The method of Claim 1, wherein the regression-based process uses multiple reconstruction sample lines.7.The method of Claim 1, wherein whether to derive the position dependent chroma fusion mode, whether to derive the final chroma predictor, or whether to encode or decode the current chroma block by using coding information comprising the final chroma predictor is determined by explicit signalling or determined implicitly.8.The method of Claim 7, wherein the position dependent chroma fusion mode is treated as an additional mode to existing chroma fusion modes.9.The method of Claim 1, wherein the training data comprises down-samples luma reconstructed samples or non-downsampled luma reconstructed samples.10.The method of Claim 1, wherein fusion weights or blending lines for one or more chroma predictors are adaptively changes according to regression-derived model resulted from the regression-based process.11.An apparatus for coding colour pictures or video, the apparatus comprising one or more electronic circuits or processors arranged to:receive input data associated with a current block comprising a current luma block and a current chroma block, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side; andderiving a position dependent chroma fusion mode for the current block, wherein said processing the current chroma block using the position dependent chroma fusion mode comprises:generate a non-LM (Linear Model) predictor for the current chroma block; anddetermining one or more chroma predictors for the current chroma block;derive one or more fusion weights for the non-LM predictor and said one or more chroma predictors by using a regression-based process with training data comprising reconstructed samples and position-dependent information;derive a final chroma predictor by blending the non-LM predictor and said one or more chroma predictors using said one or more fusion weights; andencode or decode the current chroma block by using coding information comprising the final chroma predictor.
Citation Information
Patent Citations
Method and apparatus for image encoding, and method and apparatus for image decoding
US20210321111A1
Cross component prediction of chroma samples
US20230403397A1
Method and apparatus for cross component linear model for inter prediction in video coding system
WO2023116716A1