Method and apparatus of regression-based blending for improving inter prediction in video coding system
A regression-based method for deriving adaptive blending weights addresses the inefficiencies in predicting chroma components in video coding, improving prediction accuracy and efficiency by optimizing the blending of motion-compensated and cross-component predictors.
Patent Information
- Application Number
- PCT/CN2025/071297
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-17
AI Technical Summary
Existing video coding systems face challenges in efficiently predicting chroma components using motion-compensated and cross-component predictors, lacking flexibility and accuracy in blending weights, which affects coding performance.
Implement a regression-based method to derive adaptive blending weights using location information for motion-compensated and cross-component predictors, enhancing the inter prediction process in video coding systems.
Improves coding performance by optimizing the blending of motion-compensated and cross-component predictors, leading to enhanced prediction accuracy and efficiency in video coding.
Smart Images

Figure CN2025071297_17072025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS OF REGRESSION-BASED BLENDING FOR IMPROVING INTER PREDICTION IN VIDEO CODING SYSTEMCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present invention is a Non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 618,931, filed on January 9, 2024, U.S. Provisional Patent Application No. 63 / 618,934, filed on January 9, 2024 and U.S. Provisional Patent Application No. 63 / 632,566, filed on April 11, 2024. The U.S. Provisional Patent Applications are hereby incorporated by reference in their entireties.FIELD OF THE INVENTION
[0002] The present invention relates to video coding using interCCP tool that blends motion-compensated chroma predictor with a cross-component predictor. In particular, the present invention relates to schemes to derive adaptive blending weights using a regression-based method with location information. BACKGROUND AND RELATED ART
[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter-Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, are provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.
[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.
[0006] The decoder, as shown in Fig. 1B, can use similar or portion of the same functional blocks as the encoder except for Transform 118 and Quantization 120 since the decoder only needs Inverse Quantization 124 and Inverse Transform 126. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.
[0007] According to VVC, an input picture is partitioned into non-overlapped square block regions referred as CTUs (Coding Tree Units) , similar to HEVC. Each CTU can be partitioned into one or multiple smaller size coding units (CUs) . The resulting CU partitions can be in square or rectangular shapes. Also, VVC divides a CTU into prediction units (PUs) as a unit to apply prediction process, such as Inter prediction, Intra prediction, etc.
[0008] Partitioning of the CTUs Using a Tree Structure
[0009] In HEVC, a CTU is split into CUs by using a quaternary-tree (QT) structure denoted as coding tree to adapt to various local characteristics. The decision whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the leaf CU level. Each leaf CU can be further split into one, two or four PUs according to the PU splitting type. Inside one PU, the same prediction process is applied and the relevant information is transmitted to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU splitting type, a leaf CU can be partitioned into transform units (TUs) according to another quaternary-tree structure similar to the coding tree for the CU. One of key feature of the HEVC structure is that it has the multiple partition conceptions including CU, PU, and TU.
[0010] In VVC, a quadtree with nested multi-type tree using binary and ternary splits segmentation structure replaces the concepts of multiple partition unit types, i.e. it removes the separation of the CU, PU and TU concepts except as needed for CUs that have a size too large for the maximum transform length, and supports more flexibility for CU partition shapes. In the coding tree structure, a CU can have either a square or rectangular shape. A coding tree unit (CTU) is first partitioned by a quaternary tree (a.k.a. quadtree) structure. Then the quaternary tree leaf nodes can be further partitioned by a multi-type tree structure. As shown in Fig. 2, there are four splitting types in multi-type tree structure, vertical binary splitting (SPLIT_BT_VER 210) , horizontal binary splitting (SPLIT_BT_HOR 220) , vertical ternary splitting (SPLIT_TT_VER 230) , and horizontal ternary splitting (SPLIT_TT_HOR 240) . The multi-type tree leaf nodes are called coding units (CUs) , and unless the CU is too large for the maximum transform length, this segmentation is used for prediction and transform processing without any further partitioning. This means that, in most cases, the CU, PU and TU have the same block size in the quadtree with nested multi-type tree coding block structure. The exception occurs when maximum supported transform length is smaller than the width or height of the colour component of the CU.
[0011] Fig. 3 shows a CTU divided into multiple CUs with a quadtree and nested multi-type tree coding block structure, where the bold block edges represent quadtree partitioning and the remaining edges represent multi-type tree partitioning. The quadtree with nested multi-type tree partition provides a content-adaptive coding tree structure comprised of CUs. In VVC, the maximum supported luma transform size is 64×64 and the maximum supported chroma transform size is 32×32. When the width or height of the CB is larger the maximum transform width or height, the CB is automatically split in the horizontal and / or vertical direction to meet the transform size restriction in that direction.
[0012] Intra Mode Coding with 67 Intra Prediction Modes
[0013] To capture the arbitrary edge directions presented in natural video, the number of directional intra modes in VVC is extended from 33, as used in HEVC, to 65. The new directional modes not in HEVC are depicted as dotted arrows in Fig. 4, and the planar and DC modes remain the same. These denser directional intra prediction modes apply for all block sizes and for both luma and chroma intra predictions.
[0014] In VVC, several conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for the non-square blocks.
[0015] In HEVC, every intra-coded block has a square shape and the length of each of its side is a power of 2. Thus, no division operations are required to generate an intra-predictor using DC mode. In VVC, blocks can have a rectangular shape that necessitates the use of a division operation per block in the general case. To avoid division operations for DC prediction, only the longer side is used to compute the average for non-square blocks.
[0016] Cross-Component Linear Model (CCLM) Prediction
[0017] To reduce the cross-component redundancy, a cross-component linear model (CCLM) prediction mode is used in the VVC, for which the chroma samples are predicted based on the reconstructed luma samples of the same CU by using a linear model as follows: predC (i, j) =α·recL′ (i, j) + β (1) where predC (i, j) represents the predicted chroma samples in a CU and recL' (i, j) represents the downsampled reconstructed luma samples of the same CU.
[0018] The CCLM parameters (α and β) are derived with at most four neighbouring chroma samples and their corresponding down-sampled luma samples. Suppose the current chroma block dimensions are W×H, then W’ and H’ are set as –W’ = W, H’ = H when LM_LA mode is applied; –W’ =W + H when LM_A mode is applied; –H’ = H + W when LM_L mode is applied.
[0019] The above neighbouring positions are denoted as S [0, -1] …S [W’ -1, -1] and the left neighbouring positions are denoted as S [-1, 0] …S [-1, H’ -1] . Then the four samples are selected as -S [W’ / 4, -1] , S [3 *W’ / 4, -1] , S [-1, H’ / 4] , S [-1, 3 *H’ / 4] when LM mode is applied and both above and left neighbouring samples are available; -S [W’ / 8, -1] , S [3 *W’ / 8, -1] , S [5 *W’ / 8, -1] , S [7 *W’ / 8, -1] when LM-A mode is applied or only the above neighbouring samples are available; -S [-1, H’ / 8] , S [-1, 3 *H’ / 8] , S [-1, 5 *H’ / 8] , S [-1, 7 *H’ / 8] when LM-L mode is applied or only the left neighbouring samples are available.
[0020] The four neighbouring luma samples at the selected positions are down-sampled and compared four times to find two larger values: x0A and x1A, and two smaller values: x0B and x1B. Their corresponding chroma sample values are denoted as y0A, y1A, y0B and y1B. Then xA, xB, yA and yB are derived as: Xa= (x0A + x1A +1) >>1; Xb= (x0B + x1B +1) >>1; Ya= (y0A + y1A +1) >>1; Yb= (y0B + y1B +1) >>1 (2)
[0021] Finally, the linear model parameters α and β are obtained according to the following equations. β=Yb-α·Xb (4)
[0022] Fig. 5 shows an example of the location of the left and above samples and the sample of the current block involved in the LM_LA mode. Fig. 5 shows the relative sample locations of N × N chroma block 510, the corresponding 2N × 2N luma block 520 and their neighbouring samples (shown as filled circles) .
[0023] The division operation to calculate parameter α is implemented with a look-up table. To reduce the memory required for storing the table, the diff value (difference between maximum and minimum values) and the parameter α are expressed by an exponential notation. For example, diff is approximated with a 4-bit significant part and an exponent. Consequently, the table for 1 / diff is reduced into 16 elements for 16 values of the significand as follows: DivTable [] = {0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0 } (5)
[0024] This would have a benefit of both reducing the complexity of the calculation as well as the memory size required for storing the needed tables.
[0025] Besides the above template and left template can be used to calculate the linear model coefficients together, they also can be used alternatively in the other 2 LM modes, called LM_A, and LM_L modes.
[0026] In LM_A mode, only the above template is used to calculate the linear model coefficients. To get more samples, the above template is extended to (W+H) samples. In LM_L mode, only left template are used to calculate the linear model coefficients. To get more samples, the left template is extended to (H+W) samples.
[0027] In LM_LA mode, left and above templates are used to calculate the linear model coefficients.
[0028] To match the chroma sample locations for 4: 2: 0 video sequences, two types of down-sampling filter are applied to luma samples to achieve 2 to 1 down-sampling ratio in both horizontal and vertical directions. The selection of down-sampling filter is specified by a SPS level flag. The two down-sampling filters are as follows, which are corresponding to “type-0” and “type-2” content, respectively. RecL′ (i, j) = [recL (2i-1, 2j-1) +2·recL (2i-1, 2j-1) +recL (2i+1, 2j-1) + recL (2i-1, 2j) +2·recL (2i, 2j) +recL (2i+1, 2j) +4] >>3 (6) recL′ (i, j) =recL (2i, 2j-1) +recL (2i-1, 2j) +4·recL (2i, 2j) +recL (2i+1, 2j) + recL (2i, 2j+1) +4] >>3 (7)
[0029] Note that only one luma line (general line buffer in intra prediction) is used to make the down-sampled luma samples when the upper reference line is at the CTU boundary.
[0030] This parameter computation is performed as part of the decoding process, and is not just as an encoder search operation. As a result, no syntax is used to convey the α and β values to the decoder.
[0031] For chroma intra mode coding, a total of 8 intra modes are allowed for chroma intra mode coding. Those modes include five traditional intra modes and three cross-component linear model modes (LM_LA, LM_A, and LM_L) .
[0032] Multiple Model CCLM (MMLM)
[0033] In the JEM (J. Chen, E. Alshina, G. J. Sullivan, J. -R. Ohm, and J. Boyce, Algorithm Description of Joint Exploration Test Model 7, document JVET-G1001, ITU-T / ISO / IEC Joint Video Exploration Team (JVET) , Jul. 2017) , multiple model CCLM mode (MMLM) is proposed for using two models for predicting the chroma samples from the luma samples for the whole CU. In MMLM, neighbouring luma samples and neighbouring chroma samples of the current block are classified into two groups, each group is used as a training set to derive a linear model (i.e., a particular α and β are derived for a particular group) . Furthermore, the samples of the current luma block are also classified based on the same rule for the classification of neighbouring luma samples. Three MMLM model modes (MMLM_LA, MMLM_T, and MMLM_L) are allowed for choosing the neighbouring samples from left-side and above-side, above-side only, and left-side only, respectively.
[0034] Fig. 6 shows an example of classifying the neighbouring samples into two groups. Threshold is calculated as the average value of the neighbouring reconstructed luma samples. A neighbouring sample with Rec′L [x, y] <= Threshold is classified into group 1; while a neighbouring sample with Rec′L [x, y] > Threshold is classified into group 2.
[0035] OBMC (Overlapped Block Motion Compensation)
[0036] When OBMC is applied, top and left boundary pixels of a CU are refined using motion information of neighbouring blocks with a weighted prediction as described in JVET-L0101.
[0037] Conditions of not applying OBMC are as follows: ● When OBMC is disabled at SPS level ● When current block has intra mode or IBC mode ● When current luma block area is smaller or equal to 32
[0038] Additionally, OBMC is adaptively controlled on a block level as follows: ● OBMC flag is inherited from a neighbouring affine block for affine merge mode. ● OBMC is not applied to a block if there is a neighbour block coded with IBC, palette, or BDPCM modes. ● When applying OBMC to a block, block boundary check regarding whether OBMC is applied to the boundary is further made based on the reference samples of the current block. If any absolute difference between the prediction sample and non-interpolated (integer pel) reference sample is greater than a threshold, the OBMC is not applied to that boundary.
[0039] A subblock-boundary OBMC is performed by applying the same blending to the top, left, bottom, and right subblock boundary pixels using motion information of neighbouring subblocks. It is enabled for the subblock-based coding tools: ● Affine AMVP modes; ● Affine merge modes and subblock-based temporal motion vector prediction (SbTMVP) ; ● Subblock-based bilateral matching.
[0040] When OBMC mode is used in CIIP (Combined Inter-Intra Prediction) mode with LMCS, inter blending is performed prior to LMCS mapping of inter samples. LMCS (Luma Mapping with Chroma Scaling) is applied to blended inter samples which are combined with LMCS applied intra samples in CIIP mode, where InterpredY represents the samples predicted by the motion of current block in the original domain, IntrapredY represents the samples predicted in the mapped domain, OBMCpredY represents the samples predicted by the motion of neighbouring blocks in the original domain, and w0 and w1 are the weights.
[0041] When OBMC mode is used in a LIC coded block, the LIC parameters are applied to generate the corresponding prediction samples for the OBMC of the LIC coded block. Besides, to reduce the complexity, the OBMC is only applied to the top and left CU boundaries while being always disabled for the boundaries of the internal sub-blocks of the LIC coded block.
[0042] Multi-Hypothesis Prediction (MHP)
[0043] In the multi-hypothesis inter prediction mode (JVET-M0425) , one or more additional motion-compensated prediction signals are signalled, in addition to the conventional bi-prediction signal. The resulting overall prediction signal is obtained by sample-wise weighted superposition. With the bi / uni prediction signal pbi / uni and the first additional inter prediction signal / hypothesis h3, the resulting prediction signal p3 is obtained as follows: p3= (1-α) pbi / uni+αh3
[0044] The weighting factor α is specified by the new syntax element add_hyp_weight_idx, according to the following mapping table: Table 1. Mapping between weighting factor α is specified by the new syntax element add_hyp_weight_idx.
[0045] Analogously to above, more than one additional prediction signal can be used. The resulting overall prediction signal is accumulated iteratively with each additional prediction signal. pn+1= (1-αn+1) pn+αn+1hn+1
[0046] The resulting overall prediction signal is obtained as the last pn (i.e., the pn having the largest index n) . Within this EE, up to two additional prediction signals can be used (i.e., n is limited to 2) .
[0047] The motion parameters of each additional prediction hypothesis can be signalled either explicitly by specifying the reference index, the motion vector predictor index, and the motion vector difference, or implicitly by specifying a merge index. A separate multi-hypothesis merge flag distinguishes between these two signalling modes.
[0048] For inter AMVP mode, MHP is only applied if non-equal weight in BCW is selected in bi-prediction mode.
[0049] Combination of MHP and BDOF (Bi-directional Optical Flow) is possible, however the BDOF is only applied to the bi-prediction signal part of the prediction signal (i.e., the ordinary first two hypotheses) .
[0050] InterCCCM
[0051] InterCCCM applies the CCCM method for predicting chroma samples from reconstructed luma samples when the CU uses inter prediction or intra block copy (IBC) . Fig. 7 illustrates the decoder side of the method. The cross-component filters are derived using the prediction blocks of luma and chroma. The derived filters are applied to the reconstructed luma block and blended with the prediction blocks of chroma to produce the final chroma prediction blocks. Filter coefficients are derived in step 720 for each block separately using the prediction signals (i.e., predY 710, predCb 712 and predCr 714) and the filters are applied to the reconstructed luma signal in step 730 as shown in Fig. 7. The reconstructed luma signal is formed by combining the luma prediction (PredY) 710 and residual luma signal (resY) using an adder 722. After applying the filters, the step 730 generates filtered-predicted Cb 740 and filtered-predicted Cr 750. The reconstructed Cb signal is formed by combining the filtered-predicted Cb 740 and residual Cb signal (i.e., resCb) using an adder 742. Similarly, the reconstructed Cr signal is formed by combining the filtered-predicted Cr 750 and residual Cr signal (i.e., resCr) using an adder 752. In the blending process the filtered reconstructed luma blocks use blending weight of 0.75 and chroma prediction blocks use blending weight of 0.25.
[0052] The 8-tap filter consist of 6 spatial luma samples, a nonlinear term, and a bias term. The spatial luma samples (L0, …, L5) are obtained from the luma grid, where the 6 luma samples closest to the chroma position C are selected without down sampling as shown in Fig. 8. The predicted chroma value is obtained as, predChromaVal = c0 L0+ c1L1 + c2L2 + c3L3 + c4L4 + c5L5 + c6 nonlinear ( (L0+L3+1) >> 1) + c7 B, where nonlinear is CCCM’s nonlinear operator and B is bias. The filter coefficients are derived using ECM’s division-free Gaussian elimination method and the necessary offsets are applied to samples prior to filter derivation. The offsets for division-free Gaussian elimination method are obtained using a four-point average of the luma and chroma prediction blocks, where the four points correspond to the top-left, top-right, bottom-left and bottom-right corners of the blocks. For filter coefficient derivation, at most 256 chroma samples are used.
[0053] Usage of the mode is signalled with a CABAC coded TU level flag. One new CABAC context was included to support this. The InterCCCM flag is only signalled if the TU’s luma Cbf is non-zero and the CU’s predMode is either MODE_INTER or MODE_IBC.
[0054] The encoder performs an RD decision in the transform selection loop for the chroma components when luma Cbf is non-zero and the CU’s predMode is either MODE_INTER or MODE_IBC.
[0055] Bi-prediction with CU-level Weight (BCW)
[0056] In HEVC, the bi-prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. Pbi-pred= ( (8-w) *P0+w*P1+4)>>3.
[0057] Five weights are allowed in the weighted averaging bi-prediction, w∈ {-2, 3, 4, 5, 10} . For each bi-predicted CU, the weight w is determined in one of two ways: 1) for a non-merge CU, the weight index is signalled after the motion vector difference; 2) for a merge CU, the weight index is inferred from neighbouring blocks based on the merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256) . For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (i.e., w∈ {3, 4, 5} ) are used.
[0058] At the encoder, fast search algorithms are applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows, where further details can be found in the VTM software and document JVET-L0646: –When combined with AMVR, unequal weights are only conditionally checked for 1-pel and 4-pel motion vector precisions if the current picture is a low-delay picture. –When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode. –When the two reference pictures in bi-prediction are the same, unequal weights are only conditionally checked. –Unequal weights are not searched when certain conditions are met, depending on the POC distance between current picture and its reference pictures, the coding QP, and the temporal level.
[0059] The BCW weight index is coded using one context coded bin followed by bypass coded bins. The first context coded bin indicates whether equal weight is used; and if unequal weight is used, additional bins are signalled using bypass coding to indicate which unequal weight is used.
[0060] Weighted prediction (WP) is a coding tool supported by the H. 264 / AVC and HEVC standards to efficiently code video content with fading. Support for WP is also added into the VVC standard. WP allows weighting parameters (weight and offset) to be signalled for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weight (s) and offset (s) of the corresponding reference picture (s) are applied. WP and BCW are designed for different types of video content. In order to avoid interactions between WP and BCW, which will complicate VVC decoder design, if a CU uses WP, then the BCW weight index is not signalled, and w is inferred to be 4 (i.e. equal weight is applied) . For a merge CU, the weight index is inferred from neighbouring blocks based on the merge candidate index. This can be applied to both normal merge mode and inherited affine merge mode. For constructed affine merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index for a CU using the constructed affine merge mode is simply set to the BCW index of the first control point MV.
[0061] In VVC, CIIP and BCW cannot be jointly applied for a CU. When a CU is coded with CIIP mode, the BCW index of the current CU is set to 2 (e.g. equal weight) .
[0062] Convolutional Cross-Component Model (CCCM)
[0063] In CCCM, a convolutional model is applied to improve the chroma prediction performance. The convolutional model has 7-tap filter consist of a 5-tap plus sign shape spatial component, a nonlinear term and a bias term. The input to the spatial 5-tap component of the filter consists of a centre (C) luma sample which is collocated with the chroma sample to be predicted and its above / north (N) , below / south (S) , left / west (W) and right / east (E) neighbours as illustrated in Fig. 9.
[0064] The nonlinear term (denoted as P) is represented as power of two of the centre luma sample C and scaled to the sample value range of the content: P = (C*C + midVal) >> bitDepth.
[0065] That is, for 10-bit content it is calculated as: P = (C*C + 512) >> 10
[0066] The bias term (denoted as B) represents a scalar offset between the input and output (similarly to the offset term in CCLM) and is set to middle chroma value (512 for 10-bit content) .
[0067] Output of the filter is calculated as a convolution between the filter coefficients ci and the input values and clipped to the range of valid chroma samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B
[0068] The filter coefficients ci are calculated by minimising MSE between predicted and reconstructed chroma samples in the reference area. Fig. 10 illustrates the reference area which consists of 6 lines of chroma samples above and left of the PU. Reference area extends one PU width to the right and one PU height below the PU boundaries. Area is adjusted to include only available samples. The extensions to the area shown in blue are needed to support the “side samples” of the plus shaped spatial filter and are padded when in unavailable areas.
[0069] The MSE minimization is performed by calculating autocorrelation matrix for the luma input and a cross-correlation vector between the luma input and chroma output. Autocorrelation matrix is LDL decomposed and the final filter coefficients are calculated using back-substitution. The process follows roughly the calculation of the ALF filter coefficients in ECM, however LDL decomposition was chosen instead of Cholesky decomposition to avoid using square root operations.
[0070] Regression-based GPM Blending
[0071] Regression-based GMP blending mode is designed as an additional GPM implicit mode, where the two integer blending matrices (W0 and W1) are derived from the template (1 line above, 1 column left) . The blending matrices are modelled as an affine linear function of the sample positions (x, y) in the current CU: W0 (x, y) = a·x + b·y + c and W1 (x, y) = 1 -W0 (x, y) .
[0072] The parameters (a, b, c) are derived from the reference template using the same solver (i.e., MSE minimization) as the one used for CCCM. A list of pairs of candidates is built from the regular GPM candidates and re-ordered with the template cost.
[0073] The GPM implicit mode is signalled by a CU-level flag (gpm_implicit_flag) . If gpm_implicit_flag is true, a merge-idx is coded to signal the pair of GPM candidates to be used. If gpm_implicit_flag is false, the regular GPM syntax elements are signalled.
[0074] In the present invention, methods and apparatus to improve the coding performance for video coding systems using interCCP mode are disclosed.
[0075] Fusion of Chroma Intra Prediction Modes
[0076] In ECM, two chroma intra prediction signals can be fused together. One of the two chroma intra prediction signals is predicted using one of the DM mode, DIMD chroma mode and the four default modes (non-LM mode) . The other chroma intra prediction signal is predicted using cross-component linear prediction modes (LM mode) . Two different methods are supported.
[0077] In the first method, the LM mode can be either MM-CCLM or MM-CCCM, and the final predictor is derived as follows: predC (i, j) = (w0×pred0 (i, j) +w1×pred1 (i, j) + (1<< (shift-1) ) ) >>shift where pred0 (i, j) is the predictor obtained by applying the non-LM mode, pred1 (i, j) is the predictor obtained by applying the LM mode and predC (i, j) is the final predictor of the current chroma block. The two weights, w0 and w1 are determined by the intra prediction mode of adjacent chroma blocks and shift is set equal to 2. Specifically, when the above and left adjacent blocks are both coded with LM modes, {w0, w1} = {1, 3} ; when the above and left adjacent blocks are both coded with non-LM modes, {w0, w1} = {3, 1} ; otherwise, {w0, w1} = {2, 2} . Two template costs are calculated by fusing the angular chroma prediction with MM-CCLM or MM-CCCM, respectively, and the one of the two CCPs which provides a smaller template cost is utilized to derive pred1.
[0078] In the second method, the LM mode can be either MMLM or CCLM mode, and the final predictor is derived as follows: predC (i, j) = α0×pred0 (i, j) + α1×rec′L (i, j) +α2×β where pred0 (i, j) is the predictor obtained by applying the non-LM mode, rec′L (i, j) is the set of downsampled reconstructed luma samples at co-located positions and predC (i, j) is the final predictor of the current chroma block. β is a fixed value and is set equal to 512 for 10-bit content. The three weights, α0, α1 and α2 are derived from the adjacent luma and chroma samples using the same LDL derivation method as in CCCM.
[0079] For the syntax design, one index is signaled to indicate whether fusion is applied and which method is used. It is noted that for I slices, the non-LM mode can be DM mode, DIMD chroma mode and the four default modes. For non-I slices, only DIMD chroma mode is allowed to be fused with LM modes. BRIEF SUMMARY OF THE INVENTION
[0080] A method and apparatus for video coding are disclosed. According to this method, input data associated with a current block comprising a current first-colour block and a current second-colour block is received, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side. A motion-compensated predictor is generated for the current second-colour block. A cross-component predictor for the current second-colour block is generated according to a cross-component prediction candidate selected from a candidate list. One or more adaptive blending weights are determined using one or more regression-based models comprising location information. An inter cross-component predictor is derived by blending the motion-compensated predictor and the cross-component predictor by using said one or more adaptive blending weights. The current second-colour block is encoded or decoded by using coding information comprising the inter cross-component predictor.
[0081] In one embodiment, multiple regression-based models are used. In one embodiment, a first weight associated with the inter cross-component predictor is derived from a first regression-based model using first horizontal location information and / or first vertical location information. In one embodiment, a second weight associated with the chroma motion-compensated predictor is derived from a second regression-based model using second horizontal location information and / or second vertical location information.
[0082] In one embodiment, multiple lines in a template region are used to derive said one or more adaptive blending weights.
[0083] In one embodiment, a flag is signalled or parsed to indicate whether to use said one or more adaptive blending weights or original fixed blending weights. In one embodiment, said one or more adaptive blending weights are further dependent on CU size, QP value, luma coding mode, weights of interCCCM filter, sample value, sample differences between original chroma prediction samples and filtered chroma samples, blending weight differences between 4 corners of the current block or a combination thereof.
[0084] In one embodiment, said one or more adaptive blending weights are clipped to a range.BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing.
[0086] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
[0087] Fig. 2 illustrates examples of a multi-type tree structure corresponding to vertical binary splitting (SPLIT_BT_VER) , horizontal binary splitting (SPLIT_BT_HOR) , vertical ternary splitting (SPLIT_TT_VER) , and horizontal ternary splitting (SPLIT_TT_HOR) .
[0088] Fig. 3 shows an example of a CTU divided into multiple CUs with a quadtree and nested multi-type tree coding block structure, where the bold block edges represent quadtree partitioning and the remaining edges represent multi-type tree partitioning.
[0089] Fig. 4 shows the 67 intra prediction modes as adopted by the VVC video coding standard.
[0090] Fig. 5 shows an example of the location of the left and above samples and the sample of the current block involved in the LM_LA mode.
[0091] Fig. 6 shows an example of classifying the neighbouring samples into two groups according to multiple mode CCLM.
[0092] Fig. 7 illustrates the InterCCCM method at the decoder side.
[0093] Fig. 8 illustrates the luma samples L0, …, L5 in relation to the chroma sample C.
[0094] Fig. 9 illustrates an example of spatial part of the convolutional filter.
[0095] Fig. 10 illustrates an example of reference area with paddings used to derive the filter coefficients.
[0096] Fig. 11 illustrates an example of predicting the current CU using bi-prediction.
[0097] Fig. 12 illustrates an example of predicting the current CU by consider neighbouring samples in bi-prediction.
[0098] Fig. 13 illustrates an example of non-local GPM with regression-based blending, where the initial position of W0 (x, y) and W1 (x, y) are set as the left-top position of reference CU and W0 (x, y) and W1 (x, y) are propagated from the reference CU to the current CU according to the relative x and y coordinates.
[0099] Fig. 14 illustrates an example of target sample and neighbouring sample usage during training sample collection.
[0100] Fig. 15 illustrates an example of spatial weighting WS and distance weighting WD in a template region.
[0101] Fig. 16 illustrates an example of reconstruction sample padding in model training.
[0102] Fig. 17A-C illustrates examples of centroid of the current block before and after applying spatial weighting or distance weighting for centroid at (0.5, 0.5) (Fig. 17A) , (1.25, 1.25) (Fig. 17B) , and (8.25, 8.25) (Fig. 17C) .
[0103] Fig. 18 illustrates an example of OBMC blending lines according to regression-derived model.
[0104] Fig. 19 illustrates an example of OBMC blending weightings according to regression-derived model.
[0105] Fig. 20 illustrates an example of 4 corner positions in regression-based model for OBMC weighting generation.
[0106] Fig. 21 illustrates a flowchart of an exemplary video coding system that incorporates adaptive weightings for interCCP using a regression-based method with location information according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION
[0107] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.
[0108] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.
[0109] The following methods are proposed to improve the luma / chroma inter coding prediction accuracy or coding performance. In one embodiment, whether to allow or apply the following proposed methods can depend on SPS / PPS / SH / PH syntax, or CTU or CU / PU / TU level syntax or semantic. In another embodiment, whether to allow or apply the following proposed methods can also depend on implicit conditions. For example, the following proposed methods can be applied based on block width, block height, or block area. In another embodiment, the following proposed improvements of regression-based blending can be conditionally applied to partial positions / sub-block of the current block. In another embodiment, the term “block” in this invention can refer to TU / TB, CU / CB, PU / PB, pre-defined region, or CTU / CTB. Any combination of the following proposed methods in this invention can be applied.
[0110] Bi-prediction with Regression-based Blending
[0111] For Bi-prediction case, the prediction in the current CU is generated from the two predictions P0 and P1 as shown in Fig. 11. In the prior art, taking averaging or using BCW weight to blend P0 and P1 is used. In the current proposed method, a regression-based blending method is used.
[0112] In one embodiment, the regression model is shown in the following equation: P=a*P0+b*P1+c.
[0113] In this equation, P means the prediction of current sample, P0 and P1 means the two prediction samples to be blended. The coefficients a, b, c are the values to be derived.
[0114] In one embodiment, the position-based regression model is considered. As shown in the following equation: P= (a*x+b*y+c) *P0+ (d*x+e*y+f) *P1+g*x+h*y+i
[0115] In this equation, x and y correspond to the positions in the horizontal and vertical axes. The coefficients a, …, i are the values to be derived.
[0116] In another embodiment, any constraint on the coefficients can be applied to simplify the model. For example, the constraint that sum of weights in P0 and P1 shall be equal to one can be added to the above model. For the first exemplary model, the equation becomes to P=a*P0+ (1-a) *P1+c. For the second exemplary model, the equation becomes to P= (a*x+b*y+c) *P0+ (1- (a*x+b*y+c) ) *P1+g*x+h*y+i.
[0117] For another example, some of the coefficients can be set to zero to further simplify the model. Follow the above example, coefficient c can set to zero, and the equation of first exemplary model becomes to P=a*P0+ (1-a) *P1. Only one coefficient shall be derived during the blending process.
[0118] The proposed method can combine with other embodiment mentioned in other section. For example, in one embodiment, the regression-based bi-prediction method can be jointly or separately trained for the luma or chroma component mentioned above in the section entitled OBMC with regression-based blending. For another example, in one embodiment, the neighbouring samples can be considered in the regression-based method. The following example shows the case that we consider additional four neighbouring samples in the regression model. When calculating the prediction on position A, the four neighbouring samples in the position B, C, D, E of P0 and P1 are also considered as shown in Fig. 12. Following equations shows one example of the regression-based blending model.
[0119] MHP with Regression-based Blending
[0120] In prior art of MHP, the additional hypothesis is blended with the original prediction with a fixed weight. In the proposed method, the regression-based blending is used. In one embodiment, the blending equation used in MHP is shown in the following equation: pn+1=a*pn+b*hn+1+c.
[0121] The coefficients a, b, and c are derived using the regression-based method, and the data set used for the derivation can be any of the methods mentioned above.
[0122] In another embodiment, if over than one hypothesis is used, the blending weight can be derived together instead of generating the prediction step by step: pn+1=a*pbi / uni+b*h3+c*h4+…+ m*hn+1+n.
[0123] For example, if two hypotheses are used (i.e., p4 is used for final prediction) , the step-by-step method is shown as follows and the coefficients a, b, and c are derived twice: p3=a*pbi / uni+b*h3+c, p4=a*p3+b*h4+c.
[0124] In the proposed method, the final prediction p4 is derived using following one-step equation: p4=a*p3+b*h3+c*h4+d.
[0125] The coefficients a, b, c, and d can be derived using the regression-based method, and the data set used for the derivation can be any of the methods mentioned in this disclosure.
[0126] In another embodiment, the regression-based blending methods can be any other methods mentioned in this disclosure. For example, the position-dependent model can be applied to derive the prediction.
[0127] In another invention, a new type of merge candidate can be derived by the regression-based MHP blending method and be inserted into the merge candidate list. To derive the new merge candidate, N merge candidates are blended by MHP blending methods and the blending weights are calculated by minimizing the difference between current template and blended reference template (e.g., Tblended=a*Tref0+b*Tref1+c*Tref2+d) or any regression-based blending method mentioned in this disclosure. After deriving the new merge candidate, the derived regression parameters are saved in a buffer. The candidate and the corresponding derived regression parameters are inserted into merge candidate list and reordered by ARMC (Adaptive Reordering Merge Candidates) with other merge candidates.
[0128] Non-Local GPM with Regression-based Blending
[0129] In existing GPM methods, inherited GPM mode is not used. That is, the GPM merge candidates, partition modes and blending weights cannot be inherited by other CUs. However, the GPM mode is mostly be selected by the CUs at object boundaries and the partition mode and blending weights can be similar along the object boundaries. In this invention, we propose to derive GPM partition mode and blending weights based on previous coded CUs and the derived partition mode and blending weights can be referenced by other CUs.
[0130] In one embodiment, the blending map (i.e., W0 (x, y) and W1 (x, y) ) is modelled as a linear function of the sample positions (x, y) in the current CU by minimizing the difference between the reconstruction and prediction samples of a reference CU as follows: P = P0 × W0 (x, y) + P1 × W1 (x, y) , where W0 (x, y) = ax + by + c , W1 (x, y) = 1 -W0 (x, y)
[0131] In one embodiment, the initial position (i.e., x=0 and y=0) of W0 (x, y) and W1 (x, y) is set as the left-top position of reference CU and W0 (x, y) and W1 (x, y) are propagated from the reference CU to current CU according to the relative x and y coordinates as shown in Fig. 13.
[0132] In another embodiment, the W0 (x, y) and W1 (x, y) derived from the reference CU are directly used by the current CU. That is, when deriving the W0 (x, y) and W1 (x, y) , the initial position of W0 (x, y) and W1 (x, y) is set as the left-top position of reference CU. However, when applying the W0 (x, y) and W1 (x, y) blending map, the initial position is set as the left-top position of the current CU.
[0133] In one embodiment, for merge modes, up to k0 (k0>=0) blending map candidates obtained from spatial adjacent and non-adjacent neighbours are inserted into candidate list and then reordered by ARMC.
[0134] In another embodiment, for GPM mode, up to k1 (k1>=0) blending map candidates can be obtained from spatial adjacent and non-adjacent neighbours and compared with the GPM blending map derived from templates by RDOs. A flag can be signalled after geoflag to indicate whether the regression-based GPM blending is applied. If the flag is true, a regression-based GPM candidate index is further signalled to indicate the regression-based GPM candidate in the candidate list.
[0135] In one embodiment, only the CUs with specific sizes can be used to derive blending map for reference or can apply blending map derived from regression-based blending methods.
[0136] The reference CU can be any CU with multiple (i.e., > 1) predictors and the reference CU is located in a specific search range. The search range is defined to constrain the relative position between reference CU and current CU smaller than a threshold and also constrain the value of W0 (x, y) and W1 (x, y) smaller or larger than thresholds. The search range can be the current CTU, current CTU row, current slice, or any specific range.
[0137] The blending map W0 (x, y) and W1 (x, y) can also be derived by minimizing reconstruction and prediction sample differences at template region same as method mentioned in the section entitled “Regression-based GPM Blending” and be referenced by other CUs as described in previous embodiments.
[0138] The blending map can be derived from any equation (e.g., P = P0 × W0 (x, y) + P1 × W1 (x, y) + C × W2 + N × W3 + S × W4 + …+ B × WN) as mentioned above.
[0139] BCW with Regression-based Blending
[0140] In one embodiment, the regression-based blending map (e.g., W0 (x, y) and W1 (x, y) mentioned in the section entitled “Regression-based GPM Blending” or other blending methods mentioned above) can be derived for any bi-predictive candidate in AMVP or merge modes to replace or refine original BCW blending. Specifically, the regression-based blending map is derived by minimizing reconstruction and prediction sample differences at template region or in previous coded CU as mentioned in the section entitled “Non-local GPM with regression-based blending” and applied to L0 and L1 predictors to generate a bi-predictive predictor for a bi-predictive candidate.
[0141] In one embodiment, for inter or affine merge modes including MMVD, CIIP, BM, TM, etc., the regression-based blending maps are derived for all bi-predictive candidates in the merge candidate list and the corresponding TM costs (indicated as for ith candidate) are also calculated. The can be used to compare with the TM cost derived from original the BCW blending (indicated as for ith candidate) . When is smaller than the ith candidate is blended by the regression-based blending map instead of the original BCW blending.
[0142] In another embodiment, for inter or affine merge modes including MMVD, CIIP, BM, TM, etc., a new type of bi-predictive candidate is proposed by deriving the regression-based blending map for each bi-predictive candidate in the candidate list. The bi-predictive candidate with corresponding derived regression-based blending can be viewed as a new bi-predictive candidate and be added into the candidate list.
[0143] In one embodiment, for inter or affine AMVP modes, the regression-based blending maps are derived for all bi-predictive candidates in the AMVP candidate list or derived for the bi-predictive AMVP candidate after AMVP search. If the RD cost of regression-based blending is smaller than the RD cost of original BCW blending, a flag is signalled before BCW index to indicate the usage of regression-based blending. When the flag is true, the signalling of BCW index is skipped.
[0144] Partial Regression-based Blending
[0145] In the section entitled “Regression-based GPM blending” , a regression-based blending method is proposed to blend two GPM predictors according to a position-dependent blending map (i.e., W0 (x, y) and W1 (x, y) ) when the gpm_implicit_flag is true. Although the gpm_implicit_flag is signalled after gpm flag, the proposed blending behaviour is much different from original GPM blending behaviour where only a region of CU is blended by two predictors and the remaining regions are predicted by two GPM candidates in the original GPM. In this invention, we propose to blend two predictors from bi-predictive merge candidates, GPM candidates, CIIP candidates, or other multi-predictor candidates only in some regions of the CU by regression-based blending instead of blending in the entire CU.
[0146] In one invention, the blending regions are adaptive according to the CU size, the blending weight differences between 4 corners of CU (i.e., the differences between W0 (0, 0) , W0 (w, 0) , W0 (0, h) and W0 (w, h) ) , the inherited BCW index, the sample difference between two predictors, the sample value, or other information or the blending region is predefined.
[0147] In one invention, the blending region is iteratively determined by regression. That is, initial blending maps W0 (x, y) and W1 (x, y) of two predictors are derived by the methods mentioned in the section entitled “Regression-based GPM blending” or other methods mentioned in this disclosure. Based on the initial blending map, some positions (xi, yi) in the CU can have W0 (xi, yi) and W1 (xi, yi) larger than and / or smaller than thresholds. The samples with W0 (xi, yi) and W1 (xi, yi) larger than and / or smaller than thresholds are then removed from the training samples of regression-based blending method and the second iteration of regression-based blending derivation process is performed by using the remaining training samples to obtain the second iteration regression-based blending map. The derivation process can be performed iteratively until the blending weights of all the training samples are not larger than and / or smaller than thresholds, until a maximum iteration number is reached. The maximum iteration number can be adaptive to CU size, the blending weight differences between 4 corners of CU (i.e., the differences between W0 (0, 0) , W0 (w, 0) , W0 (0, h) and W0 (w, h) ) , the inherited BCW index, or other information or the maximum iteration number can be predefined.
[0148] OBMC with Regression-based Blending
[0149] Weight for Block Boundary Blending
[0150] In ECM OBMC, fixed weightings are used in CU-boundary OBMC and subblock-boundary OBMC. The newly adopted template-matching OBMC uses the reconstruction luma samples, the current predictor and the neighbouring predictor to determine the weighting strength and weighting lines in CU-boundary. The template-matching OBMC implies that adaptive weighting brings benefits on coding efficiency but fails to provide more general weightings in OBMC. In the proposed method, a novel regression-based blending method is disclosed.
[0151] Regression-based Training Samples Collection and Model Derivation
[0152] In prior-art regression-based GPM blending, one line of reconstruction samples is used to perform model training and to derive one single model to generate position-based weighting. In the proposed method, some novel methods regarding training samples collection and model derivation are disclosed.
[0153] Joint or Separate Training for Luma and Chroma Component
[0154] In the proposed method, luma and chroma components may use a joint single or joint multiple models when the current block performs OBMC process.
[0155] In one embodiment, single model or multiple models can be derived from luma component, and chroma components reuse the luma derived models with or without refinement.
[0156] Example (1) : Extrapolation filter-based intra prediction (EIP) mode derives single model or multiple models for luma CBs. For chroma components, they can reuse the EIP derived single model or multiple models from luma CBs. To better fit chroma components, EIP luma derived models can be refined, such as by combining different model parameters together or removing model parameters or adjusting model parameters.
[0157] Example (2) : Chroma components may inherit partial luma derived models, and some model parameters may be modified according to Cb and Cr information.
[0158] In another embodiment, the luma component derives its own single model or multiple models, and chroma components derive their own single model or multiple models.
[0159] In another embodiment, chroma components Cb and Cr derive a joint single mode or joint multiple models.
[0160] Example (1) : Down-sampled or non-down-sampled luma reconstructed samples are used to derive model and neighbouring chroma reconstructed samples are used as the training target in the regression process. When calculating training cost, both neighbouring Cb and Cr chroma reconstructed samples are used and the objective function is to minimize the cost by considering Cb and Cr jointly. The jointly derived single model or multiple models are used in the prediction stage of the current block.
[0161] Example (2) : Down-sampled or non-down-sampled luma reconstructed samples are used to derive model and neighbouring chroma reconstructed samples are used as the training target in the regression process. When calculating the training cost, both neighbouring Cb and Cr chroma reconstructed samples are firstly converted into another domain using some transforms. Then, the objective function is to minimize the cost considering transformed Cb and Cr jointly. The jointly derived single model or multiple models are used in the prediction stage of the current block. For the current block, Cb and Cr components may be converted into another domain by using some transforms or omit this process.
[0162] Reconstruction Lines and Samples Collection and Samples Usage in Model Training
[0163] In existing regression-based GPM blending and template-matching-based OBMC, one reconstruction sample line is used for the model training and template-matching, respectively. In the proposed method, multiple reconstruction sample lines and different reconstruction sample collection methods during training are disclosed.
[0164] In one embodiment, multiple reconstruction sample lines are used in the regression-based method.
[0165] In another embodiment, multiple reconstruction sample lines are used to train one singe model or multiple models.
[0166] In another embodiment, subsampled reconstruction samples are collected in the model training, and subsampled reconstruction samples are collected according to some pre-defined or classification rules. For example: (1) Reconstruction line: …╳○○○╳○○○╳…
[0167] In example (1) , ○ and ╳ are uniformly distributed and reconstruction samples. ○ is used in the model training and ╳ is used in the cost calculation. (2) Reconstruction line: …╳○○○○○○○○○╳…
[0168] In example (2) , ○ and ╳ are non-uniformly distributed. The distribution of ○ and ╳ can be determined according to some pre-defined positions or classification results. The reconstruction sample ○ is used in model training and ╳ is used in cost calculation. (3) Reconstruction line: …╳○○△△○○○╳…
[0169] In example (3) , ○, △ and ╳ are non-uniformly distributed. The distribution of △, ○ and ╳can be determined according to some pre-defined positions or classification results. The reconstruction sample ○ is used in one model training, the reconstruction sample △ is used in another model training and ╳ is used in the cost calculation.
[0170] In another embodiment, when training one target sample, its neighbouring samples can be considered or omitted in the model training, as shown in Fig. 14. Sample at location A is the target sample, and samples B, C, D and E can be considered in the model training.
[0171] In another embodiment, to derive one predictor sample, one sample or multiple samples can be used.
[0172] In another embodiment, in model training, spatial weighting WS or distance weighting WD is considered, as shown in Fig. 15.
[0173] In another embodiment, in the model training, when there is only the left-template available or top-template available, either spatial weighting WS or distance weighting WD is considered.
[0174] In another embodiment, in the model training, when there is only the left-template available or top-template available, both spatial weighting WS and distance weighting WD are considered.
[0175] In another embodiment, in the model training, when there is only the left-template available or top-template available, neither spatial weighting WS nor distance weighting WD is considered.
[0176] In another embodiment, the right-top reconstruction sample and the left-bottom reconstruction sample in reconstruction sample line is padded during model training, as shown in Fig. 16. The reconstruction samples A, B, C and D are padded during the model training.
[0177] Model Derivation Methods
[0178] In the existing regression-based method, the model derivation mainly considers one model derivation using available reconstruction samples. Several new model derivation methods are proposed.
[0179] In the proposed methods, the model derivation method can further consider the centroid of one block, spatial weighting WS or distance weighting WD.
[0180] In one embodiment, the centroid of the current block is at (0.5, 0.5) as shown in Fig. 17A, and model derivation considers spatial weighting and distance weighting to make the centroid move toward the right-bottom location (1.25, 1.25) , as shown in Fig. 17B.
[0181] In another embodiment, the centroid of current block is at (0.5, 0.5) and model derivation considers spatial weighting and distance weighting. Furthermore, padding samples also consider spatial weighting and distance weighting to make centroid move toward right-bottom location (8.25, 8.25) as shown in Fig. 17C.
[0182] In another embodiment, either spatial weighting or distance weighting is applied when the centroid of current block is within a location.
[0183] Furthermore, depending on the neighbouring template, multiple models can be derived based on the left-only template, top-only template or left-top template or sub-regions inside the template region.
[0184] In one embodiment, multiple models can be derived using a different template region, such as, left-only, top-only or left-top template. Or, multiple models are derived using a sub-region inside the template.
[0185] Example (1) : 3 models are derived using the left-only, top-only and left-top template, respectively.
[0186] Example (2) : 2 models are derived using sub-regions inside the template region, where the sub-regions are determined according to some pre-defined rules or classification results. For example, the mean sample value or the median sample value can be used to separate two sub-regions.
[0187] Yet in another proposed method, derived models can be inherited from previous derived models, which are derived from earlier coded blocks or pictures. The inherited derived models can be saved as a list, like the merging candidate list, and selected by the following encoded / decoded blocks.
[0188] In one embodiment, a model list is constructed, and derived models are stored as candidates in the model list. Each encoded / decoded block can select one model candidate in the model list.
[0189] In the proposed method, after the current block is reconstructed, the current reconstructed block can be used to derive the model for the prediction of following blocks.
[0190] In one embodiment, after the current block is reconstructed, model derivation is performed using the current reconstruction and model is stored as a candidate for the prediction of following blocks.
[0191] Regression-based Adaptive Weighting Lines and Weightings
[0192] In the proposed method, based on the regression-derived model, adaptive multiple blending lines and blending weightings are generated and applied to OBMC process. Four corner positions can be used to test the weighting strength change inside the current block. To constrain the weighting strength, weighting clipping can be considered.
[0193] In one embodiment, OBMC blending lines are adaptively changed according to the regression-derived model, as shown in Fig. 18.
[0194] In another embodiment, OBMC blending weightings are adaptively applied using the regression-derived model, as shown in Fig. 19.
[0195] In another embodiment, four corner positions (0, 0) , (width, 0) , (0, height) and (width, height) are used to check the generated OBMC weightings.
[0196] Example (1) : If the weighting difference between any two corner positions is smaller than or equal to a threshold, the regression-based model is not applied to current block’s OBMC process.
[0197] Example (2) : If weighting of any two corner positions is smaller than a threshold, the regression-based model is not applied to the OBMC process of the current block.
[0198] In another embodiment, the generated OBMC weighting is clipped to be in a range. For example, the clipping range is within [1 / 32, 31 / 32] for weightings of neighbouring predictor or weightings of current predictor.
[0199] InterCCCM and Inter CCP Merge Modes with Regression-Based Blending
[0200] In the existing interCCCM method, the blending weights of original chroma prediction samples and the filtered chroma samples are fixed and set as [0.25, 0.75] which lacks the flexibility. In this invention, the blending weights of original chroma prediction samples and the filtered chroma samples are proposed to be derived from regression-based blending methods or adaptive to some auxiliary information.
[0201] In one invention, the blending weights of original chroma prediction samples and the filtered chroma samples are derived from template samples. That is, the interCCCM filter derived from luma and chroma predictors in the CU is first applied to the neighbouring luma reconstruction samples to generate neighbouring filtered chroma samples. The neighbouring filtered chroma samples are then blended with original neighbouring chroma prediction samples by Tblended=a*Tfiltered+b*Tpred+c or other blending equations to generate neighbouring blended chroma samples Tblended. The blending maps are derived by the regression-based blending methods mentioned in other sections of this disclosure by minimizing the sample differences between chroma reconstruction samples and the blended chroma samples in neighbouring regions. The derived blending maps can be applied to the samples in the CU.
[0202] In another embodiment, the blending weight of original chroma prediction samples and the filtered chroma samples are adaptive to the CU size, QP value, luma coding mode, the weights of interCCCM filter, sample value, sample differences between original chroma prediction samples and the filtered chroma samples, the blending weight differences between 4 corners of CU (i.e., the differences between W0 (0, 0) , W0 (w, 0) , W0 (0, h) and W0 (w, h) ) or other information.
[0203] In inter CCP merge mode, the final predictor of the current chroma block is formed by combining the motion-compensation predicted signals and the cross-component predicted signals derived using the selected inherited CCP model. The blending weights of inter CCP merge mode are fixed as (wCCP, winter) = (3 / 4, 1 / 4) which may lack the flexibility. To improve the performance and increase the flexibility of inter CCP merge mode, some techniques from the regression-based GPM blending can be used here.
[0204] In one embodiment, a regression-based model with location information can be used to determine the blending weight of inter CCP merge mode. For example, the blending weight wCCP of the CCP predictor can be derived from the following regression model using horizontal location information and / or vertical location information, and the blending weight wMC of the motion-compensated predictor can be derived from another regression model using horizontal location information and / or vertical location information. Two regression models can be solved together: wCCP=a0x+b0y+c0 wMC=a1x+b1y+c1
[0205] In one embodiment, if the current block is coded by inter CCP merge mode, a flag can be signalled to indicate whether the proposed regression-based blending method with location information or the original fixed weights is used for determining the blending weight of inter CCP merge mode.
[0206] In one embodiment, the weight clipping method as regression-based GPM can be used in the proposed regression-based blending method for inter CCP merge mode. After obtaining the blending weights from the regression model, a clipping operation can be further applied to the output of the regression model to clip the values within a pre-defined range. For example, if the blending weight has 5 bits for the fractional part, the clipping range can be 0 to 32, or 1 to 31.
[0207] In one embodiment, a single-line template region as the regression-based GPM mode can be used to train the regression model for the blending weight of inter CCP merge mode.
[0208] In another embodiment, multiple-line template region can be used to train the regression model for the blending weight of inter CCP merge mode.
[0209] Regression-based chroma blending for interCCCM
[0210] In one invention, the blending weights of original chroma prediction samples and the interCCCM filtered chroma samples are derived from template samples. The interCCCM filter derived from luma and chroma predictors in the CU is first applied to the neighboring luma reconstruction samples to generate neighboring filtered chroma samples. The neighboring filtered chroma predictor (denoted as Tfiltered) is then blended with original neighboring chroma predictor (denoted as Tpred) to generate blended chroma predictor (denoted as) by Tblended=w0*Tfiltered+w1*Tpred+w2 or other blending equations. The blending maps are derived by the regression-based blending methods by minimizing the sample differences between chroma reconstruction samples and the blended chroma samples in neighboring regions.
[0211] In one embodiment, the regression-based blending method can be position-based (as mentioned in the section entitled “Regression-based GPM Blending” ) or CU-based methods depending on conditions. In one example, when the differences between weights at 4 CU corners derived by position-based regression methods are smaller than a threshold, the CU-based methods are then applied or the regression-based chroma blending for interCCCM is skipped. In another example, when the weight of interCCCM filtered predictor derived by CU-based methods is smaller or larger than a threshold, the regression-based chroma blending for interCCCM is skipped.
[0212] In one embodiment, the neighboring chroma prediction samples (i.e., Tpred) are pre-stored in buffers after the reconstructions of previous coded CUs and can be fetched when deriving the regression-based blending map. In another embodiment, the neighboring chroma prediction samples (i.e., Tpred) are predicted by the motion vectors of current CU. In one example, the motion vectors in current CU boundary are used to predict the neighboring chroma prediction samples. In another example, the center motion vector of current CU is used to predict the neighboring chroma prediction samples.
[0213] In one embodiment, if a neighboring CU is coded by skip mode, the samples in such CU are excluded from the training process of regression-based chroma blending. That is the luma reconstruction and chroma prediction samples are not used as training samples for regression. In another embodiment, one of the neighboring CUs is coded by skip mode, the regression-based chroma blending is skipped for current CU.
[0214] In one embodiment, a CCCM filter with shape and tap size aligning with inter CCCM filter is derived from the CU-level luma reconstruction and chroma reconstruction or luma reconstruction and chroma prediction or luma prediction and chroma prediction and stored for each previous coded CU. That is, for each neighboring CU, a corresponding CCCM filter is derived and is applied to the neighboring luma reconstruction samples to generate filtered chroma prediction samples (i.e., Tfiltered) . The neighboring filtered chroma prediction samples are used in regression-based chroma blending process.
[0215] In another embodiment, several CCCM filters with different shapes and sizes are derived from the CU-level luma reconstruction and chroma reconstruction or luma reconstruction and chroma prediction or luma prediction and chroma prediction and stored for each previous coded CU. After applying each CCCM filter on the luma reconstruction of previous coded CU, the SSD between chroma reconstruction and the filtered chroma predictor are calculated. The CCCM filters are then reordered by the SSD. The CCCM filter used to generate the neighboring filtered chroma prediction samples (i.e., Tfiltered) for regression-based chroma blending can be the first CCCM filter with smallest SSD or can be selected by an index which is signaled in the bitstream at picture / slice / CTU / CU / PU level. When the CCCM filters used in neighboring CUs are selected, the shape and size of inter CCCM filter in current CU is determined according to the selected CCCM filters. In another example, the CCCM filter used to generate the neighboring filtered chroma prediction samples (i.e., Tfiltered) for regression-based chroma blending are determined by the SSD between the blended chroma predictor (i.e., Tblended) and the chroma reconstruction in the neighboring template regions. The CCCM filters are reordered by the SSD. The CCCM filter can be the first CCCM filter with smallest SSD or can be selected by an index which is signaled in the bitstream at picture / slice / CTU / CU / PU level.
[0216] Regression-based luma and chroma blending for interCCCM
[0217] In one invention, the luma prediction samples and chroma prediction samples of current CU are taken as the input training samples for the inter CCCM filter derivation by minimizing the difference between filtered inter CCCM prediction (i.e., Pfiltered) and chroma reconstruction of current CU and the luma reconstruction samples and chroma prediction samples of current CU are taken as the inputs when applying the inter CCCM filter.
[0218] In one embodiment, the filtered inter CCCM prediction (i.e., Pfiltered) is obtained as, Pfiltered = c0 L0+ c1L1 + c2L2 + c3L3 + c4L4 + c5L5 + c6 nonlinear ( (L0+L3+1) >> 1) + c7Pchroma + c8 B, where {L0, L1, .., L5} and B are introduced in the section entitled “InterCCCM” and Pchroma indicates the chroma predictor of current CU.
[0219] In one embodiment, the inter CCCM filter can be any shape of CCCM filter consisted of luma and / or chroma spatial, gradient, position, non-linear, bias terms, or any other term. The data of chroma terms can be collected from chroma prediction of current CU. And the data of luma terms can be collected from luma reconstruction or prediction of current CU.
[0220] InterCCCM for zero luma residual TU
[0221] In one invention, the interCCCM can be enabled for zero luma residual TU. For a zero luma residual TU, a flag is signaled to indicate usage of inter CCCM. The chroma blending weight between filtered chroma prediction and original chroma prediction can be a fixed weight aligned with or different with non-zero luma residual TUs or a weight determined by an index signaled in bitstream at picture / slice / CTU / CU / PU level or a weight derived by regression-based methods (e.g., the regression-based chroma blending introduced in the section entitled “Regression-based chroma blending for interCCCM” ) .
[0222] In one embodiment, the interCCCM is enabled implicitly for zero luma residual TU depending on conditions.
[0223] In one example, the interCCCM is enabled implicitly for zero luma residual TU depending on the QP, POC, CU size, coding mode, or any other coding information.
[0224] In another example, when the SSD between chroma prediction and the filtered chroma prediction is smaller and / or larger than a threshold, the interCCCM for the zero luma residual CU is disabled.
[0225] InterCCCM for small and large luma residual region
[0226] In one invention, the inter CCCM filter can only be applied on small or large luma residual region in a TU. That is, a inter CCCM filter is derived from the samples in small or large luma residual region of the TU and can only be applied on corresponding region.
[0227] In another invention, the two interCCCM filters are derived respectively in small and large luma residual regions and can be applied on the corresponding regions.
[0228] In one embodiment, the small luma residual region consists of zero residual regions and / or regions with all or average of the luma residuals smaller than a threshold.
[0229] In one embodiment, the large luma residual region consists of non-zero residual region and / or regions with all or average of the luma residuals larger than a threshold.
[0230] In one embodiment, the inter CCCM filter can only be derived and applied on small and / or large luma residual region at the boundary of a TU.
[0231] In one embodiment, the small and large luma residual region can be determined in sample-based or subblock-based.
[0232] In one embodiment, the chroma blending weight between filtered inter CCCM chroma prediction and original chroma prediction can be a fixed weight aligned with or different with original inter CCCM or a weight determined by an index signaled in bitstream at picture / slice / CTU / CU / PU level or a weight derived by regression-based methods (e.g., the regression-based chroma blending introduced in the section entitled “Regression-based chroma blending for interCCCM” ) .
[0233] In one embodiment, the interCCCM is enabled implicitly for small and / or high luma residual regions depending on conditions.
[0234] In one example, the interCCCM is enabled implicitly for small and / or high luma residual regions depending on the QP, POC, CU size, coding mode, or any other coding information.
[0235] In another example, when the SSD between chroma prediction and the filtered chroma prediction is smaller and / or larger than a threshold, the interCCCM for such region is disabled.
[0236] Any of the foregoing proposed methods of adaptive blending weights using a regression-based method with location information can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter / intra / prediction module of an encoder, and / or an inter / intra / prediction module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module
[0237] With reference to the encoder and decoder in Fig. 1A and Fig. 1B, any of the proposed methods of adaptive blending weights using a regression-based method with location information can be implemented in an inter / intra / prediction / transform module (e.g. Intra Pred. 110 in Fig. 1A) of an encoder, and / or an inter / intra / prediction / transform module (e.g. Intra Pred. 150 in Fig. 1B) of a decoder.
[0238] Fig. 21 illustrates a flowchart of an exemplary video coding system that incorporates adaptive weightings for interCCP using a regression-based method with location information according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side and / or decoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to this method, input data associated with a current block comprising a current first-colour block and a current second-colour block is received in step 2110, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side. A chroma motion-compensated predictor is generated for the current second-colour block in step 2120. A cross-component predictor for the current second-colour block is generated according to a cross-component prediction candidate selected from a candidate list in step 2130. One or more adaptive blending weights are determined using one or more regression-based models with location information in step 2140. An inter cross-component predictor is derived by blending the chroma motion-compensated predictor and the cross-component predictor by using said one or more adaptive blending weights in step 2150. The current second-colour block is encoded or decoded by using coding information comprising the inter cross-component predictor in step 2160.
[0239] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0240] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.
[0241] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.
[0242] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1.A method of coding colour pictures or video, the method comprising:receiving input data associated with a current block comprising a current first-colour block and a current second-colour block, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side;generating a motion-compensated predictor for the current second-colour block;generating a cross-component predictor for the current second-colour block according to a cross-component prediction candidate selected from a candidate list;determining one or more adaptive blending weights using one or more regression-based models comprising location information;deriving an inter cross-component predictor by blending the motion-compensated predictor and the cross-component predictor by using said one or more adaptive blending weights; andencoding or decoding the current second-colour block by using coding information comprising the inter cross-component predictor.2.The method of Claim 1, wherein multiple regression-based models are used.3.The method of Claim 2, wherein a first weight associated with the inter cross-component predictor is derived from a first regression-based model using first horizontal location information and / or first vertical location information.4.The method of Claim 3, wherein a second weight associated with the motion-compensated predictor is derived from a second regression-based model using second horizontal location information and / or second vertical location information.5.The method of Claim 1, wherein multiple lines in a template region are used to derive said one or more adaptive blending weights.6.The method of Claim 1, wherein a flag is signalled or parsed to indicate whether to use said one or more adaptive blending weights or original fixed blending weights.7.The method of Claim 1, wherein said one or more adaptive blending weights are further dependent on CU size, QP value, luma coding mode, weights of interCCCM filter, sample value, sample differences between original chroma prediction samples and filtered chroma samples, blending weight differences between 4 corners of the current block or a combination thereof.8.The method of Claim 1, wherein said one or more adaptive blending weights are clipped to a range.9.An apparatus for coding colour pictures or video, the apparatus comprising one or more electronic circuits or processors arranged to:receive input data associated with a current block comprising a current first-colour block and a current second-colour block, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side;generate a motion-compensated predictor for the current second-colour block;generate a cross-component predictor for the current second-colour block according to a cross-component prediction candidate selected from a candidate list;determine one or more adaptive blending weights using one or more regression-based models comprising location information;derive an inter cross-component predictor by blending the motion-compensated predictor and the cross-component predictor by using said one or more adaptive blending weights; andencode or decode the current second-colour block by using coding information comprising the inter cross-component predictor.
Citation Information
Patent Citations
A Method, An Apparatus and a Computer Program Product for Video Encoding and Video Decoding
US20230262223A1
Method and apparatus for cross-component prediction for video coding
WO2023141245A1
Method and apparatus for implicit cross-component prediction in video coding system
WO2023198142A1