Methods and apparatus of geometric partitioning mode with cross-component prediction for video coding
By employing multiple cross-component models to partition and combine prediction hypotheses for chroma components, the coding performance and accuracy of video systems are improved, addressing inefficiencies in geometric partitioning and cross-component predictions.
Patent Information
- Application Number
- PCT/CN2025/107495
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-07-08
- Publication Date
- 2026-01-15
AI Technical Summary
Existing video coding systems face challenges in improving coding performance, particularly in handling geometric partitioning modes and cross-component predictions, which affect the efficiency and accuracy of chroma component encoding and decoding.
Implementing multiple cross-component models to generate multiple hypotheses of prediction for chroma components by partitioning the current block into regions, using geometric partitioning modes and combining different prediction hypotheses with weighted averaging for improved prediction accuracy.
Enhances the coding performance and prediction accuracy of chroma components by utilizing multiple cross-component models, leading to more efficient video encoding and decoding processes.
Smart Images

Figure CN2025107495_15012026_PF_FP_ABST
Abstract
Description
METHODS AND APPARATUS OF GEOMETRIC PARTITIONING MODE WITH CROSS-COMPONENT PREDICTION FOR VIDEO CODINGCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 668,375, filed on July 8, 2024. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION
[0002] The present invention relates to video coding system. In particular, the present invention relates to geometric partitioning mode with cross-component prediction in a video coding system. BACKGROUND AND RELATED ART
[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction 110, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.
[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.
[0006] The decoder, as shown in Fig. 1B, can use some of the functional blocks as the encoder. For example, the decoder can reuse Inverse Quantization 124 and Inverse Transform 126; however, Transform 118 and Quantization 120 are not needed at the decoder. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.
[0007] In VVC, the Sequence Parameter Set (SPS) and the Picture Parameter Set (PPS) contain high-level syntax elements that apply to entire coded video sequences and pictures, respectively. The Picture Header (PH) and Slice Header (SH) contain high-level syntax elements that apply to a current coded picture and a current coded slice, respectively.
[0008] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs. The individual CTUs in a slice are processed in raster-scan order. A bi-predictive (B) slice may be decoded using intra prediction or inter prediction with at most two motion vectors and reference indices to predict the sample values of each block. A predictive (P) slice is decoded using intra prediction or inter prediction with at most one motion vector and reference index to predict the sample values of each block. An intra (I) slice is decoded using intra prediction only. RELATED ART
[0009] Technical methods related to the present invention are described in this section. Section I. 1 focuses on the introduction to geometric partitioning mode (GPM) as mentioned in JVET-T2002 and JVET-AG2025. Section I. 2 focuses on cross-component prediction (CCP) merge mode for chroma inter coding as mentioned in JVET-AF0073. Section I. 3 focuses on cross-component residual model (CCRM) for inter prediction as mentioned in JVET-AE0059.
[0010] In the present invention, methods and apparatus to improve the coding performance of video coding systems using geometric partition modes are disclosed. BRIEF SUMMARY OF THE INVENTION
[0011] A method and apparatus for video coding using coding tools including one or more cross component models related modes are disclosed. According to the method, input data associated with a current block comprising a luma component and a chroma component is received, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side. The current block is partitioned into two or more regions. First hypothesis of prediction using a first cross-component model is derived. Second hypothesis of prediction is derived. Combined prediction is generated by using the first hypothesis of prediction and the second hypothesis of prediction. The chroma component of at least one of said two or more regions is encoded or decoded using the combined prediction.
[0012] In one embodiment, the current block is partitioned into a first partition and a second partition using GPM (Geometric Partitioning Mode) . In one embodiment, the current block consists of a first region, a second region and a third region, and wherein the first region is within the first partition, the second region is within the second partition and the third region is between the first region and the second region. In one embodiment, the second hypothesis of prediction using a second cross-component model, and wherein first predictors for the first region are generated based on the first hypothesis of prediction and second predictors for the second region are generated based on the second hypothesis of prediction. In one embodiment, third predictors for the third region are generated by combining the first hypothesis of prediction and the second hypothesis of prediction.
[0013] In one embodiment, the first cross-component model is derived using motion compensation prediction associated with at least one of said two or more regions.
[0014] In one embodiment, the second hypothesis of prediction is derived using motion compensation prediction.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Fig. 1A illustrates an exemplary adaptive Inter / Intra video coding system incorporating loop processing.
[0016] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
[0017] Fig. 2 illustrates an example of GPM splits grouped by identical angles.
[0018] Fig. 3 shows an exemplary system block diagram for Cross-component residual model (CCRM) .
[0019] Fig. 4 illustrates an example of the current block using a partitioning mode as a straight line from the top-right corner to the bottom-left corner.
[0020] Fig. 5 illustrates an example of the current block using a partitioning mode as a straight line from top-left corner to the bottom-right corner, the above template of the current block, and the left template of the current block.
[0021] Fig. 6 illustrates an example of the current block using a partitioning mode as a straight line from top-right corner to the bottom-left corner and storing the cross-component model for each grid of the current block.
[0022] Fig. 7 illustrates a flowchart of an exemplary video coding system that uses multiple-mode cross-component prediction according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION
[0023] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.
[0024] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.
[0025] I. 1 Geometric Partitioning Mode (GPM)
[0026] In VVC, a geometric partitioning mode is supported for inter prediction. The geometric partitioning mode is signalled using a CU-level flag as one kind of merge mode, with other merge modes including the regular merge mode, the MMVD mode, the CIIP mode and the subblock merge mode. A total 64 of partitions are supported by geometric partitioning mode for each possible CU size w×h=2m×2n with m, n ∈ {3…6} excluding 8x64 and 64x8.
[0027] When this mode is used, a CU is split into two parts by a geometrically located straight line as shown in Fig. 2. The location of the splitting line is mathematically derived from the angle and offset parameters of a specific partition. Each part of a geometric partition in the CU is inter-predicted using its own motion; only uni-prediction is allowed for each partition, that is, each part has one motion vector and one reference index. The uni-prediction motion constraint is applied to ensure that same as the conventional bi-prediction, only two motion compensated prediction are needed for each CU.
[0028] If geometric partitioning mode is used for the current CU, then a geometric partition index indicating the partition mode of the geometric partition (including angle and offset) , and two merge indices (one for each partition) are further signalled. The number of maximum GPM candidate size is signalled explicitly in SPS and specifies syntax binarization for GPM merge indices. After predicting each part of the geometric partitions, the sample values along the geometric partition edge are adjusted using a blending processing with adaptive weights. This is the prediction signal for the whole CU, and transform and quantization process will be applied to the whole CU as in other prediction modes. Finally, the motion field of a CU predicted using the geometric partition modes is stored.
[0029] More GPM extensions can be found in JVET-AG2025. For example, GPM with merge motion vector differences (MMVD) , GPM with template matching (TM) , GPM with inter and intra prediction, bi-predictive GPM, affine motion compensation (AMC) -GPM, implicit GPM, and / or spatial geometric partitioning mode (SGPM) .
[0030] I. 2 Cross-Component Prediction Merge Mode for Chroma Inter Coding
[0031] A CCP model is implicitly selected from a CCP merge list. The final prediction of the current chroma inter block is formed by combining the motion-compensation predicted signals and the cross-component predicted signals derived using the selected CCP model. The weights for combining predictions are fixed as (wCCP, winter) = (3 / 4, 1 / 4) .
[0032] I. 3 Cross-Component Residual Model (CCRM) for Inter Prediction
[0033] According to CCRM, a cross-component residual model (CCRM) is used to predict chroma samples from reconstructed luma samples when the block uses inter prediction or intra block copy (IBC) . Fig. 3 illustrates the decoder side of the method. The cross-component filters are derived using the prediction blocks of luma and chroma. The derived filters are applied to the reconstructed luma block and blended with the prediction blocks of chroma to produce the final chroma prediction blocks. In the blending process, the filtered reconstructed luma blocks use blending weight of 0.75 and chroma prediction blocks use blending weight of 0.25. As shown in Fig. 3, the cross-component filters are derived using the prediction signals of luma and chroma. The derived filters are applied to the reconstructed luma signal producing the final chroma predictions. Filter coefficients are derived in step 320 for each chroma component separately using the prediction signals (i.e., predY 310, and predCb 312 or predCr 314) and the filters are applied to the reconstructed luma signal in step 330 as shown in Fig. 3. The reconstructed luma signal is formed by combining the luma prediction (PredY) 310 and residual luma signal (resY) using an adder 322. After applying the filters, step 330 generates filtered-predicted Cb 340 and filtered-predicted Cr 350. The reconstructed Cb signal is formed by combining the filtered-predicted Cb 340 and residual Cb signal (i.e., resCb) using an adder 342. Similarly, the reconstructed Cr signal is formed by combining the filtered-predicted Cr 350 and residual Cr signal (i.e., resCr) using an adder 352.
[0034] In order to improve the prediction accuracy or coding performance of cross-component prediction, various schemes related to geometric partitioning mode with cross-component prediction are disclosed.
[0035] II. PROPOSED METHOD
[0036] The present invention proposes to use more than one cross-component models to generate multiple-hypotheses of prediction for chroma components of the current block when the current block partitions into multiple prediction blocks. For example, when Geometric Partitioning Mode (GPM) is applied to the current block, a partitioning mode (which can be denoted as an angle and / or distance / offset) is selected to split the current block into two prediction blocks. For each prediction block, one prediction mode is determined to generate one hypothesis of prediction. A first prediction mode (e.g. a first inter motion candidate) and / or a second prediction mode (e.g. a second inter motion candidate) can be selected from a candidate list (e.g. an inter merge candidate list containing multiple inter merging motion candidates) . As shown in Fig. 4, the current 8x8 block uses a partitioning mode as a straight line from the top-right corner to the bottom-left corner. A first prediction mode is determined to generate a first hypothesis of prediction and a second prediction mode is determined to generate a second hypothesis of prediction. The first and second hypotheses of prediction are combined according to weights associated with the partitioning mode. w (i, j) denotes the weight value for the second hypothesis of prediction at sample (i, j) . If the summation of the weight for the first hypothesis of prediction and the weight for the second hypothesis of prediction is 8, in this example, w (i, j) can be as follows.
[0037] For the samples far away from the partitioning line (e.g. the first region and the second region) , the weight for one of the first and second hypotheses of prediction is 0. In the first region, the weight for the second hypothesis of prediction is 0. That is, the weight for the first hypothesis of prediction is 8. In the second region, the weight for the first hypothesis of prediction is 0. That is, the weight for the second hypothesis of prediction is 8. For the samples near the partitioning line (e.g. the third region) , the weights for the first and second hypotheses of prediction are smaller than 8 and larger than 0. That is, for the third region, the first and second hypotheses of prediction are weighted averaged. Moreover, in the first part of the third region which is near the first region, the weight for the first hypothesis of prediction is larger than or equal to the weight for the second hypothesis of prediction. In the second part of the third region which is near the second region, the weight for the second hypothesis of prediction is larger than or equal to the weight for the first hypothesis of prediction. The first prediction block includes the first region and the first part of the third region and the second prediction block includes the second region and the second part of the third region.
[0038] Several methods are proposed to improve the prediction accuracy of chroma components for the partitioned block. Multiple cross-component models are used to generate multiple hypotheses of cross-component prediction (CCP) (for chroma components) for a partitioned block, respectively. For example, the partitioned block is a GPM coding block containing two prediction blocks. The first cross-component model is determined to generate a first hypothesis of CCP as the first hypothesis of prediction (for chroma components) . The second cross-component model is determined to generate a second hypothesis of CCP as the second hypothesis of prediction (for chroma components) .
[0039] In one embodiment, for chroma components, the predictors in the partitioned block are formed by using CCP. For a GPM coding block, in the first region, predictors of the chroma components are formed by using the first hypothesis of CCP. In the second region, predictors of the chroma components are formed by using the second hypothesis of CCP. In the third region, predictors of the chroma components are formed by combining (using the weights associated with the partitioning mode) the first and second hypotheses of CCP.
[0040] In another embodiment, for chroma components, the predictors in the partitioned block are formed by using CCP and existing prediction (e.g. motion compensation prediction) . In one embodiment, for a GPM coding block, in the first region, predictors of the chroma components are formed by using weighted average (e.g. {3: 1} ) of the first hypothesis of CCP and a first hypothesis of motion compensation prediction (which is generated using the first prediction mode, such as a first inter motion candidate) . In the second region, predictors of the chroma components are formed by using weighted average (e.g. {3: 1} ) of the second hypothesis of CCP and a second hypothesis of motion compensation prediction (which is generated using the second prediction mode, such as a second inter motion candidate) . In the third region, predictors of the chroma components are formed by combining (using the weights associated with the partitioning mode) (1) weighted average, for example, {3: 1} , of the first hypothesis of CCP and the first hypothesis of motion compensation prediction and (2) weighted average, for example, {3: 1} , the second hypothesis of CCP and the second hypothesis of motion compensation prediction. In another embodiment, motion compensation results in the first, second, and third regions are determined. Also, cross-component results in the first, second, and third regions are determined. Then, the cross-component results and motion compensation results are combined using weighted average (e.g. {3: 1} ) . When generating the motion compensation results, the first hypothesis of motion compensation prediction is used in the first region, the second hypothesis of motion compensation prediction is used in the second region, and the first and second hypotheses of motion compensation prediction are combined using the weights associated with the partitioning mode in the third region. When generating the cross-component results, the first hypothesis of CCP is used in the first region, the second hypothesis of CCP is used in the second region, and the first and second hypotheses of CCP are combined using the weights associated with the partitioning mode in the third region.
[0041] In one embodiment, when cross-component prediction merge mode (CCP merge mode) is applied to the current block and the current block is a partitioned block, multiple cross-component models are used.
[0042] In one case, multiple cross-component models are from multiple candidate lists, respectively. For example, a first CCP merge list is constructed and the first cross-component model for the first and third regions is selected from the first CCP merge list. A second CCP merge list is constructed and the second cross-component model for the second and third regions is selected from the second CCP merge list.
[0043] In one sub-embodiment, when selecting the cross-component model from a list, the selection can depend on an implicit rule or explicit rule.
[0044] For an example of the implicit rule, the selection can depend on template costs, partitioning mode, block width, block height, block position, or a combination thereof. A template cost can be calculated for each candidate in the list and the template cost for a candidate implies the distortion between the reconstructed samples in the template and the predicted samples (generated using the candidate) in the template. The template can include above template, left template, or both. In one method, the template is fixed to include both above template and left template. In another method, the template varies with the block position and / or partitioning mode. When the above or left template of the current block is not available according to the block position, only the available left or above template is used for calculating the template costs. When the first and second prediction blocks are determined according to the partitioning mode, the template associated with the first prediction block is used to calculate the template costs of candidates for the first cross-component model and the template associated with the second prediction block is used to calculate the template costs of candidates for the second cross-component model. As shown in Fig. 5, the first prediction block (including the first region and the first part of the third region) uses the left template, and the second prediction block (including the second region and the second part of the third region) uses the above template.
[0045] For an example of the explicit rule, the selection can depend on the explicit indication of at least one of GPM index (e.g. a merge index) for indicating the first prediction mode or second prediction mode (e.g. the first or second inter motion candidate) . For an example of the explicit rule, a cross-component model index for indicating the cross-component model from the list is signalled or parsed.
[0046] In one sub-embodiment, when constructing a candidate list, the list construction can depend on the partitioning mode, block width, block height, block area, and / or block position. As shown in Fig. 5, when the first and second prediction blocks are determined according to the partitioning mode, the first list construction (of candidates for the first cross-component model) comprises to include the cross-component models inherited from the left-oriented coded region and / or derived using the reconstructed or predicted samples from the left-oriented coded region. The second list construction (of candidates for the second cross-component model) comprises the cross-component models inherited from the above-oriented coded region and / or derived using the reconstructed or predicted samples from the above-oriented coded region. The above-oriented coded region has the position y smaller than the top-left position of the current block and the left-oriented coded region has the position x smaller than the top-left position of the current block.
[0047] In one sub-embodiment, the first cross-component model should be different from the second cross-component model. A redundancy check is performed. If the second cross-component model is redundant with the first cross-component model, the second cross-component model is modified or replaced with an alternative candidate cross-component model.
[0048] In another case, multiple cross-component models are from a single candidate list. For example, the single CCP merge list is constructed and the first and second cross-component models are selected from different candidates in the single CCP merge list.
[0049] In one sub-embodiment, the selection can depend on an implicit rule or explicit rule.
[0050] For an example of the implicit rule, the selection can depend on template costs, partitioning mode, block width, block height, block position, or a combination thereof. A template cost can be calculated for each candidate in the list and the template cost for a candidate implies the distortion between the reconstructed samples in the template and the predicted samples (generated using the candidate) in the template. For calculating above template costs, the template includes above template. For calculating left template costs, the template includes left template. The selection of the first cross-component model used for the first and third regions depends on either the above template costs or the left template costs, based on which template is associated with the first prediction block. The selection of the second cross-component model used for the second and third regions depends on either the above template costs or the left template costs, depending on which template is associated with the second prediction block. The first cross-component model for the first and third regions should be different from the second cross-component model for the second and third regions. A redundancy check is performed. If one cross-component model is redundant with the other, one of the two cross-component models is modified or replaced with an alternative candidate cross-component model. As shown in Fig. 5, the first cross-component model is selected using the left template costs, and the second cross-component model is selected using the above template costs. For example, the selected first cross-component model is used as the candidate with the smallest left template cost and the selected second cross-component model is used as the candidate with the smallest above template cost. If the selected second cross-component model is redundant with the selected first cross-component model, the second cross-component model is changed to use the cross-component model with the second smallest above template costs according to one embodiment. In another embodiment, the cross-component model with the larger cost between the smallest above and left template costs is changed. That is, when the smallest above template cost is larger than the smallest left template cost, the second cross-component model is changed to use the cross-component model with the second smallest above template cost. When the smallest left template cost is larger than the smallest above template cost, the first cross-component model is changed to use the cross-component model with the second smallest left template cost.
[0051] For an example of the explicit rule, the selection can depend on the explicit indication of at least one of GPM indexes (e.g. a merge index) for indicating the first prediction mode or second prediction mode (for example, the first or second inter motion candidate) . For an example of the explicit rule, one or more cross-component model indices for indicating the cross-component models from the list are signalled or parsed. If one cross-component model index is signalled or parsed for indicating the first cross-component model and another cross-component model index is signalled or parsed for indicating the second cross-component model, for indicating the second cross-component model from the list, the indication depends on the signalling or indication associated with the first cross-component model.
[0052] In one sub-embodiment, when constructing the single candidate list, the list construction can depend on the partitioning mode, block width, block height, block area, and / or block position.
[0053] In one embodiment, when cross-component residual model (CCRM) is applied to the current block and the current block is a partitioned block, multiple cross-component models are used. The cross-component filters are derived using the prediction blocks of luma and chroma.
[0054] In one case, the first cross-component model (filter) for the first and the third regions is derived using the prediction blocks (referring to motion compensation results) of luma and chroma in all or subset of the first prediction block (including the first region and the first part of the third region) or in all or subset of the first and third regions. The second cross-component model (filter) for the second and the third regions is derived using the prediction blocks (referring to motion compensation results) of luma and chroma in all or subset of the second prediction block (including the second region and the second part of the third region) or in all or subset of the second and third regions.
[0055] In another case, the first cross-component model (filter) for the first and the third regions is derived using the prediction blocks (referring to the first hypothesis of motion compensation prediction (which is generated using the first prediction mode, such as a first inter motion candidate) ) of luma and chroma in all or subset of the first prediction block (including the first region and the first part of the third region) or in all or subset of the first and third regions or in all or subset of the whole current block. The second cross-component model (i.e., filter) for the second and third regions is derived using the prediction blocks (referring to the second hypothesis of motion compensation prediction (which is generated using the second prediction mode, such as a second inter motion candidate) ) of luma and chroma in all or subset of the second prediction block (including the second region and the second part of the third region) or in all or subset of the second and third regions or in all or subset of the whole current block.
[0056] In one embodiment, the storage of cross-component models depends on the partitioning mode, block width, block height, and / or sample position. For each grid of the current block, a cross-component model is stored and / or can be referenced by the following coding blocks. The grid can be 1x1, 2x2, 4x4, or nxn, where n is a positive integer in luma or chroma samples. When determining which cross-component model is stored for the current grid, the centre position of the current grid is used. For example, the centre position of the current grid is at the right-bottom centre position of the current grid with (x, y) = ( (grid_w *grid_x + (grid_w>>1) ) , (grid_h *grid_y + (grid_h>>1) ) ) in the current block. If the centre position of the current grid is associated with the first cross-component model, the first cross-component model is stored. If the centre position of the current grid is associated with the second cross-component model, the second cross-component model is stored. If the centre position of the current grid is associated with both the first and second cross-component models, either the first cross-component model or the second cross-component model is stored according to a pre-defined rule. In one sub-embodiment, the pre-defined rule is fixed to store the second cross-component model. In another sub-embodiment, the pre-defined rule is fixed to store the first cross-component model. In another sub-embodiment, the pre-defined rule is to store the first or second cross-component model according to the partitioning mode and / or sample position and / or weighting. As shown in Fig. 6, the grid (grid_w x grid_h) is 4 x 4. Sample position in the current block with (x, y) = ( (4 *0 + 2) , (4 *0 + 2) ) is the right-bottom centre in grid (0, 0) . Sample position in the current block with (x, y) = ( (4 *1 + 2) , (4 *0 + 2) ) is the right-bottom centre in grid (1, 0) . Sample position in the current block with (x, y) = ( (4 *0 +2) , (4 *1 + 2) ) is the right-bottom centre in grid (0, 1) . Sample position in the current block with (x, y) = ( (4 *1 + 2) , (4 *1 + 2) ) is the right-bottom centre in grid (1, 1) . The stored cross-component models for grid (0, 0) , grid (1, 0) , grid (0, 1) , and grid (1, 1) can be the second, second, second, and first cross-component models, respectively, in this example.
[0057] The term “block” in this invention can refer to TU / TB, CU / CB, PU / PB, pre-defined region, or CTU / CTB.
[0058] Any combination of the proposed methods in this invention can be applied.
[0059] The proposed methods in this invention can be enabled and / or disabled according to implicit rules (for example, block width, height, or area) or according to explicit rules (for example, syntax on block, tile, slice, picture, SPS, or PPS level) . For example, the proposed method is applied when the block area is smaller / larger than a threshold.
[0060] Any of the foregoing proposed methods can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an intra / inter / IBC / prediction module of an encoder, and / or an intra / inter / IBC / prediction module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the intra / inter / IBC / prediction module of the encoder and / or the intra / inter / IBC / prediction module of the decoder, so as to provide the information needed by the intra / inter / IBC / prediction module.
[0061] The proposed methods as described above can be implemented in an encoder side or a decoder side. For example, any of the proposed methods can be implemented in an Intra / Inter coding module (e.g. Intra Pred. 150 / MC 152 in Fig. 1B) in a decoder or an Intra / Inter coding module in an encoder (e.g. Intra Pred. 110 / Inter Pred. 112 in Fig. 1A) . Any of the proposed methods can also be implemented as a circuit coupled to the intra / inter coding module at the decoder or the encoder. However, the decoder or encoder may also use additional processing unit to implement the required prediction processing. While the Intra Pred. units (e.g. unit 110 / 112 in Fig. 1A and unit 150 / 152 in Fig. 1B) are shown as individual processing units, they may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .
[0062] Fig. 7 illustrates a flowchart of an exemplary video coding system that uses multiple-mode cross-component prediction according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side or at the decoder side. The steps shown in the flowchart may also be implemented based on hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, input data associated with a current block comprising a luma component and a chroma component is received in step 710, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side. The current block is partitioned into two or more regions in step 720. First hypothesis of prediction using a first cross-component model is derived in step 730. Second hypothesis of prediction is derived in step 740. Combined prediction is generated by using the first hypothesis of prediction and the second hypothesis of prediction in step 750. The chroma component of at least one of said two or more regions is encoded or decoded using the combined prediction in step 760.
[0063] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0064] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.
[0065] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.
[0066] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1.A method of coding colour pictures, the method comprising:receiving input data associated with a current block comprising a luma component and a chroma component, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side;partitioning the current block into two or more regions;deriving first hypothesis of prediction using a first cross-component model;deriving second hypothesis of prediction;generating combined prediction by using the first hypothesis of prediction and the second hypothesis of prediction; andencoding or decoding the chroma component of at least one of said two or more regions using the combined prediction.2.The method of Claim 1, wherein the current block is partitioned into a first partition and a second partition using GPM (Geometric Partitioning Mode) .3.The method of Claim 2, wherein the current block consists of a first region, a second region and a third region, and wherein the first region is within the first partition, the second region is within the second partition and the third region is between the first region and the second region.4.The method of Claim 3, wherein the second hypothesis of prediction using a second cross-component model, and wherein first predictors for the first region are generated based on the first hypothesis of prediction and second predictors for the second region are generated based on the second hypothesis of prediction.5.The method of Claim 4, wherein third predictors for the third region are generated by combining the first hypothesis of prediction and the second hypothesis of prediction.6.The method of Claim 1, wherein the first cross-component model is derived using motion compensation prediction associated with at least one of said two or more regions.7.The method of Claim 1, wherein the second hypothesis of prediction is derived using motion compensation prediction.8.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block comprising a luma component and a chroma component, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side;partition the current block into two or more regions;derive first hypothesis of prediction using a first cross-component model;derive second hypothesis of prediction;generate combined prediction by using the first hypothesis of prediction and the second hypothesis of prediction; andencode or decode the chroma component of at least one of said two or more regions using the combined prediction.
Citation Information
Patent Citations
Method and apparatus of localized luma prediction mode inheritance for chroma prediction in video coding
CN108702501A
Inter-frame prediction method, encoder, decoder and storage medium
CN115315953A
Device and method for decoding video data
US20230412801A1
Method and apparatus for blending prediction in video coding system
WO2023207646A1