Geometric partitioning mode extensions

Geometric partitioning mode with regression-based blending addresses the challenge of complex motion in video coding by partitioning blocks and deriving blending weights, enhancing compression efficiency and prediction accuracy.

WO2025152999A1PCT designated stage expired Publication Date: 2025-07-24MEDIATEK INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/072649
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-17
Filing Date
2025-01-16
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in efficiently handling complex motion patterns in video frames, particularly in partitioning and predicting pixel blocks, leading to suboptimal compression and decoding performance.

Method used

Implementing geometric partitioning mode (GPM) with regression-based blending, where the current block is partitioned into multiple parts, and blending weights are derived using regression from neighboring template regions to generate predictors for these partitions, allowing for improved prediction and encoding.

Benefits of technology

Enhances video coding efficiency by better accommodating complex motion patterns, reducing redundancy, and improving compression performance through refined partitioning and prediction techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025072649_24072025_PF_FP_ABST
    Figure CN2025072649_24072025_PF_FP_ABST
Patent Text Reader

Abstract

A method of using geometric partitioning mode (GPM) with regression-based blending is provided. The current block is partitioned into at least two GPM partitions. The video coder derives a blending weight by regression using samples of a template region neighboring the current block for the partitioning. The video coder selects a pairing of candidates from a list of pairings of candidates. A merge index may be signaled to select a pairing of candidates from the list when a flag is signaled to indicate that the blending weight is derived by regression. The video coder applies the selected pairing of candidates to the two GPM partitions with the derived blending weight to generate a predictor of the current block. The video coder encodes or decodes the current block using the generated predictor.
Need to check novelty before this filing date? Find Prior Art

Description

GEOMETRIC PARTITIONING MODE EXTENSIONSCROSS REFERENCE TO RELATED PATENT APPLICATION (S)

[0001] The present disclosure is part of a non-provisional application that claims the priority benefit of U.S. Provisional Patent Application No. 63 / 621,649, filed on 17 January 2024. Content of above-listed application is herein incorporated by reference.TECHNICAL FIELD

[0002] The present disclosure relates generally to video coding. In particular, the present disclosure relates to methods of coding pixel blocks by geometric partitioning mode (GPM) .BACKGROUND

[0003] Unless otherwise indicated herein, approaches described in this section are not prior art to the claims listed below and are not admitted as prior art by inclusion in this section.

[0004] High-Efficiency Video Coding (HEVC) is an international video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) . HEVC is based on the hybrid block-based motion-compensated DCT-like transform coding architecture. The basic unit for compression, termed coding unit (CU) , is a 2Nx2N square block of pixels, and each CU can be recursively split into four smaller CUs until the predefined minimum size is reached. Each CU contains one or multiple prediction units (PUs) .

[0005] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Expert Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11. The input video signal is predicted from the reconstructed signal, which is derived from the coded picture regions. The prediction residual signal is processed by a block transform. The transform coefficients are quantized and entropy coded together with other side information in the bitstream. The reconstructed signal is generated from the prediction signal and the reconstructed residual signal after inverse transform on the de-quantized transform coefficients. The reconstructed signal is further processed by in-loop filtering for removing coding artifacts. The decoded pictures are stored in the frame buffer for predicting the future pictures in the input video signal.

[0006] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . The leaf nodes of a coding tree correspond to the coding units (CUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs. The individual CTUs in a slice are processed in raster-scan order. A bi-predictive (B) slice may be decoded using intra prediction or inter prediction with at most two motion vectors (MVs) and reference indices to predict the sample values of each block. A predictive (P) slice is decoded using intra prediction or inter prediction with at most one motion vector and reference index to predict the sample values of each block. An intra (I) slice is decoded using intra prediction only.

[0007] A CTU can be partitioned into one or multiple non-overlapped coding units (CUs) using the quadtree (QT) with nested multi-type-tree (MTT) structure to adapt to various local motion and texture characteristics. A CU can be further split into smaller CUs using one of the five split types: quad-tree partitioning, vertical binary tree partitioning, horizontal binary tree partitioning, vertical center-side triple-tree partitioning, horizontal center-side triple-tree partitioning.

[0008] Each CU contains one or more prediction units (PUs) . The prediction unit, together with the associated CU syntax, works as a basic unit for signaling the predictor information. The specified prediction process is employed to predict the values of the associated pixel samples inside the PU. Each CU may contain one or more transform units (TUs) for representing the prediction residual blocks. A transform unit (TU) is comprised of a transform block (TB) of luma samples and two corresponding transform blocks of chroma samples and each TB correspond to one residual block of samples from one color component. An integer transform is applied to a transform block. The level values of quantized coefficients together with other side information are entropy coded in the bitstream. The terms coding tree block (CTB) , coding block (CB) , prediction block (PB) , and transform block (TB) are defined to specify the 2-D sample array of one-color component associated with CTU, CU, PU, and TU, respectively. Thus, a CTU consists of one luma CTB, two chroma CTBs, and associated syntax elements. A similar relationship is valid for CU, PU, and TU.

[0009] For each inter-predicted CU, motion parameters consisting of motion vectors, reference picture indices and reference picture list usage index, and additional information are used for inter-predicted sample generation. The motion parameter can be signalled in an explicit or implicit manner. When a CU is coded with skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector delta or reference picture index. A merge mode is specified whereby the motion parameters for the current CU are obtained from neighbouring CUs, including spatial and temporal candidates, and additional schedules introduced in VVC. The merge mode can be applied to any inter-predicted CU. The alternative to merge mode is the explicit transmission of motion parameters, where motion vector, corresponding reference picture index for each reference picture list and reference picture list usage flag and other needed information are signalled explicitly per each CU.

[0010] Intra block copy (IBC) or current picture referencing (CPR) refer to coding pixel blocks by referencing pixel positions within same current picture as the current block by using block vectors.

[0011] In advanced motion vector prediction (AMVP) mode, a motion vector predictor (MVP) candidate is determined based on template matching (TM) error to select the one that reaches the minimum difference between the current block template and the reference block template, and then TM is performed only for this particular MVP candidate for MV refinement. The TM process may refine this MVP candidate using iterative search according to an adaptive motion vector resolution (AMVR) mode search pattern.SUMMARY

[0012] The following summary is illustrative only and is not intended to be limiting in any way. That is, the following summary is provided to introduce concepts, highlights, benefits and advantages of the novel and non-obvious techniques described herein. Select and not all implementations are further described below in the detailed description. Thus, the following summary is not intended to identify essential features of the claimed subject matter, nor is it intended for use in determining the scope of the claimed subject matter.

[0013] Some embodiments of the disclosure provide a method of using geometric partitioning mode (GPM) with regression-based blending is provided. The current block is partitioned into at least two GPM partitions. The video coder derives a blending weight by regression using samples of a template region neighboring the current block for the partitioning. The video coder selects a pairing of candidates from a list of pairings of candidates. A merge index may be signaled to select a pairing of candidates from the list when a flag is signaled to indicate that the blending weight is derived by regression. The video coder applies the selected pairing of candidates to the two GPM partitions with the derived blending weight to generate a predictor of the current block. The video coder encodes or decodes the current block using the generated predictor.

[0014] In some embodiments, the video coder selects the pairing of candidates by selecting from among a plurality of lists of pairings of candidates, with each pairing of candidates of a particular list including at least one candidate of a particular coding tool. In some embodiments, the video coder selects a pairing of candidates by selecting from only one list of pairings of candidates that includes pairings of any two candidates for any of two or more coding tools. For example, the list of pairings of candidates may include different pairings of candidates from inter merge candidates and affine merge candidates. For another example, the list of pairings of candidates may include different pairings of candidates from inter merge candidates, affine merge candidates, intra merge candidates, and intra block copy (IBC) merge candidates.

[0015] In some embodiments, the list of pairings of candidates is reordered according to template cost and each pairing of candidates in the list is assigned a corresponding merge index, such that the video coder may select a pairing of candidates by signaling a merge index to indicate the selected pairing of candidates in the list. In some embodiments, the merge index for the list of pairings of candidates is signaled when a flag (e.g., gpm_implicit_flag) is signaled to indicate that the blending weight is derived by regression.

[0016] In some embodiments, when only a first GPM partition of the at least two GPM partitions of the current block has a LIC enabling flag and a set of inherited LIC parameters, the video coder refines the set of inherited LIC parameters and applies the refined set of LIC parameters to a second GPM partition of the at least two GPM partitions. A neighboring template of the current block and a corresponding reference block may be used to refine the inherited LIC parameters for the second GPM partition.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are included to provide a further understanding of the present disclosure, and are incorporated in and constitute a part of the present disclosure. The drawings illustrate implementations of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. It is appreciable that the drawings are not necessarily in scale as some components may be shown to be out of proportion than the size in actual implementation in order to clearly illustrate the concept of the present disclosure.

[0018] FIG. 1 shows spatial and temporal candidates for merge mode.

[0019] FIG. 2 shows the control point motion vectors (CPMVs) of a current block that is coded by affine motion field.

[0020] FIG. 3 illustrates affine motion field per subblock.

[0021] FIG. 4 conceptually illustrates CPMV inheritance.

[0022] FIG. 5 illustrates the partitioning of a CU by the geometric partitioning mode (GPM) .

[0023] FIG. 6 illustrates an example uni-prediction candidate list for a GPM partition and the selection of a uni-prediction MV for GPM.

[0024] FIG. 7 illustrates an example partition edge blending process for GPM for a CU.

[0025] FIG. 8 illustrates the ramp function for weights for GPM blending based on displacement from a predicted sample position to the GPM partitioning boundary and the blending area.

[0026] FIG. 9 illustrates the GPM partitioning edge extended from the current CU to partition the template region.

[0027] FIG. 10 conceptually illustrates one list of pairings of candidates that includes all possible pairings of merge candidates.

[0028] FIG. 11 conceptually illustrates multiple lists of pairings of candidates that correspond to different coding tools.

[0029] FIG. 12 illustrates an example video encoder that may implement GPM.

[0030] FIG. 13 illustrates portions of the video encoder that implement GPM with blending by regression.

[0031] FIG. 14 conceptually illustrates a process for performing GPM prediction with regression-based blending weight using a list of pairings of candidates.

[0032] FIG. 15 illustrates an example video decoder that may implement GPM.

[0033] FIG. 16 illustrates portions of the video decoder that implement GPM with blending by regression.

[0034] FIG. 17 conceptually illustrates a process for performing GPM prediction with regression-based blending weight using a list of pairings of candidates.

[0035] FIG. 18 conceptually illustrates an electronic system with which some embodiments of the present disclosure are implemented.DETAILED DESCRIPTION

[0036] In the following detailed description, numerous specific details are set forth by way of examples in order to provide a thorough understanding of the relevant teachings. Any variations, derivatives and / or extensions based on teachings described herein are within the protective scope of the present disclosure. In some instances, well-known methods, procedures, components, and / or circuitry pertaining to one or more example implementations disclosed herein may be described at a relatively high level without detail, in order to avoid unnecessarily obscuring aspects of teachings of the present disclosure. I. Inter Merge Mode

[0037] Skip and Merge modes obtain the motion information from spatially neighboring blocks (spatial candidates) or a temporal co-located block (temporal candidate) . When a PU is Skip or Merge mode, no motion information is coded, instead, only the index of the selected candidate is coded. For Skip mode, the residual signal is forced to be zero and not coded. If a particular block is encoded as Skip or Merge, a candidate index is signaled to indicate which candidate among the candidate set is used for merging. Each merged PU reuses the MV, prediction direction, and reference picture index of the selected candidate.

[0038] FIG. 1 shows spatial and temporal candidates for merge mode. As illustrated, up to four spatial MV candidates are derived from A0, A1, B0 and B1 spatial neighbors, and one temporal MV candidate is derived from TBR or TCTR (TBR is used first, if TBR is not available, TCTR is used instead) . If any of the four spatial MV candidates is not available, the position B2 is then used to derive MV candidate as a replacement. After the derivation process of the four spatial MV candidates and one temporal MV candidate, removing redundancy (pruning) is applied to remove redundant MV candidates. If after removing redundancy (pruning) , the number of available MV candidates is smaller than five, three types of additional candidates are derived and are added to the candidate set (candidate list) . The encoder selects one final candidate within the candidate set for Skip, or Merge modes based on the rate-distortion optimization (RDO) decision, and transmits the index to the decoder. II. Affine Prediction

[0039] A. Affine Motion Field

[0040] An object in a video may have different types of motion, including translation motions, zoom in / out motions, rotation motions, perspective motions and the other irregular motions. In some embodiments, a block-based affine transform motion compensation prediction is used to account for these various types of motion. A block-based affine transform motion compensation prediction may be used. Specifically, the affine motion field mvx, mvy of the current block at position (x, y) is in the form of a linear model: mvx = a*x + b*y + c mvy = d*x + e*y + f               (0)

[0041] The coefficients {a, b, c, d, e, f } are parameters of the linear model. In some embodiments, the affine motion field at position (x, y) can be described by motion information of two control points (CPs) (at e.g., top-right and top-left corners of the block) (4-parameter model) or motion information of three control points (at e.g., top-right, top-left, and bottom-left corners of the block) (6-parameter model) .

[0042] For 4-parameter affine motion model, eq. 0 (motion vector at sample location (x, y) in a block) can be written as:

[0043] For 6-parameter affine motion model, eq. 0 can be written as:

[0044] Where (mv0x, mv0y) is motion vector of the top-left corner control point (top-left corner CPMV, or mv0) , (mv1x, mv1y) is the motion vector of the top-right corner control point (top-right corner CPMV, or mv1) , and (mv2x, mv2y) is the motion vector of the bottom-left corner control point (bottom-left corner CPMV, or mv2) .

[0045] In some embodiments, a 2-parameter model can be used to refine translation inter-prediction candidate (e.g., regular merge candidates) : MVxregress = MVxorig + biasX MVyregress = MVyorig + biasY                (3)

[0046] FIG. 2 shows the control point motion vectors (CPMVs) of a current block that is coded by affine motion field. The current block has CPMVs at top-left corner (mv0) , top-right-corner (mv1) , and bottom-left corner (mv2) . The affine motion field mv' at positions (x, y) in the current block #6200 can be derived using an affine motion model such as eq. 1 (4-parameter affine model) or eq. 2 (6-parameter affine model) .

[0047] In order to simplify the motion compensation prediction, block based affine transform prediction is applied. FIG. 3 illustrates affine motion field per subblock. The figure illustrates a motion field of motion vectors for a block having 16 4x4 subblocks. To derive the motion vector of each 4×4 luma subblock, the motion vector of the center sample of each subblock. is calculated according to eq. 1 or eq. 2, and rounded to 1 / 16 fraction accuracy. A motion compensation interpolation filters can be applied to generate the prediction of each subblock with derived motion vector. The subblock size of chroma-components is also set to be 4×4. The MV of a 4×4 chroma subblock is calculated as the average of the MVs of the top-left and bottom-right luma subblocks in the collocated 8x8 luma region.

[0048] B. Affine Merge Mode

[0049] In affine merge mode, the motion vectors at the control points (CPMVs) of the current CU are generated based on the motion information of the spatial neighboring CUs. There can be up to five CPMVP candidates, and an index is signalled to indicate the one to be used for the current CU. The following three types of CPMV candidates are used to form the affine merge candidate list: (1) inherited affine merge candidates that are extrapolated from the CPMVs of the neighbour CUs; (2) constructed affine merge candidates CPMVPs that are derived using the translational MVs of the neighbour CUs; (3) zero MVs.

[0050] An inherited affine candidate inherits an affine model from a neighboring block by directly obtaining the CPMVs from the neighboring blocks (one from left neighboring CUs and one from above neighboring CUs) . The candidate neighboring blocks for inheriting affine candidates are as shown in FIG. 1 above. For the left predictor, the scan order of the candidate blocks is A1→A0, and for the above predictor, the scan order for the candidate blocks is B1→B0→B2. When a neighboring affine CU is identified, its control point motion vectors are used to derive the CPMVP candidate in the affine merge list of the current CU.

[0051] FIG. 4 conceptually illustrates control point motion vector inheritance. As illustrated, for a current block 410, if a left-bottom neighboring block A is coded in affine mode, the motion vectors mv2, mv3, and mv4 of the top left corner, above right corner and left bottom corner of a CU 420 that contains the block A can be inherited by the current block 410. When block A is coded with 4-parameter affine model, the two CPMVs of the current CU 410 can be calculated according to mv2 and mv3. In case that block A is coded with 6-parameter affine model, the three CPMVs of the current CU 410 may be calculated according to mv2, mv3, and mv4.

[0052] Constructed affine candidate means the candidate is constructed by combining the neighbor translational motion information of each control point. The motion information for the control points is derived from the specified spatial neighbors and temporal neighbor shown in FIG. 1. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, the B2→B3→A2 blocks are checked and the MV of the first available block is used. For CPMV2, the B1→B0 blocks are checked and for CPMV3, the A1→A0 blocks are checked. For TMVP is used as CPMV4 if it’s available. After MVs of four control points are attained, affine merge candidates are constructed based on those motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3} , {CPMV1, CPMV2, CPMV4} , {CPMV1, CPMV3, CPMV4} , {CPMV2, CPMV3, CPMV4} , {CPMV1, CPMV2} , {CPMV1, CPMV3} . The combination of 3 CPMVs constructs a 6-parameter affine merge candidate and the combination of 2 CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling process, if the reference indices of control points are different, the related combination of control point MVs is discarded. After inherited affine merge candidates and constructed affine merge candidate are checked, if the list is still not full, zero MVs are inserted to the end of the list.

[0053] C. Affine AMVP Prediction

[0054] Affine AMVP mode can be applied to CUs with both width and height larger than or equal to 16. An affine flag in CU level is signalled in the bitstream to indicate whether affine AMVP mode is used and then another flag is signalled to indicate whether 4-parameter affine or 6-parameter affine. In this mode, the difference of the CPMVs of current CU and their predictors CPMVPs is signalled in the bitstream. The affine AVMP candidate list size is 2 and it is generated by using the following four types of CPMV candidate in order: (1) Inherited affine AMVP candidates that extrapolated from the CPMVs of the neighbour CUs, (2) Constructed affine AMVP candidates CPMVPs that are derived using the translational MVs of the neighbour CUs, (3) Translational MVs from neighboring CUs, and (4) Zero MVs.

[0055] The checking order of inherited affine AMVP candidates is same to the checking order of inherited affine merge candidates. The only difference is that, for AVMP candidate, only the affine CU that has the same reference picture as in current block is considered. No pruning process is applied when inserting an inherited affine motion predictor into the candidate list.

[0056] Constructed AMVP candidate is derived from the specified spatial neighbors shown in FIG. 1 above. The same checking order is used as in affine merge candidate construction. In addition, reference picture index of the neighboring block is also checked. The first block in the checking order is inter coded and uses the same reference picture as in current CUs. When the current CU is coded with 4-parameter affine mode, and mv0 and mv1 are both available, they are added as one candidate in the affine AMVP list. When the current CU is coded with 6-parameter affine mode, and all three CPMVs are available, they are added as one candidate in the affine AMVP list. Otherwise, constructed AMVP candidate is set as unavailable.

[0057] If the affine AMVP candidates list still has less than 2 candidates after valid inherited affine AMVP candidates and constructed AMVP candidate are inserted, mv0, mv1, and mv2 will be added, in order, as the translational MVs to predict all control point MVs of the current CU, when available. Finally, zero MVs are used to fill the affine AMVP candidates list if the list is still not full. III. Geometric Partitioning (GPM) and Extensions

[0058] A. GPM

[0059] The geometric partitioning mode (GPM) is signalled using a CU-level flag as one kind of merge mode, with other merge modes that includes the regular merge mode, the MMVD mode, the CIIP mode, and the subblock merge mode. In total 64 partition modes are supported by geometric partitioning mode for each possible CU size w×h=2m×2n with m, n ∈ {3…6} excluding 8x64 and 64x8.

[0060] FIG. 5 illustrates the partitioning of a CU by the geometric partitioning mode (GPM) . Each GPM partitioning or GPM split is a partition mode characterized by a distance-angle pairing that defines a bisecting or segmenting line. The figure illustrates examples of the GPM splits grouped by identical angles. As illustrated, when GPM is used, a CU is split into at least two parts by a geometrically located straight line. The location of the splitting line is mathematically derived from the angle and offset parameters of a specific partition.

[0061] Each partition in the CU formed by a partition mode of GPM is inter-predicted using its own motion (vector) . In some embodiments, only uni-prediction is allowed for each partition, that is, each part has one motion vector and one reference index. The uni-prediction motion constraint is applied to ensure that, similar to conventional bi-prediction, only two motion compensated prediction are performed for each CU.

[0062] If GPM is used for the current CU, then a geometric partition index indicating the partition mode of the geometric partitioning (angle and offset) and two merge indices (one for each partition) are further signalled. Each of the at least two partitions created by the geometric partitioning according to a partition mode may be assigned a merge index to select a candidate from a uni-prediction candidate list (also referred to as the GPM candidate list) . The pair of merge indices of the two partitions therefore select a pair of merge candidates. The maximum number of candidates in the GPM candidate list may be signalled explicitly in SPS to specify syntax binarization for GPM merge indices. After predicting each of the at least two partitions, the sample values along the geometric partitioning edge are adjusted using a blending processing with adaptive weights. This is the prediction signal for the whole CU, and transform and quantization process will be applied to the whole CU as in other prediction modes. The motion field of the CU as predicted by GPM is then stored.

[0063] The uni-prediction candidate list for a GPM partition (the GPM candidate list) may be derived directly from the merge candidate list of the current CU. FIG. 6 illustrates an example uni-prediction candidate list 600 for a GPM partition and the selection of a uni-prediction MV for GPM. The GPM candidate list 600 is constructed in an even-odd manner with only uni-prediction candidates that alternates between L0 MV and L1 MV. Let n be the index of the uni-prediction motion in the uni-prediction candidate list for GPM. The LX (i.e., L0 or L1) motion vector of the n-th extended merge candidate, with X equal to the parity of n, is used as the n-th uni-prediction motion vector for GPM. (These motion vectors are marked with “x” in the figure. ) In case a corresponding LX motion vector of the n-th extended merge candidate does not exist, the L (1 -X) motion vector of the same candidate is used instead as the uni-prediction motion vector for GPM.

[0064] As mentioned, the sample values along the geometric partition edge are adjusted using a blending processing with adaptive weights. Specifically, after predicting each part of a geometric partition using its own motion, blending is applied to the at least two prediction signals to derive samples around geometric partition edge. The blending weight for each position of the CU are derived based on the distance between the individual position and the partition edge. The distance for a position (x, y) to the partition edge are derived as:

[0065] where i, j are the indices for angle and offset of a geometric partition, which depend on the signaled geometric partition index. The sign of ρx, j and ρy, j depend on angle index i. The weights for each part of a geometric partition are derived as following: wIdxL (x, y) =partIdx ? 32+d (x, y) : 32-d (x, y)    (6) w1 (x, y) =1-w0 (x, y)      (8)

[0066] The variable partIdx depends on the angle index i. FIG. 7 illustrates an example partition edge blending process for GPM for a CU 700. In the figure, blending weights are generated based on an initial blending weight w0.

[0067] As mentioned, the motion field of a CU predicted using GPM is stored. Specifically, Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition and a combined Mv of Mv1 and Mv2 are stored in the motion field of the GPM coded CU. The stored motion vector type for each individual position in the motion filed are determined as: sType = abs (motionIdx) < 32 ? 2: (motionIdx ≤ 0 ? (1-partIdx) : partIdx)   (9)

[0068] where motionIdx is equal to d (4x+2, 4y+2) , which is recalculated from equation (2) . The partIdx depends on the angle index i. If sType is equal to 0 or 1, Mv0 or Mv1 are stored in the corresponding motion field, otherwise if sType is equal to 2, a combined Mv from Mv0 and Mv2 are stored. The combined Mv are generated using the following process: (i) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1) , then Mv1 and Mv2 are simply combined to form the bi-prediction motion vectors; (ii) otherwise, if Mv1 and Mv2 are from the same list, only uni-prediction motion Mv2 is stored.

[0069] B. GPM with Merge Vector Difference (GPM-MMVD)

[0070] In some embodiments, GPM is extended by applying motion vector refinement to existing GPM uni-directional MVs. A flag is first signaled for a GPM CU, to specify whether this mode is used. If the mode is used, each geometric partition of a GPM CU can further decide whether to signal motion vector difference (MVD) or not. If MVD is signaled for a geometric partition, after a GPM merge candidate is selected, the motion of the partition is further refined by the signaled MVDs information. All other procedures are kept the same as in GPM.

[0071] The MVD is signaled as a pair of distance and direction, similar as in MMVD. There are nine candidate distances (1 / 4-pel, 1 / 2-pel, 1-pel, 2-pel, 3-pel, 4-pel, 6-pel, 8-pel, 16-pel) , and eight candidate directions (four horizontal / vertical directions and four diagonal directions) involved in GPM with MMVD (GPM-MMVD) . In addition, when pic_fpel_mmvd_enabled_flag is equal to 1, the MVD is left shifted by 2 as in MMVD.

[0072] C. GPM with Template Matching (GPM-TM)

[0073] In some embodiments, template matching (TM) may be applied to refine MVs of GPM partitions. When GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether TM is applied to both geometric partitions. Motion information for each geometric partition is refined using TM. When TM is chosen, a template is constructed using left, above, or left and above neighboring samples according to partition angle. Table 1 below shows Template for the first and second geometric partitions, where A represents using above samples, L represents using left samples, and L+A represents using both left and above samples.

[0074] Table 1:

[0075] The motion is then refined by minimizing the difference between the current template and the template in the reference picture using the same search pattern of merge mode with half-pel interpolation filter disabled. A GPM candidate list is constructed as follows: (1) the video coder derives interleaved List-0 MV candidates and List-1 MV candidates directly from the regular merge candidate list, where List-0 MV candidates are higher priority than List-1 MV candidates. A pruning method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates; (2) the video coder further derives interleaved List-1 MV candidates and List-0 MV candidates directly from the regular merge candidate list, where List-1 MV candidates are higher priority than List-0 MV candidates. The same pruning method with the adaptive threshold is also applied to remove redundant MV candidates; and (3) the video coder pads the GPM candidate list with zero MV candidates until the GPM candidate list is full.

[0076] In some embodiments, the GPM-MMVD and GPM-TM are exclusively enabled to one CU for which GPM is used. In some embodiments, this is done by firstly signaling the GPM-MMVD syntax. When both of the two GPM-MMVD control flags are set to false (i.e., the GPM-MMVD are disabled for two GPM partitions) , the GPM-TM flag is signaled to indicate whether the template matching refinement is applied to the GPM partitions. Otherwise (at least one GPM-MMVD flag is set to true) , the value of the GPM-TM flag is inferred to be false.

[0077] D. GPM with Adaptive Blending

[0078] For GPM prediction, the final prediction samples are generated with by blending the prediction of the two prediction signals using weighted average. Two integer blending matrices (W0 and W1) are used. In some embodiments, the weights in the GPM blending matrices may be derived from a ramp function based on the displacement from a predicted sample position to the GPM partitioning boundary. The blending area size is fixed to two (2 samples on each side of the GPM partition split boundary) . FIG. 8 illustrates the ramp function for weights for GPM blending based on displacement (d) from a predicted sample position to the GPM partitioning boundary and the blending area (τ) . The figure shows four extra blending area sizes (quarter, half, double, and quadrupole of the existing area size) that are defined for the blending process. A CU level flag may be coded to signal the selected blending area size.

[0079] E. Template Matching Based Reordering for GPM Split Modes

[0080] In some embodiments, given the motion information of the current GPM block, the respective TM cost values of GPM split modes are computed. Then, all GPM split modes are reordered in ascending ordering based on the TM cost values. Instead of sending GPM split mode, an index using Golomb-Rice code is signaled to indicate where the exact GPM split mode is located in a reordering list. The reordering method for GPM split modes is a two-step process performed after the respective reference templates of the two GPM partitions in a coding unit are generated, as follows: (1) extending GPM partition edge into the reference templates of the two GPM partitions, resulting in 64 reference templates; (2) computing the respective TM cost for each of the 64 reference templates; (3) reordering GPM split modes based on their TM cost values in ascending order and marking the best 32 split modes as available split modes. FIG. 9 illustrates the GPM partitioning edge extended from the current CU to partition the template region. GPM blending process is not used in the template region around the edge.

[0081] F. GPM with Affine Prediction

[0082] In some embodiments, GPM is extended to enable affine motion compensation (AMC) . Therefore, a GPM partition can be predicted by AMC inter-prediction, non-AMC inter-prediction or intra-prediction. In addition, a GPM partition predicted by AMC can be combined with the other GPM partition predicted by AMC, non-AMC, or intra-prediction.

[0083] When AMC is applied, a uni-prediction affine merge candidate list is constructed from the subblock-based merge candidate list after discarding sub-TMVP candidates, similar to the uni-prediction merge candidate list construction for GPM in VVC. AMC is performed for a GPM partition using the control point motion vectors (CPMVs) of a merge candidate in the uni-prediction affine merge candidate list.

[0084] A gpm_affine_flag is signaled for each GPM partition to indicate whether AMC is applied for the GPM partition. A merge candidate index for the GPM partition is signaled using different arithmetic context models depending on whether AMC or non-AMC is applied.

[0085] G. Regression-Based GPM Blending

[0086] For some embodiments, regression-based GPM blending mode is provided as an additional GPM implicit mode, where the two integer blending matrices (W0 and W1) are derived from a template region (1 line above, 1 column left of the current block) . The blending matrices are modelled as an affine linear function of the sample positions (x, y) in the current CU: W0 (x, y) = a. x + b. y + c and W1 (x, y) = 1 -W0 (x, y)

[0087] The parameters (a, b, c) are derived from the reference template using a that minimizes mean-square-error (MSE) between corresponding input samples and output samples in the template region neighboring the current block. The MSE minimization is performed by calculating autocorrelation matrix for the input samples and a cross-correlation vector (with parameters a, b, c) between the input and output samples. Autocorrelation matrix is LDL decomposed and the final filter coefficients are calculated using back-substitution. The autocorrelation matrix is calculated using the reconstructed values of luma and / or chroma samples.

[0088] A list of pair of candidates is built from the regular GPM candidates and re-ordered with the template cost. The GPM implicit mode is signaled by a CU-level flag (gpm_implicit_flag) . If gpm_implicit_flag is true, a merge-idx is coded to signal the pair of GPM candidates to be used. If gpm_implicit_flag is false, the regular GPM syntax elements are signaled.

[0089] In some embodiments, an additional GPM implicit mode, where the two integer blending matrices (i.e., same as W0 and W1 described above) are derived from the template region, can be selected when GPM with affine prediction is enabled. In some embodiments, a flag is signaled after the gpm_affine_flag of both GPM partitions to indicate whether the regression-based GPM blending is enabled. When the flag and gpm_affine_flag of at least one GPM partition are enabled, the reference templates of both GPM partitions are used to derive W0 and W1 by minimizing the sample differences in template region and the GPM partitions predicted by AMC and non-AMC are blended according to W0 and W1 in the CU. For one or more than one GPM partitions predicted by AMC, the reference template is predicted in a subblock-based manner. That is, the reference template samples are fetched from the top or left reconstruction region of each affine prediction subblock. When the flag is enabled and the gpm_affine_flags of both GPM partitions are disabled, both GPM partitions predicted by non-AMC are blended according to W0 and W1 in the CU.

[0090] In some embodiments, GPM with affine prediction mode can be selected when regression-based GPM blending is enabled.

[0091] H. Lists of Pairings of Candidates for GPM

[0092] In some embodiments, a list of pairings of merge candidates is built, with the two candidates in each pairing being used to encode or decode the two GPM partitions respectively. In some embodiments, the list of pairings of candidates is built from the inter merge candidates and affine merge candidates. The pairings in the list may include pairings of two inter merge candidates, pairings of two affine merge candidates, and pairings of one inter merge candidate and one affine merge candidate, to be used for the derivation of regression-based blending weights. The list may be re-ordered by the template costs. If gpm_implicit_flag is true, a merge-idx is coded to signal a pairing of candidates for GPM prediction. If gpm_implicit_flag is false, the regular GPM syntax elements are signaled.

[0093] In some embodiments, two (first and second) lists of pairings of candidates are built from the inter merge candidates and affine merge candidates. The first list is built from the inter merge candidates and the second list is built from the inter merge and affine merge candidates excluding the pair of two inter merge candidates. If gpm_implicit_flag is true, an additional flag is signaled to indicate which list is used. A merge-idx is coded after the flag to signal the pair of GPM candidates in the selected list to be used for the derivation of regression-based blending weights. If gpm_implicit_flag is false, the regular GPM syntax elements are signaled.

[0094] In some embodiments, two lists of pairings of candidates are built from the inter merge candidates and affine merge candidates. The first list is built from the inter merge candidates and the second list is built from the affine merge candidates. If gpm_implicit_flag is true, an additional flag is signaled to indicate which list is used. A merge-idx (merge index) is coded after the flag to signal the pairing of GPM candidates in the selected list to be used for the derivation of regression-based blending weights.

[0095] In some embodiments, GPM with inter, intra, IBC, affine prediction mode can be selected when regression-based GPM blending is enabled. All the pairings of inter merge, intra, IBC and affine merge candidates are separated into several lists. If gpm_implicit_flag is true, several bits are signaled to indicate which list is used and a following merge-idx is signaled to select the candidate pair in the list for the derivation of regression-based blending weights.

[0096] In some embodiment, a list of pair of candidates is built from the inter merge, intra, IBC and affine merge candidates. The list is re-ordered by the template costs. If gpm_implicit_flag is true, a merge-idx is coded to signal the pair of GPM candidates, including the pair of two inter merge candidates, the pair of two intra candidates, the pair of two IBC candidates, the pair of two affine merge candidates, the pair of one inter merge candidate and one intra, IBC or affine merge candidate, the pair of one intra candidate and one inter merge, IBC or affine merge candidate, the pair of one IBC candidate and one inter merge, intra or affine merge candidate and the pair of one affine merge candidate and one inter merge, intra or IBC candidate, to be used for the derivation of regression-based blending weights. If gpm_implicit_flag is false, the regular GPM syntax elements are signaled.

[0097] In some embodiments, two lists of pair of candidates are built from the inter merge, intra, IBC and affine merge candidates. The 1st list is built from the inter merge candidates and the 2nd list is built from the inter merge, intra, IBC and affine merge candidates excluding the pair of two inter merge candidates. If gpm_implicit_flag is true, an additional flag is signaled to indicate which list is used.

[0098] In some embodiment, four lists of pair of candidates are built from the inter merge, intra, IBC and affine merge candidates. The first list is built from the inter merge candidates, the second list is built from intra candidates, the third list is built from IBC candidates and the fourth list is built from the affine merge candidates. If gpm_implicit_flag is true, two additional bits are used to indicate which list is used.

[0099] FIG. 10 conceptually illustrates one list of pairings of candidates that includes all possible pairings of merge candidates. Each entry of the list 1000 is a pairing of two merge candidates that may come from any of three coding tools A, B, and C. (Candidate A: 1, A: 2, etc. may represent inter merge candidates; B: 1, B: 2, etc. may represent affine merge candidates; C: 1, C: 2, etc. may represent IBC merge candidates) The list 1000 includes all or most possible pairings among the merge candidates. The candidates in the list 1000 may be further reordered and assigned indices based on template cost.

[0100] FIG. 11 conceptually illustrates multiple lists of pairings of candidates that correspond to different coding tools. The list 1110 includes all or most possible pairings with merge candidates of coding tools A. The list 1120 includes all or most possible pairings with merge candidates of coding tool B. The list 1130 includes all or most possible pairings with merge candidates of coding tool C. The candidates in the lists 1110, 1120, and 1130 may be further reordered and assigned indices based on template cost.

[0101] J. GPM with Local Illumination Compensation (LIC)

[0102] Local illumination Compensation (LIC) is an inter prediction technique to model local illumination variation between current block and its prediction block as a function of that between current block template and reference block template. The parameters of the function can be denoted by a scale α and an offset β, which forms a linear equation, that is, α*p [x] +β to compensate illumination changes, where p [x] is a reference sample pointed to by MV at a location x on reference picture. In some embodiments, since the parameters α and β can be derived based on current block template and reference block template, no signaling overhead is required for them. The video encoder may signal an LIC flag to enable or disable the use of LIC.

[0103] In some embodiments, GPM can be enabled with LIC. In that, the motion information of two GPM partitions includes LIC flag and LIC parameters. If the LIC process of two GPM partitions are both enabled, the corresponding LIC parameters can be applied respectively to two partitions without re-derived LIC parameters again in MC process. In some embodiments, if only one GPM partition contains enabling LIC flag and the corresponding LIC parameters, LIC can be directly applied to two GPM partitions. In that, the same LIC parameters will be applied on both GPM partitions.

[0104] In some embodiments, if only one GPM partition contains enabling LIC flag and the corresponding LIC parameters (inherited LIC parameters) . It can be applied to the corresponding GPM partition. And for the other GPM partition, a refined LIC parameters will be applied. The refined LIC parameters are derived by refining the inherited LIC parameters. In some embodiments, templates of current block (e.g., above and left neighboring regions) and the corresponding reference blocks are used to refine the inherited LIC parameters.

[0105] In some embodiments, a flag is signaled in the bitstream to indicate whether GPM can be enabled with LIC (i.e., cu_gpm_lic_flag) . If cu_gpm_lic_flag is enabled, only if the merge candidates include enabled LIC flag and the corresponding LIC parameters in the merge information can be inserted into the merge list for GPM.

[0106] Any of the foregoing proposed methods can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter / intra / prediction module of an encoder, and / or an inter / intra / prediction module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module. IV. Example Video Encoder

[0107] FIG. 12 illustrates an example video encoder 1200 that may implement GPM. As illustrated, the video encoder 1200 receives input video signal from a video source 1205 and encodes the signal into bitstream 1295. The video encoder 1200 has several components or modules for encoding the signal from the video source 1205, at least including some components selected from a transform module 1210, a quantization module 1211, an inverse quantization module 1214, an inverse transform module 1215, an intra-picture estimation module 1224, an intra-prediction module 1225, a motion compensation module 1230, a motion estimation module 1235, an in-loop filter 1245, a reconstructed picture buffer 1250, a MV buffer 1265, and a MV prediction module 1275, and an entropy encoder 1290. The motion compensation module 1230 and the motion estimation module 1235 are part of an inter-prediction module 1240. The intra-prediction module 1225 and the intra-prediction estimation module 1224 are part of a current picture prediction module 1220, which uses current picture reconstructed samples as reference samples for prediction of the current block.

[0108] In some embodiments, the modules 1210 –1290 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device or electronic apparatus. In some embodiments, the modules 1210 –1290 are modules of hardware circuits implemented by one or more integrated circuits (ICs) of an electronic apparatus. Though the modules 1210 –1290 are illustrated as being separate modules, some of the modules can be combined into a single module.

[0109] The video source 1205 provides a raw video signal that presents pixel data of each video frame without compression. A subtractor 1208 computes the difference between the raw video pixel data of the video source 1205 and the predicted pixel data 1213 from the motion compensation module 1230 or intra-prediction module 1225 as prediction residual 1209. The transform module 1210 converts the difference (or the residual pixel data or residual signal 1208) into transform coefficients (e.g., by performing Discrete Cosine Transform, or DCT) . The quantization module 1211 quantizes the transform coefficients into quantized data (or quantized coefficients) 1212, which is encoded into the bitstream 1295 by the entropy encoder 1290.

[0110] The inverse quantization module 1214 de-quantizes the quantized data (or quantized coefficients) 1212 to obtain transform coefficients 1218, and the inverse transform module 1215 performs inverse transform on the transform coefficients 1218 to produce reconstructed residual 1219. The reconstructed residual 1219 is added with the predicted pixel data 1213 to produce reconstructed pixel data 1217. In some embodiments, the reconstructed pixel data 1217 is temporarily stored in a line buffer 1227 (or intra prediction buffer) for intra-picture prediction and spatial MV prediction. The reconstructed pixels are filtered by the in-loop filter 1245 and stored in the reconstructed picture buffer 1250. In some embodiments, the reconstructed picture buffer 1250 is a storage external to the video encoder 1200. In some embodiments, the reconstructed picture buffer 1250 is a storage internal to the video encoder 1200.

[0111] The intra-picture estimation module 1224 performs intra-prediction based on the reconstructed pixel data 1217 to produce intra prediction data. The intra-prediction data is provided to the entropy encoder 1290 to be encoded into bitstream 1295. The intra-prediction data is also used by the intra-prediction module 1225 to produce the predicted pixel data 1213.

[0112] The motion estimation module 1235 performs inter-prediction by producing MVs to reference pixel data of previously decoded frames stored in the reconstructed picture buffer 1250. These MVs are provided to the motion compensation module 1230 to produce predicted pixel data.

[0113] Instead of encoding the complete actual MVs in the bitstream, the video encoder 1200 uses MV prediction to generate predicted MVs, and the difference between the MVs used for motion compensation and the predicted MVs is encoded as residual motion data and stored in the bitstream 1295.

[0114] The MV prediction module 1275 generates the predicted MVs based on reference MVs that were generated for encoding previously video frames, i.e., the motion compensation MVs that were used to perform motion compensation. The MV prediction module 1275 retrieves reference MVs from previous video frames from the MV buffer 1265. The video encoder 1200 stores the MVs generated for the current video frame in the MV buffer 1265 as reference MVs for generating predicted MVs.

[0115] The MV prediction module 1275 uses the reference MVs to create the predicted MVs. The predicted MVs can be computed by spatial MV prediction or temporal MV prediction. The difference between the predicted MVs and the motion compensation MVs (MC MVs) of the current frame (residual motion data) are encoded into the bitstream 1295 by the entropy encoder 1290.

[0116] The entropy encoder 1290 encodes various parameters and data into the bitstream 1295 by using entropy-coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman encoding. The entropy encoder 1290 encodes various header elements, flags, along with the quantized transform coefficients 1212, and the residual motion data as syntax elements into the bitstream 1295. The bitstream 1295 is in turn stored in a storage device or transmitted to a decoder over a communications medium such as a network.

[0117] The in-loop filter 1245 performs filtering or smoothing operations on the reconstructed pixel data 1217 to reduce the artifacts of coding, particularly at boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 1245 include deblock filter (DBF) , sample adaptive offset (SAO) , and / or adaptive loop filter (ALF) . In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.

[0118] FIG. 13 illustrates portions of the video encoder 1200 that implement GPM with blending by regression. The figure illustrates the components that generates a GPM prediction as the predicted pixel data 1213. As illustrated, at least two predictions 1310 and 1320 are generated for at least two GPM partitions. Each partition prediction may be generated by intra prediction 1220, or inter prediction 1240, or current picture referencing (IBC) prediction 1312. Inter-prediction may be further modified by local illumination compensation (LIC) module 1314. The predictions may be performed at sub-block level to implement affine merge mode or affine AMVP mode for each GPM partition.

[0119] Each partition prediction is generated according to a candidate for a coding tool. Thus, for example, the first partition prediction 1310 may be generated by an affine merge candidate, while the second partition prediction 1320 may be generated by an inter merge candidate. The encoder selects a pairing of candidates from a list (or lists) of pairing candidates 1330 by providing a merge index, as the list of pairings of candidates 1330 may be reordered according to template costs. The merge index may be provided by the entropy encoder 1290 when regression-based GPM blending is enabled.

[0120] The two partition predictions 1310 and 1320 are brought together according to GPM partition parameters (provided by the entropy encoder 1290) to form the GPM prediction by a GPM blending module 1350. The samples along the GPM partitioning boundary  / edge are blended according to a weighting factor (blending weight) that is provided by a regression module 1360. The regression module 1360 uses corresponding input and output samples in the template region to perform MSE minimization to determine the weighting factor.

[0121] When LIC parameters is available for only one of the two GPM partitions, a LIC parameter refinement module 1340 may refine the LIC parameters for the other GPM partition based on samples of neighboring template and corresponding reference blocks stored in the reconstructed picture buffer 1250 or the line buffer 1227.

[0122] FIG. 14 conceptually illustrates a process 1400 for performing GPM prediction with regression-based blending weight using a list of pairings of candidates. In some embodiments, one or more processing units (e.g., a processor) of a computing device implementing the encoder 1200 performs the process 1400 by executing instructions stored in a computer readable medium. In some embodiments, an electronic apparatus implementing the encoder 1200 performs the process 1400.

[0123] The encoder receives (at block 1410) data to be encoded as a current block of pixels in a current picture. The encoder partitions (at block 1420) the current block into at least two geometric partition mode (GPM) partitions.

[0124] The encoder derives (at block 1430) a blending weight by regression using samples of a template region neighboring the current block for the partitioning. In some embodiments, the encoder performs MSE minimization based on corresponding input and output samples in the template region to determine the weighting factor.

[0125] The encoder selects (at block 1440) a pairing of candidates from a list of pairings of candidates. In some embodiments, the encoder selects the pairing of candidates by selecting from among a plurality of lists of pairings of candidates, with each pairing of candidates of a particular list including at least one candidate of a particular coding tool. In some embodiments, the encoder selects a pairing of candidates by selecting from only one list of pairings of candidates that includes pairings of any two candidates for any of two or more coding tools. For example, the list of pairings of candidates may include different pairings of candidates from inter merge candidates and affine merge candidates. For another example, the list of pairings of candidates may include different pairings of candidates from inter merge candidates, affine merge candidates, intra merge candidates, and intra block copy (IBC) merge candidates.

[0126] In some embodiments, the list of pairings of candidates is reordered according to template cost and each pairing of candidates in the list is assigned a corresponding merge index, such that the encoder may select a pairing of candidates by signaling a merge index to indicate the selected pairing of candidates in the list. In some embodiments, the merge index for the list of pairings of candidates is signaled when a flag (e.g., gpm_implicit_flag) is signaled to indicate that the blending weight is derived by regression.

[0127] In some embodiments, when only a first GPM partition of the at least two GPM partitions of the current block has a LIC enabling flag and a set of inherited LIC parameters, the encoder refines the set of inherited LIC parameters and applies the refined set of LIC parameters to a second GPM partition of the at least two GPM partitions. A neighboring template of the current block and a corresponding reference block may be used to refine the inherited LIC parameters for the second GPM partition.

[0128] The encoder applies (at block 1450) the selected pairing of candidates to the two GPM partitions with the derived blending weight to generate a predictor of the current block. The encoder encodes (at block 1460) the current block using the generated predictor to produce prediction residuals. V. Example Video Decoder

[0129] In some embodiments, an encoder may signal (or generate) one or more syntax element in a bitstream, such that a decoder may parse said one or more syntax element from the bitstream.

[0130] FIG. 15 illustrates an example video decoder 1500 that may implement GPM. As illustrated, the video decoder 1500 is an image-decoding or video-decoding circuit that receives a bitstream 1595 and decodes the content of the bitstream into pixel data of video frames for display. The video decoder 1500 has several components or modules for decoding the bitstream 1595, including some components selected from an inverse quantization module 1514, an inverse transform module 1515, an intra-prediction module 1525, a motion compensation module 1530, an in-loop filter 1545, a decoded picture buffer 1550, a MV buffer 1565, a MV prediction module 1575, and a parser 1590. The motion compensation module 1530 is part of an inter-prediction module 1540. The intra-prediction module 1525 is part of a current picture prediction module 1520, which uses current picture reconstructed samples as reference samples for prediction of the current block.

[0131] In some embodiments, the modules 1514 –1590 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device. In some embodiments, the modules 1514 –1590 are modules of hardware circuits implemented by one or more ICs of an electronic apparatus. Though the modules 1514 –1590 are illustrated as being separate modules, some of the modules can be combined into a single module.

[0132] The parser 1590 (or entropy decoder) receives the bitstream 1595 and performs initial parsing according to the syntax defined by a video-coding or image-coding standard. The parsed syntax element includes various header elements, flags, as well as quantized data (or quantized coefficients) 1512. The parser 1590 parses out the various syntax elements by using entropy-coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman encoding.

[0133] The inverse quantization module 1514 de-quantizes the quantized data (or quantized coefficients) 1512 to obtain transform coefficients, and the inverse transform module 1515 performs inverse transform on the transform coefficients 1518 to produce reconstructed residual signal 1519. The reconstructed residual signal 1519 is added with predicted pixel data 1513 from the intra-prediction module 1525 or the motion compensation module 1530 to produce decoded pixel data 1517. The decoded pixels data are filtered by the in-loop filter 1545 and stored in the decoded picture buffer 1550. In some embodiments, the decoded picture buffer 1550 is a storage external to the video decoder 1500. In some embodiments, the decoded picture buffer 1550 is a storage internal to the video decoder 1500.

[0134] The intra-prediction module 1525 receives intra-prediction data from bitstream 1595 and according to which, produces the predicted pixel data 1513 from the decoded pixel data 1517 stored in the decoded picture buffer 1550. In some embodiments, the decoded pixel data 1517 is also stored in a line buffer 1527 (or intra prediction buffer) for intra-picture prediction and spatial MV prediction.

[0135] In some embodiments, the content of the decoded picture buffer 1550 is used for display. A display device 1505 either retrieves the content of the decoded picture buffer 1550 for display directly, or retrieves the content of the decoded picture buffer to a display buffer. In some embodiments, the display device receives pixel values from the decoded picture buffer 1550 through a pixel transport.

[0136] The motion compensation module 1530 produces predicted pixel data 1513 from the decoded pixel data 1517 stored in the decoded picture buffer 1550 according to motion compensation MVs (MC MVs) . These motion compensation MVs are decoded by adding the residual motion data received from the bitstream 1595 with predicted MVs received from the MV prediction module 1575.

[0137] The MV prediction module 1575 generates the predicted MVs based on reference MVs that were generated for decoding previous video frames, e.g., the motion compensation MVs that were used to perform motion compensation. The MV prediction module 1575 retrieves the reference MVs of previous video frames from the MV buffer 1565. The video decoder 1500 stores the motion compensation MVs generated for decoding the current video frame in the MV buffer 1565 as reference MVs for producing predicted MVs.

[0138] The in-loop filter 1545 performs filtering or smoothing operations on the decoded pixel data 1517 to reduce the artifacts of coding, particularly at boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 1545 include deblock filter (DBF) , sample adaptive offset (SAO) , and / or adaptive loop filter (ALF) . In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.

[0139] FIG. 16 illustrates portions of the video decoder 1500 that implement GPM with blending by regression. The figure illustrates the components that generates a GPM prediction as the predicted pixel data 1513. As illustrated, at least two predictions 1610 and 1620 are generated for at least two GPM partitions. Each partition prediction may be generated by intra prediction 1520, or inter prediction 1540, or current picture referencing (IBC) prediction 1612. Inter-prediction may be further modified by local illumination compensation (LIC) module 1614. The predictions may be performed at sub-block level to implement affine merge mode or affine AMVP mode for each GPM partition.

[0140] Each partition prediction is generated according to a candidate for a coding tool. Thus, for example, the first partition prediction 1610 may be generated by an affine merge candidate, while the second partition prediction 1620 may be generated by an inter merge candidate. The decoder selects a pairing of candidates from a list (or lists) of pairing candidates 1630 by providing a merge index, as the list of pairings of candidates 1630 may be reordered according to template costs. The merge index may be provided by the entropy decoder 1590 when regression-based GPM blending is enabled.

[0141] The two partition predictions 1610 and 1620 are brought together according to GPM partition parameters (provided by the entropy decoder 1590) to form the GPM prediction by a GPM blending module 1650. The samples along the GPM partitioning boundary  / edge are blended according to a weighting factor (blending weight) that is provided by a regression module 1660. The regression module 1660 uses corresponding input and output samples in the template region to perform MSE minimization to determine the weighting factor.

[0142] When LIC parameters is available for only one of the two GPM partitions, a LIC parameter refinement module 1640 may refine the LIC parameters for the other GPM partition based on samples of neighboring template and corresponding reference blocks stored in the reconstructed picture buffer 1550 or the line buffer 1527.

[0143] FIG. 17 conceptually illustrates a process 1700 for performing GPM prediction with regression-based blending weight using a list of pairings of candidates. In some embodiments, one or more processing units (e.g., a processor) of a computing device implementing the decoder 1500 performs the process 1700 by executing instructions stored in a computer readable medium. In some embodiments, an electronic apparatus implementing the decoder 1500 performs the process 1700.

[0144] The decoder receives (at block 1710) data to be decoded as a current block of pixels in a current picture. The decoder partitions (at block 1720) the current block into at least two geometric partition mode (GPM) partitions.

[0145] The decoder derives (at block 1730) a blending weight by regression using samples of a template region neighboring the current block for the partitioning. In some embodiments, the decoder performs MSE minimization based on corresponding input and output samples in the template region to determine the weighting factor.

[0146] The decoder selects (at block 1740) a pairing of candidates from a list of pairings of candidates. In some embodiments, the decoder selects the pairing of candidates by selecting from among a plurality of lists of pairings of candidates, with each pairing of candidates of a particular list including at least one candidate of a particular coding tool. In some embodiments, the decoder selects a pairing of candidates by selecting from only one list of pairings of candidates that includes pairings of any two candidates for any of two or more coding tools. For example, the list of pairings of candidates may include different pairings of candidates from inter merge candidates and affine merge candidates. For another example, the list of pairings of candidates may include different pairings of candidates that each of the candidates is selected from inter merge candidates, affine merge candidates, intra merge candidates, and intra block copy (IBC) merge candidates.

[0147] In some embodiments, the list of pairings of candidates is reordered according to template cost and each pairing of candidates in the list is assigned a corresponding merge index, such that the decoder may select a pairing of candidates by signaling a merge index to indicate the selected pairing of candidates in the list. In some embodiments, the merge index for the list of pairings of candidates is signaled when a flag (e.g., gpm_implicit_flag) is signaled to indicate that the blending weight is derived by regression.

[0148] In some embodiments, when only a first GPM partition of the at least two GPM partitions of the current block has a LIC enabling flag and a set of inherited LIC parameters, the decoder refines the set of inherited LIC parameters and applies the refined set of LIC parameters to a second GPM partition of the at least two GPM partitions. A neighboring template of the current block and a corresponding reference block may be used to refine the inherited LIC parameters for the second GPM partition.

[0149] The decoder applies (at block 1750) the selected pairing of candidates to the two GPM partitions with the derived blending weight to generate a predictor of the current block. The decoder reconstructs (at block 1760) the current block by combining the generated predictor with prediction residuals. The decoder may then provide the reconstructed current block for display as part of the reconstructed current picture. VI. Example Electronic System

[0150] Many of the above-described features and applications are implemented as software processes that are specified as a set of instructions recorded on a computer readable storage medium (also referred to as computer readable medium) . When these instructions are executed by one or more computational or processing unit (s) (e.g., one or more processors, cores of processors, or other processing units) , they cause the processing unit (s) to perform the actions indicated in the instructions. Examples of computer readable media include, but are not limited to, CD-ROMs, flash drives, random-access memory (RAM) chips, hard drives, erasable programmable read only memories (EPROMs) , electrically erasable programmable read-only memories (EEPROMs) , etc. The computer readable media does not include carrier waves and electronic signals passing wirelessly or over wired connections.

[0151] In this specification, the term “software” is meant to include firmware residing in read-only memory or applications stored in magnetic storage which can be read into memory for processing by a processor. Also, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while remaining distinct software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Finally, any combination of separate programs that together implement a software invention described here is within the scope of the present disclosure. In some embodiments, the software programs, when installed to operate on one or more electronic systems, define one or more specific machine implementations that execute and perform the operations of the software programs.

[0152] FIG. 18 conceptually illustrates an electronic system 1800 with which some embodiments of the present disclosure are implemented. The electronic system 1800 may be a computer (e.g., a desktop computer, personal computer, tablet computer, etc. ) , phone, PDA, or any other sort of electronic device. Such an electronic system includes various types of computer readable media and interfaces for various other types of computer readable media. Electronic system 1800 includes a bus 1805, processing unit (s) 1810, a graphics-processing unit (GPU) 1815, a system memory 1820, a network 1825, a read-only memory 1830, a permanent storage device 1835, input devices 1840, and output devices 1845.

[0153] The bus 1805 collectively represents all system, peripheral, and chipset buses that communicatively connect the numerous internal devices of the electronic system 1800. For instance, the bus 1805 communicatively connects the processing unit (s) 1810 with the GPU 1815, the read-only memory 1830, the system memory 1820, and the permanent storage device 1835.

[0154] From these various memory units, the processing unit (s) 1810 retrieves instructions to execute and data to process in order to execute the processes of the present disclosure. The processing unit (s) may be a single processor or a multi-core processor in different embodiments. Some instructions are passed to and executed by the GPU 1815. The GPU 1815 can offload various computations or complement the image processing provided by the processing unit (s) 1810.

[0155] The read-only-memory (ROM) 1830 stores static data and instructions that are used by the processing unit (s) 1810 and other modules of the electronic system. The permanent storage device 1835, on the other hand, is a read-and-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the electronic system 1800 is off. Some embodiments of the present disclosure use a mass-storage device (such as a magnetic or optical disk and its corresponding disk drive) as the permanent storage device 1835.

[0156] Other embodiments use a removable storage device (such as a floppy disk, flash memory device, etc., and its corresponding disk drive) as the permanent storage device. Like the permanent storage device 1835, the system memory 1820 is a read-and-write memory device. However, unlike storage device 1835, the system memory 1820 is a volatile read-and-write memory, such a random access memory. The system memory 1820 stores some of the instructions and data that the processor uses at runtime. In some embodiments, processes in accordance with the present disclosure are stored in the system memory 1820, the permanent storage device 1835, and / or the read-only memory 1830. For example, the various memory units include instructions for processing multimedia clips in accordance with some embodiments. From these various memory units, the processing unit (s) 1810 retrieves instructions to execute and data to process in order to execute the processes of some embodiments.

[0157] The bus 1805 also connects to the input and output devices 1840 and 1845. The input devices 1840 enable the user to communicate information and select commands to the electronic system. The input devices 1840 include alphanumeric keyboards and pointing devices (also called “cursor control devices” ) , cameras (e.g., webcams) , microphones or similar devices for receiving voice commands, etc. The output devices 1845 display images generated by the electronic system or otherwise output data. The output devices 1845 include printers and display devices, such as cathode ray tubes (CRT) or liquid crystal displays (LCD) , as well as speakers or similar audio output devices. Some embodiments include devices such as a touchscreen that function as both input and output devices.

[0158] Finally, as shown in FIG. 18, bus 1805 also couples electronic system 1800 to a network 1825 through a network adapter (not shown) . In this manner, the computer can be a part of a network of computers (such as a local area network ( “LAN” ) , a wide area network ( “WAN” ) , or an Intranet, or a network of networks, such as the Internet. Any or all components of electronic system 1800 may be used in conjunction with the present disclosure.

[0159] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a machine-readable or computer-readable medium (alternatively referred to as computer-readable storage media, machine-readable media, or machine-readable storage media) . Some examples of such computer-readable media include RAM, ROM, read-only compact discs (CD-ROM) , recordable compact discs (CD-R) , rewritable compact discs (CD-RW) , read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM) , a variety of recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc. ) , flash memory (e.g., SD cards, mini-SD cards, micro-SD cards, etc. ) , magnetic and / or solid state hard drives, read-only and recordable discs, ultra-density optical discs, any other optical or magnetic media, and floppy disks. The computer-readable media may store a computer program that is executable by at least one processing unit and includes sets of instructions for performing various operations. Examples of computer programs or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter.

[0160] While the above discussion primarily refers to microprocessor or multi-core processors that execute software, many of the above-described features and applications are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) . In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In addition, some embodiments execute software stored in programmable logic devices (PLDs) , ROM, or RAM devices.

[0161] As used in this specification and any claims of this application, the terms “computer” , “server” , “processor” , and “memory” all refer to electronic or other technological devices. These terms exclude people or groups of people. For the purposes of the specification, the terms display or displaying means displaying on an electronic device. As used in this specification and any claims of this application, the terms “computer readable medium, ” “computer readable media, ” and “machine readable medium” are entirely restricted to tangible, physical objects that store information in a form that is readable by a computer. These terms exclude any wireless signals, wired download signals, and any other ephemeral signals.

[0162] While the present disclosure has been described with reference to numerous specific details, one of ordinary skill in the art will recognize that the present disclosure can be embodied in other specific forms without departing from the spirit of the present disclosure. In addition, a number of the figures (including FIG. 14 and FIG. 17) conceptually illustrate processes. The specific operations of these processes may not be performed in the exact order shown and described. The specific operations may not be performed in one continuous series of operations, and different specific operations may be performed in different embodiments. Furthermore, the process could be implemented using several sub-processes, or as part of a larger macro process. Thus, one of ordinary skill in the art would understand that the present disclosure is not to be limited by the foregoing illustrative details, but rather is to be defined by the appended claims. Additional Notes

[0163] The herein-described subject matter sometimes illustrates different components contained within, or connected with, different other components. It is to be understood that such depicted architectures are merely examples, and that in fact many other architectures can be implemented which achieve the same functionality. In a conceptual sense, any arrangement of components to achieve the same functionality is effectively "associated" such that the desired functionality is achieved. Hence, any two components herein combined to achieve a particular functionality can be seen as "associated with" each other such that the desired functionality is achieved, irrespective of architectures or intermediate components. Likewise, any two components so associated can also be viewed as being "operably connected" , or "operably coupled" , to each other to achieve the desired functionality, and any two components capable of being so associated can also be viewed as being "operably couplable" , to each other to achieve the desired functionality. Specific examples of operably couplable include but are not limited to physically mateable and / or physically interacting components and / or wirelessly interactable and / or wirelessly interacting components and / or logically interacting and / or logically interactable components.

[0164] Further, with respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity.

[0165] Moreover, it will be understood by those skilled in the art that, in general, terms used herein, and especially in the appended claims, e.g., bodies of the appended claims, are generally intended as “open” terms, e.g., the term “including” should be interpreted as “including but not limited to, ” the term “having” should be interpreted as “having at least, ” the term “includes” should be interpreted as “includes but is not limited to, ” etc. It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles "a" or "an" limits any particular claim containing such introduced claim recitation to implementations containing only one such recitation, even when the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an, " e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more; ” the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number, e.g., the bare recitation of "two recitations, " without other modifiers, means at least two recitations, or two or more recitations. Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention, e.g., “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. In those instances where a convention analogous to “at least one of A, B, or C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention, e.g., “a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B. ”

[0166] From the foregoing, it will be appreciated that various implementations of the present disclosure have been described herein for purposes of illustration, and that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various implementations disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

1.A video coding method comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video;partitioning the current block into first and second geometric partition mode (GPM) partitions;deriving a blending weight by regression using samples of a template region neighboring the current block for the partitioning;selecting a pairing of candidates from a list of pairings of candidates;applying the selected pairing of candidates to the first and second GPM partitions with the derived blending weight to generate a predictor of the current block; andencoding or decoding the current block using the generated predictor.2.The video coding method of claim 1, wherein the list of pairings of candidates comprises different pairings of candidates from inter merge candidates and affine merge candidates.3.The video coding method of claim 1, wherein the list of pairings of candidates comprises different pairings of candidates from inter merge candidates, affine merge candidates, intra merge candidates, and intra block copy (IBC) merge candidates.4.The video coding method of claim 1, wherein the list of pairings of candidates is reordered according to template cost and each pairing of candidates in the list is assigned a corresponding merge index.5.The video coding method of claim 1, wherein selecting a pairing of candidates comprises signaling a merge index to indicate the pairing of candidates in the list of pairings of candidates.6.The video coding method of claim 5, wherein the merge index is signaled when a flag is signaled to indicate that the blending weight is derived by regression.7.The video coding method of claim 1, wherein selecting a pairing of candidates comprises selecting from among a plurality of lists of pairings of candidates, wherein each pairing of candidates of a particular list includes at least one candidate of a particular coding tool.8.The video coding method of claim 1, wherein selecting a pairing of candidates comprises selecting from only one list of pairings of candidates that includes pairings of any two candidates for any of two or more coding tools.9.The video coding method of claim 1, wherein only the first GPM partition of the current block has a LIC enabling flag and a set of inherited LIC parameters, the method further comprising refining the set of inherited LIC parameters and applying the refined set of LIC parameters to the second GPM partition.10.The video coding method of claim 9, wherein a neighboring template of the current block and a corresponding reference block are used to refine the inherited LIC parameters for the second GPM partition.11.An electronic apparatus comprising:a video coder circuit configured to perform operations comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video;partitioning the current block into first and second geometric partition mode (GPM) partitions;deriving a blending weight by regression using samples of a template region neighboring the current block for the partitioning;selecting a pairing of candidates from a list of pairings of candidates;applying the selected pairing of candidates to the first and second GPM partitions with the derived blending weight to generate a predictor of the current block; andencoding or decoding the current block using the generated predictor.12.A video decoding method comprising:receiving data to be decoded as a current block of pixels of a current picture of a video;partitioning the current block into first and second geometric partition mode (GPM) partitions;deriving a blending weight by regression using samples of a template region neighboring the current block for the partitioning;selecting a pairing of candidates from a list of pairings of candidates;applying the selected pairing of candidates to the first and second GPM partitions with the derived blending weight to generate a predictor of the current block; andreconstructing the current block using the generated predictor.13.A video encoding method comprising:receiving data to be encoded as a current block of pixels of a current picture of a video;partitioning the current block into first and second geometric partition mode (GPM) partitions;deriving a blending weight by regression using samples of a template region neighboring the current block for the partitioning;selecting a pairing of candidates from a list of pairings of candidates;applying the selected pairing of candidates to the first and second GPM partitions with the derived blending weight to generate a predictor of the current block; andencoding the current block using the generated predictor.

Citation Information

Patent Citations

  • Image encoding / decoding method and device, and recording medium storing bit stream

    CN114731409A

  • Geometric partition mode with motion vector refinement

    WO2022245876A1

  • GPM-based image coding method and device

    WO2023055126A1