Method and apparatus for prioritized initial motion vector for decoder-side motion refinement in video coding
By prioritizing the initial motion vectors on the decoder side and combining multiple encoding and decoding tools, the problem of low motion vector prediction efficiency in VVC is solved, achieving a more efficient video encoding and decoding process and improved quality.
Patent Information
- Application Number
- CN202480025460.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-24
- Filing Date
- 2024-02-07
- Publication Date
- 2025-11-21
AI Technical Summary
Existing video coding and decoding technologies suffer from inefficiency in motion vector prediction, especially in the Multifunctional Video Coding (VVC) standard, where the selection and refinement of motion vectors have not been fully optimized, resulting in low coding and decoding efficiency.
By prioritizing the initial motion vector on the decoder side, and by biasing the initial motion vector among the motion vector candidates, combined with various encoding and decoding tools such as bilateral matching, template matching, and overlapping block motion compensation, the accuracy of motion vectors and encoding and decoding efficiency are improved.
It improves the accuracy of motion vector prediction and the efficiency of encoding and decoding in the video encoding and decoding process, thereby enhancing video quality and compression performance.
Smart Images

Figure CN121002845A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a non-provisional of U.S. Provisional Patent Application No. 63 / 484,546 (filed February 13, 2023), U.S. Provisional Patent Application No. 63 / 495,779 (filed April 13, 2023), and U.S. Provisional Patent Application No. 63 / 497,754 (filed April 24, 2023), and claims priority thereto. The above U.S. Provisional Patent Applications are incorporated herein by reference in their entirety. TECHNICAL FIELD The present disclosure relates to video coding. In particular, the present disclosure relates to prioritizing an initial motion vector in a video coding system to bias towards the initial motion vector in a set of motion vector candidates. BACKGROUND Versatile Video Coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the International Telecommunication Union Video Coding Experts Group (VCEG) and ISO / IEC Moving Picture Experts Group (MPEG). The standard has been published as an ISO standard: ISO / IEC 23090-3:2021, Information technology - Coding of immersive media representations - Part 3: Versatile Video Coding, published in February 2021. VVC is developed based on its predecessor, High Efficiency Video Coding (HEVC), by adding more coding tools to improve coding efficiency and handle various types of video sources, including three-dimensional (3D) video signals.
[0004] Figure 1AAn exemplary adaptive inter / intra video coding system incorporating in-loop processing is shown. For intra prediction 110, prediction data is derived based on previously coded video data in the current picture. For inter prediction 112, motion estimation (ME) is performed at the encoder side and motion compensation (MC) is performed based on the results of ME to provide prediction data derived from other pictures and motion data. A switch 114 selects either intra prediction 110 or inter prediction 112, and the selected prediction data is provided to a summer 116 to form prediction error, also known as residual. The prediction error then goes through transform (T) 118 followed by quantization (Q) 120 processing. The transformed and quantized residual is then coded by an entropy coder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with additional information, such as motion and coding modes associated with intra and inter predictions, and other information associated with in-loop filters applied to regions of the underlying pictures. The additional information associated with intra prediction 110, inter prediction 112, and in-loop filters 130 is provided to entropy coder 122 as shown. Figure 1A When inter prediction mode is used, the reference picture or pictures also have to be reconstructed at the encoder side. Therefore, the transformed and quantized residual goes through inverse quantization (IQ) 124 and inverse transform (IT) 126 processing to recover the residual. The residual is then added back to the prediction data 136 at reconstruction (REC) 128 to reconstruct the video data. The reconstructed video data can be stored in a reference picture buffer 134 and used for prediction of other frames.
[0005] As Figure 1AAs shown, the input video data undergoes a series of processes in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to these processes. Therefore, a loop filter 130 is typically applied to the reconstructed video data before it is stored in the reference picture buffer 134 to improve video quality. For example, a deblocking filter (DF), Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF) may be used. Loop filter information may need to be included in the bitstream so that the decoder can correctly recover the required information. Therefore, loop filter information is also provided to the entropy encoder 122 for inclusion in the bitstream. Figure 1A In the process, the loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference image buffer 134. Figure 1A The system described is intended to demonstrate an exemplary architecture of a typical video encoder. It may correspond to a High Efficiency Video Codec (HEVC) system, VP8, VP9, H.264, or VVC.
[0006] like Figure 1B As shown, the decoder can use the same or partially the same functional modules as the encoder, except for transform 118 and quantization 120, because the decoder only needs inverse quantization 124 and inverse transform 126. The decoder uses entropy decoder 140 to decode the video bitstream into quantized transform coefficients and the required encoding / decoding information (e.g., ILPF information, intra-frame prediction information, and inter-frame prediction information). Intra-frame prediction 150 at the decoder end does not need to perform mode search. Instead, the decoder only needs to generate intra-frame predictions based on the intra-frame prediction information received from entropy decoder 140. Furthermore, for inter-frame prediction, the decoder only needs to perform motion compensation (MC 152) based on the inter-frame prediction information received from entropy decoder 140, without performing motion estimation.
[0007] In High Efficiency Video Coding (HEVC), a picture is divided into a series of Coding Tree Units (CTUs). A CTU consists of one NxN block of luma samples and two corresponding blocks of chroma samples for pictures with three sample arrays, or one NxN block of monochrome planar samples for pictures coded using three independent color planes. The concept of CTU is roughly analogous to the macroblock in previous standards such as Advanced Video Coding (AVC). In the main profile, the maximum allowed size of the luma block in a CTU is specified as 64x64. A CTU is partitioned using a quad-tree structure (called coding tree) to adapt to various local characteristics. The decision whether to use inter (temporal) or intra (spatial) prediction to code a picture region is made at the leaf CU level. Each leaf CU can be further partitioned into one, two or four Prediction Units (PUs) according to the PU partition type. Within a PU, the same prediction process is applied and the related information is transmitted to the decoder on a PU basis. After obtaining the residual block after applying the prediction process according to the PU partition type, a leaf CU can be partitioned into Transform Units (TUs) according to another quad-tree structure similar to the coding tree of the CU. One key feature of the HEVC structure is that it has a multi- partition concept that includes CUs, PUs and TUs.
[0008] In VVC, the concept of multiple partition unit types is replaced by quad-tree with binary and ternary splitting structure and nested multi-type tree, i.e., the separation of CU, PU and TU concepts is removed, except for the need to handle CU of excessive size to accommodate the maximum transform length, and more flexibility in CU partition shape is supported. In the coding tree structure, a CU can be square or rectangular. A Coding Tree Unit (CTU) is first partitioned by a quad-tree (also called quad-tree) structure. Then, the quad-tree leaf nodes can be further partitioned by a multi-type tree structure. As Figure 2As shown, there are four types of splitting in the multi-type tree structure, vertical binary splitting (SPLIT_BT_VER 210), horizontal binary splitting (SPLIT_BT_HOR 220), vertical ternary splitting (SPLIT_TT_VER 230), and horizontal ternary splitting (SPLIT_TT_HOR 240). Multi-type tree leaf nodes are referred to as coding units (CUs), and this type of splitting is used for prediction and transform processing without further partitioning, unless the CU is too large to exceed the maximum transform length. This means that, in most cases, CUs, PUs, and TUs have the same block size in the quadtree plus nested multi-type tree coding block structure. The exception occurs when the maximum supported transform length is smaller than the width or height of the color component of the CU.
[0009] Figure 3 A CTU is shown being divided into CUs using a quadtree and nested multi-type tree coding block structure, where the bold block edges represent quadtree partitioning and the remaining edges represent multi-type tree partitioning. The quadtree with nested multi-type tree partitioning provides a content-adaptive coding tree structure consisting of CUs. The size of a CU can be as large as a CTU or as small as 4x4 in luma samples. For a 4:2:0 chroma format, the maximum chroma CB size is 64x64, and the smallest size chroma CB consists of 16 chroma samples.
[0010] In VVC, the maximum supported luma transform size is 64x64, and the maximum supported chroma transform size is 32x32. When the width or height of a CB is larger than the maximum transform width or height, the CB is automatically split in the horizontal and / or vertical direction to meet the transform size limit in that direction.
[0011] In VVC, the coding tree scheme supports independent block tree structures for luma and chroma. For P and B slices, the luma and chroma CTBs in a CTU must share the same coding tree structure. However, for I slices, luma and chroma can have independent block tree structures. When the independent block tree mode is applied, the luma CTBs are partitioned into CUs by one coding tree structure, while the chroma CTBs are partitioned into chroma CUs by another coding tree structure. This means that in an I slice, a CU can consist of either a coding block of luma components or two coding blocks of chroma components, while in a P or B slice, a CU always consists of coding blocks of all three color components, unless the video is monochrome.
[0012] For each inter prediction CU, the motion parameters consist of the motion vector, the reference picture index and the reference picture list usage index, and additional information needed for the VVC new coding features for inter prediction sample generation. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded in skip mode, the CU is associated with one PU and there is no significant residual coefficient, no coded motion vector delta or reference picture index. A merge mode is specified, in which the motion parameters of the current CU are derived from neighboring CUs, including spatial and temporal candidates, and additional plans introduced in VVC. The merge mode can be applied to any inter prediction CU, not only in skip mode. The alternative to the merge mode is the explicit transmission of the motion parameters, in which the motion vector, the corresponding reference picture index for each reference picture list and the reference picture list usage flag, and other needed information are explicitly signaled in each CU.
[0013] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 5) are studying the potential need for standardization of future video coding technology with compression capability significantly above that of current VVC standard. The enhanced compression model (ECM) reference software aims to demonstrate the reference implementation of coding techniques and decoding processes for the exploration of VVC enhancement compression beyond VVC capability. The reference software is accessible through https: / / vcgit.hhi.fraunhofer.de / ecm / ECM.git. The ECM is basically a successor of VVC, so it shares many common parts with VVC.
[0014] In HEVC, for each inter PU, one of the three prediction modes, including inter, skip and merge, can be selected. In general, a motion vector competition (MVC for short) scheme is introduced to select one motion candidate from a given candidate set, including spatial and temporal motion candidates. Multiple reference motion estimation allows finding the best reference in two possible reconstructed reference picture lists, i.e., list 0 and list 1. For the inter mode (informally named AMVP mode, where AMVP stands for advanced motion vector prediction), the inter prediction indicator (list 0, list 1 or bi-prediction), the reference index, the motion candidate index, the motion vector difference (MVD) and the prediction residual are transmitted. For the skip and merge modes, only the merge index is transmitted, and the current PU inherits the inter prediction indicator, the reference index and the motion vector from the neighboring PU referred by the coded merge index. For the CU coded in skip, the residual signal is also omitted.
[0015] In VVC, AMVP mode is further improved by new modes such as symmetric motion vector difference (SMVD) mode, adaptive motion vector resolution (AMVR) and affine AMVP mode; merge / skip mode is further improved by enhanced merge candidates, combined inter and intra prediction (CIIP), affine merge mode, subblock temporal motion vector predictor (SbTMVP), merge mode with motion vector difference (MMVD) and geometric partition mode (GPM). In VVC, decoder-side motion vector refinement (DMVR), bi-directional optical flow (BDOF) and prediction refinement with optical flow (PROF) are used to refine the decoder-side motion vector or motion compensation predictor.
[0016] In VVC, several new coding tools are developed to further improve AMVP, merge and skip modes, such as bilateral matching AMVP-merge mode, multi-hypothesis prediction (MHP), overlapped block motion compensation (OBMC), etc. In addition, decoder-side motion vector refinement based on template matching is proposed to enhance the coding efficiency of inter prediction.
[0017] In addition to the inter coding functions in HEVC, VVC includes a variety of new and improved inter prediction coding tools listed as follows: - Extended merge prediction - Merge mode with motion vector difference (MMVD) - Symmetric motion vector difference (SMVD) signaling - Affine motion compensation prediction - Subblock-based temporal motion vector prediction (SbTMVP) - Adaptive motion vector resolution (AMVR) - Motion field storage: 1 / 16 luma sample motion vector storage and 8x8 motion field compression - Bi-prediction with CU-level Weight (BCW) - Bi-directional optical flow (BDOF) - Decoder-side motion vector refinement (DMVR) - Geometry partition mode (GPM) - Combined inter and intra prediction (CIIP) After the completion of the VVC standardization, the Enhanced Compression Model (ECM) reference software was developed to study the potential needs of future video coding technology standardization. In the current ECM, several inter prediction coding tools are included to further save BD rate: - Local Illumination Compensation (LIC) - Non-adjacent spatial candidates - Template Matching (TM) - Overlapped Block Motion Compensation (OBMC) - Multi-hypothesis prediction (MHP) - Bi-side matching AMVP-merge mode - And some other tools under development The following text provides detailed information on the partial inter prediction methods specified in VVC and ECM.
[0018] Extended Merge Prediction In VVC, the merge candidate list is constructed in order by including the following five types of candidates: - Spatial MVP from spatial neighboring CUs - Temporal MVP from co-located CUs - History-based MVP from FIFO table - Pairwise average MVP - Zero motion vector (Zero MV) The size of the merge list is signaled in the sequence parameter set header and the maximum allowed size of the merge list is 6. For each CU coded in merge mode, the index of the best merge candidate is coded using Truncated Unary Binarization (TU). The first bin of the merge index is coded using context coding and the remaining bins are coded using bypass coding.
[0019] The derivation process of each type of merge candidate is provided in this section. As in HEVC, VVC also supports parallel derivation of the merge candidate list (or called merge candidate list) for all CUs within a certain size of region.
[0020] History-based merge candidate derivation The history-based MVP (History-based MVP, HMVP) merge candidate is added to the merge list after the spatial MVP and TMVP. In this method, the motion information of previously coded blocks is stored in a table and used as the MVP for the current CU. A table containing multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (emptied) when a new CTU row is encountered. Whenever there is a non-subblock inter coded CU, the associated motion information is added as a new HMVP candidate to the last entry of the table.
[0021] The size S of the HMVP table is set to 6, which means that up to 5 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained First-In-First-Out (FIFO) rule is used, where a redundancy check is first applied to find if there is an identical HMVP in the table. If found, the identical HMVP is removed from the table and all the HMVP candidates after it are moved forward, and the identical HMVP is inserted to the last entry of the table.
[0022] The HMVP candidates can be used in the construction process of the merge candidate list. The latest few HMVP candidates in the table are checked in order and inserted to the candidate list after the TMVP candidates. Redundancy check is applied to the HMVP candidates with the spatial or temporal merge candidates.
[0023] To reduce the number of redundancy check operations, the following simplifications are introduced: - The last two entries in the table are checked for redundancy with the A1 and B1 spatial candidates, respectively.
[0024] - The HMVP-based merge candidate list construction process is terminated once the total number of available merge candidates reaches the maximum allowed merge candidate number minus 1.
[0025] Pairwise average merge candidate derivation A pairwise average candidate is generated by averaging a predefined pair of candidates in the existing merge candidate list using the first two merge candidates. The first merge candidate is defined as p0Cand and the second merge candidate is defined as p1Cand. The average motion vector is calculated according to the availability of the motion vectors p0Cand and p1Cand from each reference list, respectively. If both motion vectors in one list are available, they are averaged even if they point to different reference pictures, and their reference pictures are set to the reference picture in p0Cand; if only one motion vector is available, it is used directly; if no motion vector is available, this list is kept invalid. In addition, if the half-pel interpolation filter indices of p0Cand and p1Cand are different, they are set to 0.
[0026] When the merge list is not full after the pairwise average merge candidate is added, zero-MVPs are inserted at the end until the maximum number of merge candidates is reached.
[0027] Merge mode with MVD (MMVD) In addition to the merge mode, where the implicitly derived motion information is directly used for the prediction sample generation of the current CU, VVC also introduces a merge mode based on motion vector difference (MMVD). After the regular merge flag is sent, an MMVD flag is sent immediately to specify whether MMVD mode is used for the CU.
[0028] In MMVD, after the selection of the merge candidate, it is further refined by the sent MVD information. The further information includes a merge candidate flag, an index for specifying the motion magnitude, and an index for indicating the motion direction. In MMVD mode, one of the first two candidates in the merge list is selected as the MV base. The MMVD candidate flag is sent to specify which one between the first and second merge candidate is used.
[0029] The distance index specifies the motion magnitude information and indicates the predefined offset of the L0 reference block 410 and the L1 reference block 420 relative to the starting point (412 and 422). As shown in Figure 4 , the offset is added to the horizontal component or the vertical component of the starting motion magnitude, where different styles of small circles correspond to different offsets relative to the center. The relationship between the distance index and the predefined offset is shown in Table 1.
[0030]
[0031] The direction index indicates the MVD direction relative to the starting point. The direction index can represent the four directions shown in Table 2. It is important to note that the meaning of the MVD symbol can vary depending on the information of the starting MV. When the starting MV is an unprediction MV or a bidirectional predicted MV and both lists point to the same side of the current image (i.e., both reference POCs are greater than or less than the current image's POC), the symbol in Table 2 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional predicted MV and the two MVs point to different sides of the current image (i.e., one reference POC is greater than the current image's POC, and the other reference POC is less than the current image's POC), and the POC difference in list 0 is greater than the difference in list 1, the symbol in Table 2 specifies the sign of the MV offset added to the list0 MV component of the starting MV, while the sign of the list1 MV has the opposite value. Otherwise, if the POC difference in list 1 is greater than that in list 0, the symbol in Table 2 specifies the sign of the MV offset added to the list1 MV component of the starting MV, while the sign of the list0 MV has the opposite value.
[0032] The MVD is scaled based on the POC difference in each direction. If the POC differences are the same in both lists, no scaling is needed. Otherwise, if the POC difference in list 0 is greater than the POC difference in list 1, the MVD of list 1 is scaled as described above by defining the POC difference of L0 as td and the POC difference of L1 as tb. If the POC difference of L1 is greater than that of L0, the MVD of list 0 is scaled in the same way. If the initial MV is unidirectionally predicted, the MVD is added to the available MV.
[0033]
[0034] Affine Motion Compensation Prediction In HEVC, only a translational motion model is applied for motion compensation prediction (MCP). However, in the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is used. Figure 5A As shown in -B, the affine motion field of block 510 is composed of... Figure 5A Motion information of two control points (4 parameters) or Figure 5B The motion information description of the motion vectors (6 parameters) of the three control points in the model.
[0035] For the 4-parameter affine motion model, the motion vector at sample position (x, y) in the block is derived as follows: , (1) For the 6-parameter affine motion model, the motion vector at sample position (x, y) in the block is derived as follows: , (2) where (mv 0x , mv 0y ) is the motion vector of the top-left control point, (mv 1x , mv 1y ) is the motion vector of the top-right control point, and (mv 2x , mv 2y ) is the motion vector of the bottom-left control point.
[0036] To simplify the motion-compensation prediction, a block-based affine transform prediction is adopted. To derive the motion vector for each 4x4 luma sub-block, as shown in FIG. 6, the motion vector of each sub-block center sample is calculated according to the above equation and rounded to 1 / 16 fractional precision. Then, a motion-compensation interpolation filter is applied to generate the prediction for each sub-block using the derived motion vector. The sub-block size for chroma components is also set to 4x4. The motion vector (MV) for a 4x4 chroma sub-block is calculated as the average of the motion vectors of the top-left and bottom-right luma sub-blocks in the collocated 8x8 luma region.
[0037] Similar to translational inter prediction, there are two modes for affine inter prediction: affine merge mode and affine AMVP mode.
[0038] Affine Merge Prediction The AF M ERGE mode can be applied to a CU whose width and height are both greater than or equal to 8. In this mode, the CPMVs (control point MVs) of the current CU are generated based on the motion information of spatial neighboring CUs. There can be up to five CPMV candidates, and an index is signaled to indicate the candidate used for the current CU. The following three types of CPMV candidates are used to form the affine merge candidate list: - inherited affine merge candidate extrapolated from the CPMVs of neighboring CUs - constructed affine merge candidate CPMVP derived using the translational MVs of neighboring CUs - zero MV In VVC, there are up to two inherited affine candidates, which are derived from the affine motion model of neighboring blocks, one from the left neighboring CU and one from the top neighboring CU. Figure 7 shows the candidate blocks of the current block 710. For the left predictor, the scan order is A0->A1; for the top predictor, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No prune check is performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vectors are used to derive the CPMVP candidates in the affine merge list of the current CU. As shown in Figure 8, if the bottom-left neighboring block A of the current block 810 is coded with affine mode, the motion vectors v2, v3 and v4 of the top-left, top-right and bottom-left corners of the CU 820 containing block A are available. When block A is coded with 4-parameter affine model, two CPMVs (i.e. v_0 and v_1) of the current CU are calculated from v2 and v3; when block A is coded with 6-parameter affine model, three CPMVs of the current CU are calculated from v2, v3 and v4.
[0039] Constructing affine candidates refers to the candidates constructed by combining the neighboring translational motion information of each control point. The motion information of the control points is derived from the specified spatial and temporal neighborhoods of the current block 910, as shown in Figure 9 k (k = 1, 2, 3, 4) denotes the k-th control point. For CPMV1, the B2->B3->A2 block is checked and the motion vector of the first available block is used. For CPMV2, the B1->B0 block is checked; for CPMV3, the A1->A0 block is checked. For TMVP, if available, it is used as CPMV4.
[0040] After obtaining the motion vectors of the four control points, affine merge candidates are constructed based on the motion information. The following combinations of control point motion vectors will be used in order to construct: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4},{CPMV2, CPMV3, CPMV4}, { CPMV1, CPMV2}, { CPMV1, CPMV3} The combination of 3 CPMVs constitutes a 6-parameter affine merge candidate, and the combination of 2 CPMVs constitutes a 4-parameter affine merge candidate. To avoid the motion scaling process, if the reference indices of the control points are different, the related combination of control point motion vectors is discarded.
[0041] After checking the inherited affine merge candidates and constructed affine merge candidates, if the list is still not full, a zero MV is inserted at the end of the list.
[0042] Decoder-side motion vector refinement (DMVR) in VVC To improve the accuracy of the motion vector in merge mode, a decoder-side motion vector refinement method based on bilateral matching (BM) is applied in VVC. In the bi-prediction operation, the refined motion vector of the current block 1020 in the current picture 1010 is searched around the initial motion vectors (1032 and 1034) in the reference picture list L0 1012 and the reference picture list L1 1014. As shown in FIG. 10, the collocated blocks 1022 and 1024 in L0 and L1 are determined according to the initial motion vectors (1030 and 1032) and the position of the current block 1020 in the current picture. The BM method calculates the distortion between the two candidate blocks (1042 and 1044) in the reference picture list L0 and the list L1. The positions of the two candidate blocks (1042 and 1044) are determined by adding two opposite offsets (1062 and 1064) to the two initial motion vectors (1032 and 1034) to obtain two candidate motion vectors (1052 and 1054). As shown in FIG. 10, based on each motion vector candidate around the initial motion vector (1032 or 1034), the sum of absolute differences (SAD) between the candidate blocks (1042 and 1044) is calculated. The motion vector candidate (1052 or 1054) with the lowest SAD will be the refined motion vector and used to generate the bi-prediction signal.
[0043] In VVC, the application of DMVR is limited and only applicable to CUs coded with the following modes and features: - CU-level merge mode with bi-predictive MV - One reference picture is before the current picture and the other reference picture is after the current picture - The distance (i.e., POC difference) of the two reference pictures to the current picture is the same - Both reference pictures are short-term reference pictures - The CU has more than 64 luma samples - The height and width of the CU are both greater than or equal to 8 luma samples - The BCW weight index indicates equal weights - The current block is not enabled with WP (Weighted Prediction) - The current block is not using CIIP mode The refined MV derived by the DMVR process is used to generate inter prediction samples and for temporal motion vector prediction of future picture coding. The original MV is used for deblocking and for spatial motion vector prediction of future CU coding.
[0044] Additional features of DMVR are mentioned in the following subclauses.
[0045] Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode is supported for inter prediction. Geometric partitioning mode is signaled using a CU-level flag as one of the merge modes, including regular merge mode, MMVD mode, CIIP mode and subblock merge mode. For each possible CU size, geometric partitioning mode supports 64 partitions in total for each possible CU size where 8x64 and 64x8 are not included.
[0046] When this mode is used, a CU is partitioned into two parts by a geometrically positioned straight line as shown in FIG. 11. The position of the partitioning line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the CU using geometric partitioning is inter predicted using its own motion; each partition is allowed only unidirectional prediction, i.e. each part has one motion vector and one reference index. The unidirectional prediction motion constraint is applied to ensure the same as traditional bi-prediction, i.e. each CU only needs two motion-compensated predictions. The unidirectional prediction motion of each partition is derived.
[0047] If the current CU uses geometric partitioning mode, a geometric partitioning index is also signaled, indicating the partition mode (angle and offset) of the geometric partitioning, and two merge indices (one for each partition). The maximum size of the GPM candidate set is explicitly specified in SPS, and the syntax binarization of the GPM merge indices is specified. After the prediction of each part of the geometric partitioning, a blending process with adaptive weights is used to adjust the sample values at the geometric partitioning edge. This is the prediction signal of the whole CU, and the transform and quantization processes are applied to the whole CU as other prediction modes. Finally, the motion field of the CU predicted using geometric partitioning mode is stored.
[0048] Combined Inter and Intra Prediction (CIIP) In VVC, when a CU is coded in merge mode, an extra flag is signaled to indicate whether inter / intra combination prediction (CIIP) mode is applied to the current CU if the CU contains at least 64 luma samples (i.e., CU width times CU height is equal to or larger than 64) and both CU width and CU height are smaller than 128 luma samples. As the name suggests, CIIP prediction combines inter prediction signal with intra prediction signal. The inter prediction signal P_inter under CIIP mode is derived using the same inter prediction process as regular merge mode; the intra prediction signal P_intra is derived using planar mode following regular intra prediction process. Then, the intra and inter prediction signals are combined using a weighted average method, where the weight value wt is calculated according to the coding mode of the top and left neighboring blocks (as shown in FIG. 12) of the current CU 1210 as follows: - If top neighbor is available and intra coded, set isIntraTop to 1, otherwise set it to 0; - If left neighbor is available and intra coded, set islntraLeft to 1, otherwise set it to 0; - If (islntraLeft + isIntraTop) is equal to 2, set wt to 3; - Otherwise, if (islntraLeft + isIntraTop) is equal to 1, set wt to 2; --- Otherwise, set wt to 1.
[0049] CIIP prediction is formed as follows: , (3) Template Matching (TM) Template matching (TM) is a decoder-side motion vector (MV) derivation method that refines the motion information of a current CU by finding the closest match between a template (i.e., the top and / or left neighboring blocks of the current CU) in the current picture and a block (i.e., of the same size as the template) in the reference picture. As shown in FIG. 13, a better motion vector (MV) is searched around the initial motion of the current CU within a [-8, +8] pixel search range. In FIG. 13, the pixel row 1314 above the current block and the pixel column 1316 left of the current block 1312 in the current picture 1310 (only a partial picture area is shown) are selected as the template. The search starts from the initial position (identified by the initial motion vector 1330) in the reference picture. As shown in FIG. 13, the corresponding pixel row 1324 above the reference block 1322 and the corresponding pixel column 1326 left of the reference block 1322 in the reference picture 1320 (only a partial picture area is shown) are identified. The [-8, +8] pixel search range 1340 is represented by the dashed squares.
[0050] The template matching method in JVET-J0021 (Yi-Wen Chen et al., “Description of SDR, HDR and 360° video coding technology proposal by Qualcomm and Technicolor - low and high complexity versions,” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 10th Meeting: San Diego, US, 10-20 April 2018, Paper: JVET-J0021) is used with the following modifications: the search step size is determined based on the AMVR mode, and TM can be cascaded with the bilateral matching process in merge mode.
[0051] In AMVP mode, MVP candidates are determined based on template matching error to select the candidate with the smallest difference between the current block template and the reference block template, and then TM is only performed for that specific MVP candidate to refine the MV. TM starts from full-pel MVD precision (or 4-pixel for 4-pixel AMVR mode) and refines the MVP candidate using an iterative diamond search within a search range of [-8, +8] pixels. AMVP candidates can be further refined by a cross search that starts from full-pel MVD precision (or 4-pixel for 4-pixel AMVR mode) and then proceeds with half-pel and quarter-pel searches in turn, depending on the AMVR mode as shown in Table 3. This search process ensures that the MVP candidate remains at the same MV precision as indicated by the AMVR mode after the TM process. During the search process, the search process is terminated if the difference between the previous minimum cost in the iteration and the current minimum cost is less than or equal to the area of the block.
[0052]
[0053] In merge mode, a similar search method is applied to the merge candidate indicated by the merge index. TM can be performed to 1 / 8-pel MVD precision or skip beyond half-pel MVD precision depending on whether an alternative interpolation filter is used according to the merge motion information (when AMVR is in half-pel mode), as shown in Table 3. In addition, when TM mode is enabled, template matching can be applied as an additional MV refinement process between the independent process or the block-level and subblock-level bilateral matching (BM) method, depending on whether BM is enabled according to its enabling condition check.
[0054] Overlapped Block Motion Compensation (OBMC) When OBMC is applied, the top and left boundary pixels of a CU are predicted with weighted refinement using motion information of neighboring blocks, as described in JVET-L0101 (Zhi-Yi Lin et al., “CE10.2.1: OBMC,” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 12th Meeting: Macao, CN, 3-12 Oct, 2018, Doc: JVET-L0101).
[0055] The condition for not applying OBMC is as follows: - When OBMC is disabled at the SPS level - When the current block has Intra mode or IBC mode - When the current block applies LIC - When the current luma block region is smaller than or equal to 32 Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom and right sub-block boundary pixels using the motion information of the neighboring sub-blocks. It is applied to sub-block based coding tools: - Affine AMVP mode; - Affine merge mode and sub-block based temporal motion vector prediction (SbTMVP); - Sub-block based bi-prediction.
[0056] When the OBMC mode is combined with LMCS for CIIP mode, the inter blending is performed before the LMCS mapping of the inter samples. The LMCS is applied to the blended inter samples that are combined with the intra samples for which LMCS is applied in CIIP mode, , , where denotes the sample in the original domain predicted by the motion of the current block, denotes the predicted sample in the mapped domain, denotes the sample in the original domain predicted by the motion of the neighboring block, and are the weights.
[0057] Template matching based OBMC In the template matching based OBMC scheme, the predictor of the CU boundary samples is not directly derived using weighted prediction, but the derivation method depends on the template matching cost, including using only the motion information of the current block, or using the motion information of the neighboring blocks and a hybrid mode.
[0058] In this scheme, for each block of size 4x4 at the top CU boundary, the above template size is equal to 4x1. If N neighboring blocks have the same motion information, the above template size is enlarged to 4Nx1 because the MC (Motion Compensation) operation can be processed at once. For each left block of size 4x4 at the left CU boundary, the left template size is equal to 1x4 or 1x4N as shown in Figure 14 In Figure 14 , block 1410 corresponds to one CU. If N neighboring blocks have the same motion information, the above template size is enlarged to 4Nx1 because the MC operation can be processed at once, in the same way as in ECM-OBMC. For each left block of size 4x4 at the left CU boundary, the left template size is equal to 1x4 or 1x4N.
[0059] For each 4x4 top block (or N 4x4 block group), the predictor of the boundary samples is derived according to the following steps: - Take block A as the current block and its above neighboring block AboveNeighbour_A as an example. The operation of the left neighboring block is done in the same way.
[0060] - First, three template matching costs (Cost1, Cost2 and Cost3) are measured by SAD (Sum of Absolute Difference) between the reconstructed samples of the template and the corresponding reference samples derived by the MC process according to the following three types of motion information: 、 、 ) : 1. Cost1 is calculated according to the motion information of A.
[0061] 2. Cost2 is calculated according to the motion information of AboveNeighbour_A.
[0062] 3. Cost3 is calculated according to the weighted prediction of the motion information of A and AboveNeighbour_A, with the weighting factors being 3 / 4 and , respectively.
[0063] - Second, the final prediction result of the boundary samples is calculated by selecting one of the three methods by comparing Cost1, Cost2 and Cost3.
[0064] The original MC result using the motion information of the current block is denoted as , the MC result using the motion information of the neighboring block is denoted as . The final prediction result is denoted as .
[0065] - If Cost1 is the minimum value, then .
[0066] - If (Cost2 + (Cost2 » 2) + (Cost2 » 3)) <= Cost1, then use blending mode 1.
[0067] - For luma blocks, the number of blended pixel rows is 4.
[0068] , , , , For chroma blocks, the number of blended pixel rows is 1.
[0069] , - If Cost1<= Cost2, use hybrid mode 2.
[0070] For luma blocks, the number of mixed pixel rows is 2.
[0071] , , For chroma blocks, the number of mixed pixel rows / columns is 1.
[0072] , - Otherwise, use hybrid mode 3.
[0073] For luma blocks, the number of mixed pixel rows is 4.
[0074] , , , For chroma blocks, the number of mixed pixel rows is 1.
[0075] .
[0076] Geometric partitioning mode (GPM) with template matching (TM) Template matching is applied for GPM. When GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether TM is applied for the two geometric partitions. The motion information of each geometric partition is refined using TM. After TM is selected, a template is constructed using the left, top or top-left neighboring samples according to the partition angle, as shown in Table 4. Then, the motion is refined by minimizing the difference between the current template and the template in the reference picture using the same search mode as in merge mode (disable half-pel interpolation filter).
[0077]
[0078] The GPM candidate list is constructed as follows: 1. Interleaved List-0 MV candidates and List-1 MV candidates are directly obtained from the regular merge candidate list, with List-0 MV candidates having higher priority than List-1 MV candidates. An adaptive threshold pruning method based on the current CU size is adopted to remove redundant MV candidates.
[0079] 2. Interleaved List-1 MV candidates and List-0 MV candidates are further derived directly from the regular merge candidate list, with List-1 MV candidates having higher priority than List-0 MV candidates. Again, an adaptive threshold pruning method based on the current CU size is employed to remove redundant MV candidates.
[0080] 3. Zero-padding MV candidates until the GPM candidate list is full.
[0081] GPM-MMVD and GPM-TM are enabled for one GPM CU only. This is achieved by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are equal to false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied for both GPM partitions. Otherwise (at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is inferred to be false.
[0082] GPM partition mode reordering based on template matching In GPM partition mode reordering based on template matching, given the motion information of the current GPM block, the TM cost value of each GPM partition mode is calculated. Then, all GPM partition modes are reordered in ascending order according to the TM cost value. Instead of signaling the GPM partition mode, an index of Golomb-Rice code is used to indicate the exact position of the GPM partition mode in the reordered list.
[0083] The reordering method of GPM partition modes is a two-step process, which is performed after the corresponding reference templates of the two GPM partitions in the generated coding unit are generated, as follows: • The GPM partition edges are extended into the reference templates of the two GPM partitions to obtain 64 reference templates, and the TM cost of each reference template is calculated; • The GPM partition modes are reordered in ascending order according to the TM cost value, and the best 32 partition modes are marked as available partition modes.
[0084] As shown in FIG. 15, the edges on the template are extended from the edges of the current CU, but GPM blending processing is not used in the template area that crosses the edges. In FIG. 15, block 1510 corresponds to the current block, block 1520 corresponds to the top template, and block 1530 corresponds to the left template.
[0085] After reordering the TM costs in ascending order, a signaling index is used.
[0086] Current picture reference Motion compensation is one of the key techniques in hybrid video coding, which explores the pixel correlation between neighboring pictures. It is generally assumed that a pattern corresponding to an object or background in a frame will be displaced to form a corresponding object in a subsequent frame, or be associated with other patterns within the current frame. By estimating such displacement (e.g., using block matching techniques), the pattern can be substantially reproduced without re-encoding the pattern. Similarly, block matching and copying are also attempted to allow selection of a reference block from the same picture. It is observed that this concept is not efficient when applied to camera-captured videos. Part of the reason is that text patterns in spatial neighboring regions can be similar to the current coding block, but usually have some gradual transition in space. Therefore, it is difficult to find a perfect match for a block in the same picture of a camera-captured video, thus limiting the improvement of coding performance.
[0087] However, for screen content, the spatial correlation between pixels within the same picture is different. For typical videos containing text and graphics, there are usually repeated patterns within the same picture. Therefore, it has been observed that intra-picture (picture) block compensation is very effective. Thus, a new prediction mode, the intra block copy (IBC) mode or called current picture reference (CPR), is introduced for screen content coding to exploit this property. Under the CPR mode, a prediction unit (PU) is predicted from a previously reconstructed block within the same picture. In addition, a displacement vector (called block vector or BV) is used to represent the relative displacement from the current block position to the reference block position. The prediction error is then coded using transform, quantization, and entropy coding. Fig. 16 shows an example of CPR compensation, where block 1612 is the corresponding block of block 1610, and block 1622 is the corresponding block of block 1620. In this technique, the reference samples correspond to the reconstructed samples of the current decoded picture before the in-loop filtering operations (deblocking filter and sample adaptive offset (SAO) filter in HEVC).
[0088] When using CPR, only a part of the current picture can be used as the reference picture. In order to regulate the valid motion vector (MV) values that reference the current picture, some bitstream conformance constraints need to be imposed.
[0089] First, one of the following two conditions must be true: BV_x + offsetX + nPbSw + xPbs - xCbs <= 0, (4) BV_y + offsetY + nPbSh + yPbs - yCbs <= 0, (5) Second, the following WPP condition must be true: ( xPbs + BV_x + offsetX + nPbSw - 1 ) / CtbSizeY - xCbs / CtbSizeY <= yCbs / CtbSizeY - ( yPbs + BV_y + offsetY + nPbSh - 1 ) / CtbSizeY, (6) In equations (4) to (6), (BV_x, BV_y) is the luma block vector (motion vector of the CPR) of the current PU; nPbSw and nPbSh are the width and height of the current PU; (xPbs, yPbs) is the position of the current PU relative to the top-left pixel of the current picture; (xCbs, yCbs) is the position of the current CU relative to the top-left pixel of the current picture; CtbSizeY is the size of the CTU. offsetX and offsetY are two adjustment offsets considering the chroma sample interpolation of the CPR mode.
[0090] offsetX = BVC_x & 0x7? 2 : 0, (7) offsetY = BVC_y & 0x7? 2 : 0, (8) (BVC_x, BVC_y) is the chroma block vector with 1 / 8 pixel resolution in HEVC.
[0091] Third, the reference block of the CPR must be within the same tile / slice boundary.
[0092] IBC with template matchingBoth IBC merge mode and IBC AMVP mode use template matching. Compared to regular IBC merge mode, the IBC-TM merge list is modified to select the candidates according to the motion distance between the candidates, just like in regular TM merge mode. The zero motion at the end is replaced with the motion vectors of the left (-W, 0), top (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is its height.
[0093]
[0094] In IBC-TM merge mode, the selected candidate object is refined using a template matching method before RDO or the decoding process. IBC-TM merge mode has been competing with regular IBC merge mode, and a TM merge flag is signaled.
[0095] In IBC-TM AMVP mode, up to 3 candidate objects can be selected from the IBC-TM merge list. The 3 selected candidate objects are refined using a template matching method and sorted according to their template matching cost. Then, only the top 2 candidate objects are considered in the motion estimation process as usual.
[0096] The template matching refinement for IBC-TM merge and AMVP modes is very simple because IBC motion vectors are restricted to (i) integer and (ii) within the reference region, as shown in FIG. 17, which shows the IBC reference region depending on the current CU position in various cases (1710-1740). Therefore, in IBC-TM merge mode, all refinements are performed with integer precision, while in IBC-TM AMVP mode, they are performed with integer or 4-pixel precision, depending on the AMVR value. This refinement only accesses samples without interpolation. In both cases, the refined motion vector and the template used in each refinement step must follow the constraints of the reference region.
[0097] Multi-Pass Decoder-Side Motion Vector Refinement (MP-DMVR) A multi-pass decoder-side motion vector refinement method is employed. In the first pass, bi-directional matching (BM) is applied to the coded block. In the second pass, BM is applied to each 16x16 sub-block within the coded block. In the third pass, the motion vectors (MVs) in each 8x8 sub-block are refined by applying bi-directional optical flow (BDOF). The refined motion vectors (MVs) will be stored for spatial and temporal motion vector prediction.
[0098] First Pass - Block-Based Bi-Directional Matching Motion Vector Refinement In the first stage, a refined motion vector (MV) is obtained by applying BM to the coding block. Similar to the decoder-side motion vector refinement (DMVR), in the bi-prediction operation, a refined motion vector (MV) is searched around two initial motion vectors (MV0 and MV1) in the reference picture lists L0 and L1. Based on the minimum bi-lateral matching cost between two reference blocks in L0 and L1, the refined motion vectors (MV0_pass1 and MV1_pass1) are derived around the initial motion vectors (MV).
[0099] The BM performs a local search to obtain the integer sample precision intDeltaMV. The local search adopts a 3x3 square search pattern, with a search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum value of sHor and sVer is 8.
[0100] The bi-lateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the MRSAD cost function is applied to eliminate the DC effect of the inter-reference block distortion. The intDeltaMV local search terminates when the bilCost of the 3x3 search pattern center point has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3x3 search pattern, and the search for the minimum cost continues until the end of the search range is reached.
[0101] Further, the existing fractional sample refinement is applied to derive the final deltaMV. The refined MV after the first stage is as follows: MV0_pass1 = MV0 + deltaMV, MV1_pass1 = MV1 - deltaMV, Second stage - sub-block based bi-lateral matching MV refinement In the second stage, a refined MV is derived by applying BM to the 16x16 grid sub-blocks. For each sub-block, a refined motion vector (MV) is searched around two motion vectors (MV) obtained in the first stage (e.g., MV0_pass1 and MV1_pass1) in the reference picture lists L0 and L1. The refined motion vectors (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bi-lateral matching cost between two reference sub-blocks in L0 and L1.
[0102] For each sub-block, BM performs a full search to derive the integer sample precision intDeltaMV. The horizontal direction search range of the full search is [-sHor, sHor] and the vertical direction search range is [-sVer, sVer], where the values of sHor and sVer are determined by the block size, and the maximum value of sHor and sVer is 8.
[0103] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between the two reference sub-blocks, with the formula: bilCost = satdCost * costFactor. The search area (2*sHor + 1) * (2*sVer + 1) is divided into 5 diamond-shaped search areas, as shown in Figure 18 where the 5 search areas are shown in 5 different shades. Each search area is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond-shaped area is processed in order starting from the center of the search area. In each area, the search points are processed in raster scan order, starting from the top-left corner of the area until the bottom-right corner. When the minimum bilCost within the current search area is less than or equal to the threshold of sbW * sbH, the integer-pel full search terminates; otherwise, the integer-pel full search continues to the next search area until all search points are checked. In addition, if the difference between the last minimum cost and the current minimum cost in the iteration is less than or equal to the threshold of the block area, the search process terminates.
[0104] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV(sbIdx2). The refined MV in the second round is then derived as: MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2), MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2), Third stage - sub-block based bi-directional optical flow MV refinement In the third stage, the refined motion vector (MV) is derived by applying BDOF to the 8x8 grid sub-blocks. For each 8x8 sub-block, BDOF refinement is applied, starting from the refined motion vector of the second pass parent-child block, to derive the scaled Vx and Vy without clipping. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped to between -32 and 32.
[0105] The refined motion vectors of the third stage (e.g., MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as follows: MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv, MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) - bioMv, Bilateral matching AMVP-merge mode The bi-directional predictor consists of an AMVP predictor in one direction and a merge predictor in the other direction. This mode can be enabled for a coded block when the selected merge predictor and AMVP predictor satisfy the DMVR condition. The DMVR condition means that there are at least one reference picture from the past and one reference picture from the future with respect to the current picture and the distance of the two reference pictures to the current picture is the same, then bilateral matching MV refinement is applied to the merge MV candidate and AMVP MVP as a starting point. Otherwise, if the template matching function is enabled, template matching MV refinement is applied to the merge predictor or the AMVP predictor with higher template matching cost.
[0106] The AMVP part of this mode is signaled in the form of regular uni-directional AMVP, i.e., reference index and MVD are sent, and if template matching is used, the derived MVP index is sent; if template matching is disabled, the MVP index is sent.
[0107] For the AMVP direction LX, X can be 0 or 1, the merge part of the other direction (1 - LX) is derived implicitly by minimizing the bilateral matching cost between the AMVP predictor and the merge predictor (i.e., for a pair of AMVP and merge motion vectors). For each merge candidate in the merge candidate list with a motion vector in the other direction (1 - LX), the bilateral matching cost is calculated using the merge candidate motion vector and the AMVP motion vector. The merge candidate with the minimum cost is selected. The bilateral matching refinement is applied to the coded block using the selected merge candidate motion vector and the AMVP motion vector as a starting point.
[0108] The third stage of multi-stage DMVR (i.e., 8x8 sub-PU BDOF refinement of multi-stage DMVR) is enabled for the AMVP-merge mode coded block.
[0109] The mode is indicated by a flag, and if the mode is enabled, the AMVP direction LX is also indicated by a flag.
[0110] When the current block uses bi-match (BM) AMVP merge mode and template matching is enabled, the MVD signal is not signaled. A pair of additional AMVP merge MVPs is introduced. The merge candidate list is sorted in ascending order based on the BM cost. An index (0 or 1) is signaled to indicate which merge candidate is used in the sorted merge candidate list. When there is only one candidate in the merge candidate list, the pair of AMVP MVP and merge MVP without bi-match MV refinement is padded.
[0111] MVD sign prediction in ECM In this method, the possible MVD sign combinations are sorted according to the template matching cost, and the index corresponding to the true MVD sign is derived and context coded. At the decoder side, the derivation process of MVD sign is as follows: 1. Parse the magnitude of the MVD component 2. Parse the context coded MVD sign prediction index 3. Construct MV candidates by combining possible signs with the absolute MVD value and add them to the MV predictor 4. Derive the MVD sign prediction cost for each derived MV according to the template matching cost and sort them 5. Use the MVD sign prediction index to pick the true MVD sign MVD sign prediction is applied to inter AMVP, affine AMVP, MMVD and affine MMVD modes. Note that when wrap-around motion compensation is enabled, the candidate motion vectors should be clipped considering the wrap-around offset.
[0112] As mentioned above, the decoder-side motion vector derivation based on template matching has been used in various video coding tools. This disclosure discloses a prioritized initial motion vector to reduce the template matching cost related to the initial motion vector, so as to prioritize the use of the initial motion vector in the motion vector refinement process. SUMMARY A method and apparatus for video coding are disclosed. According to the method, input data associated with a current block is received, wherein the input data comprises pixel data of the current block to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side. An initial MV (motion vector) of the current block is determined. One or more MV candidates are also determined. An initial TM cost associated with the initial MV is derived, wherein the initial TM cost comprises a first metric associated with a current template of the current block and a first template of a first reference block pointed by the initial MV. A modified initial TM cost is derived by adding an offset value to the initial TM cost. One or more candidate TM costs are determined for the one or more MV candidates, wherein the one or more candidate TM costs respectively comprise one or more second metrics associated with the current template of the current block and one or more second templates, and wherein each of the one or more second templates is associated with a second reference block pointed by one of the one or more MV candidates. The initial MV and the one or more MV candidates are reordered according to the modified initial TM cost and the one or more candidate TM costs.
[0114] In one embodiment, the offset value corresponds to a positive value or a negative value. In one embodiment, the offset value corresponds to a fixed value or the offset value depends on the initial TM cost. In one embodiment, the offset value is signaled or parsed from a bitstream. In one embodiment, the offset value is determined according to coding information. For example, the coding information comprises a template region, a QP (quantization parameter), or both, and wherein the template region is associated with the current template, the first template, the one or more second templates, or a combination thereof.
[0115] In one embodiment, the one or more MV candidates are associated with inter directions of motion compensation, and the inter directions of motion compensation comprise uni-directional L0 prediction, uni-directional L1 prediction, bi-directional prediction, more than two hypotheses of prediction, or a combination thereof, and if one MV candidate has a same inter direction as the initial MV, a respective TM cost of the one MV candidate is reduced by a second offset value, and the second offset value corresponds to a positive value.
[0116] In one embodiment, the one or more MV candidates are associated with weights of BCW (bi-directional prediction with CU-level weights), and if one MV candidate has a same BCW weight as the initial MV, a respective TM cost of the one MV candidate is reduced by a second offset value, and the second offset value corresponds to a positive value.
[0117] In one embodiment, each of the one or more candidate TM costs further comprises a weighted candidate MV difference between the initial MV and one of the one or more MV candidates. In one embodiment, a weighting factor of the weighted candidate MV difference corresponds to a fixed value or depends on coding information. In another embodiment, the weighting factor of the weighted candidate MV difference is predefined, or explicitly signaled or parsed from a bitstream.
[0118] In one embodiment, the first metric and the one or more second metrics correspond to SAD (sum of absolute difference), SSD (sum of squared difference), or SATD (sum of absolute transformed difference).
[0119] According to another method, input data related to a current block is received, wherein the input data comprises pixel data of the current block to be encoded at an encoder side or data related to the current block to be decoded at a decoder side. An initial MV (motion vector) of the current block is determined. One or more MV candidates are also determined. An initial TM cost related to the initial MV is derived based on a current template of the current block and a first template of a first reference block pointed by the initial MV. One or more candidate TM costs of the one or more MV candidates are respectively determined based on the current template of the current block and one or more second templates, wherein the one or more candidate TM costs comprise one or more second metrics, and wherein each of the one or more second templates is associated with a second reference block pointed by one of the one or more MV candidates. One or more weighted candidate TM costs are derived, wherein each of the one or more weighted candidate TM costs is derived by weighting one of the one or more candidate TM costs with a weighting factor, and the weighting factor is larger for a target MV candidate that is further away from the initial MV, or one or more modified candidate TM costs are derived by reducing one or more target candidate TM costs with a smaller step size by an offset value. The initial MV and the one or more MV candidates are reordered according to the initial TM cost and the one or more weighted candidate TM costs, or according to the initial TM cost and the one or more modified candidate TM costs.
[0120] In one embodiment, the one or more MV candidates comprise inter and affine MMVD (merge mode and motion vector difference) candidates with respective step sizes and directions. BRIEF DESCRIPTION OF DRAWINGS FIG. 1A illustrates an exemplary adaptive inter / intra video coding system including a loop process.
[0122] FIG. IB illustrates a corresponding decoder of the encoder in FIG. 1A.
[0123] FIG. 2 shows four types of splitting in a multi-type tree structure, including vertical binary tree splitting (SPLIT_BT_VER), horizontal binary tree splitting (SPLIT_BT_HOR), vertical ternary tree splitting (SPLIT_TT_VER), and horizontal ternary tree splitting (SPLIT_TT_HOR).
[0124] FIG. 3 shows an example of splitting a CTU into CUs using quad-tree and nested multi-type tree coding block structure, where bold block edges represent quad-tree splitting and the rest represent multi-type tree splitting.
[0125] FIG. 4 shows an example of MVD (MMVD) merge mode with a predefined offset for the starting point of L0 and L1 reference blocks.
[0126] FIGS. 5A-B show examples of affine motion field for a current block, where the motion is described by motion information of two control points (4 parameters) in FIG. 5A or motion information of three control point motion vectors (6 parameters) in FIG. 5B.
[0127] FIG. 6 shows an example of block-based affine transform prediction, where the motion vector of each 4x4 luma sub-block is derived according to an affine motion model.
[0128] FIG. 7 shows a spatial neighboring candidate block used to derive an inherited affine candidate block.
[0129] FIG. 8 shows an example of deriving control point MVPs for an affine coded current block, where the neighboring CUs on the left are also coded in affine mode.
[0130] FIG. 9 shows an example of constructing an affine candidate by combining the neighboring translational motion information of each control point.
[0131] FIG. 10 shows an example of decoder-side motion vector refinement (DMVR) in VVC.
[0132] FIG. 11 shows a partitioning mode according to
[0061] geometry partitioning mode (GPM), where a CU is split into two parts by a geometrically positioned straight line.
[0133] FIG. 12 shows top and left neighboring blocks used to calculate inter and intra combination prediction (CIIP) weights.
[0134] FIG. 13 shows a template matching (TM) example used to refine the current CU motion information in a decoder-side motion vector (MV) derivation method.
[0135] FIG. 14 shows an example of OBMC based on template matching, where for each left block of size 4x4 at the left CU boundary, the left template size is equal to 1x4 or 1x4N. Similar process also applies to each top block of size 4x4 at the top CU boundary.
[0136] FIG. 15 shows an example of GPM partition mode reordering based on template matching, where the edges on the template are extended from the edges of the current CU.
[0137] FIG. 16 shows an example of CPR (current reference picture) compensation.
[0138] FIG. 17 shows an example of IBC motion vector constraint using template matching.
[0139] FIG. 18 shows 5 diamond search regions used in the second pass bilateral matching motion vector refinement in the multi-pass decoder-side motion vector refinement (MP-DMVR) process.
[0140] FIG. 19 shows an example of MVD candidate reordering by motion vector matching according to an embodiment of the present application.
[0141] FIG. 20 shows a flowchart of an example video coding system using a preferred initial motion vector according to an embodiment of the present application.
[0142] FIG. 21 shows a flowchart of another example video coding system using a preferred initial motion vector according to an embodiment of the present application.
DETAILED DESCRIPTION
[0144] Moreover, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the application can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures or operations are not shown or described in detail in order to avoid obscuring aspects of the application. Reference will now be made to the drawings wherein like numerals refer to like components throughout the specification. The following description merely is exemplary and is not intended to limit the application claimed herein to certain selected embodiments.
[0145] The concept of template matching (TM) is to refine the initial motion vector (MV) by performing template matching at the decoder, where the template is a predefined region formed by reconstructed luma and / or chroma samples as shown in FIG. 13. Based on the initial motion vector (MV), the template points to a corresponding reference block in the reference picture.
[0146] Various matching metrics (or called similarity metrics or matching costs) can be used to evaluate the similarity between the template and the reference block. The matching metrics can include, but are not limited to, sum of absolute difference (SAD), sum of squared error (SSE), and sum of transformed difference (SATD). Starting from the initial motion vector (MV), the template matching is performed to search for the best matching reference block in the reference picture within a predefined search range. After visiting all the predefined search points and computing the corresponding matching metrics, the motion vector between the template and the best matching reference block (e.g., the reference block with the minimum template matching cost) is considered as the refined motion vector.
[0147] The TM process is used to refine the motion vector (MV) based on a given initial motion vector at the decoder side. Since the motion vector refinement is done at the decoder side, there is no guarantee that the refined motion vector derived from the TM process is always better than the initial MV in terms of the accuracy of motion compensation (MC) prediction. Moreover, since the initial MV is derived from the previous block, it is somewhat reliable. This disclosure proposes several methods to make the TM process more inclined to the initial MV.
[0148] This disclosure proposes several methods to improve TM. It is noted that the term "block" in this disclosure can be a coding unit (CU), a coding block (CB), a prediction unit (PU), a prediction block (PB), a transform unit (TU), a transform block (TB), or a block of any size.
[0149] In a first embodiment of the disclosure, the cost of the initial motion vector (MV) in the TM motion refinement process is reduced by an offset value. The offset value can be a fixed value or a certain ratio of the TM cost of the initial motion vector (MV) (e.g., 1 / 16 of the TM cost of the initial motion vector (MV)). The offset value can be determined by at least one syntax element signaled in a video level (e.g., video parameter set), a sequence level (e.g., sequence parameter set), a picture level (e.g., picture parameter set, picture header), a slice level (e.g., slice header), a coding tree unit (CTU), and / or a block level (e.g., prediction unit, transform unit).
[0150] In a second embodiment of the disclosure, the cost of the initial MV in the TM motion refinement process is reduced by an offset value. The offset value can be a fixed value or a certain ratio of the TM cost of the initial MV (e.g., 1 / 16 of the TM cost of the initial MV). The offset value can be determined according to the coding information.
[0151] In one example, the offset value is determined according to the template area, as follows.
[0152] If the template area is less than TH_1 the offset value is set to 1 / N_1 of the TM cost of the initial MV Otherwise, if the template area is less than TH_2 the offset value is set to 1 / N_2 of the TM cost of the initial MV Otherwise, if the template area is less than TH_3 the offset value is set to 1 / N_3 of the TM cost of the initial MV … … … Otherwise, if the template area is less than TH_M, a signal is sent the offset value is set to 1 / N_M of the TM cost of the initial MV Otherwise the offset value is set to 1 / N_M+1 of the TM cost of the initial MV Endif In another example, the offset value is determined according to the QP used for the current block, as follows.
[0153] If the QP of the current block is less than TH_1 the offset value is set to 1 / N_1 of the TM cost of the initial MV Otherwise, if the QP of the current block is less than TH_2 the offset value is set to 1 / N_2 of the TM cost of the initial MV Otherwise if QP of current block is smaller than TH_3 Offset value is set to 1 / N_3 of TM cost of initial MV … … … Otherwise if QP of current block is smaller than TH_M Offset value is set to 1 / N_M of TM cost of initial MV Otherwise Offset value is set to 1 / N_M+1 of TM cost of initial MV Endif The offset value (e.g., N_1, N_2, …, N_M in the above embodiments) can be predefined, or can be determined by at least one syntax element signaled in video level (e.g., video parameter set), sequence level (e.g., sequence parameter set), picture level (e.g., picture parameter set, picture header), slice level (e.g., slice header), coding tree unit (CTU), and / or block level (e.g., prediction unit, transform unit).
[0154] It is worth noting that the refinement cost can be modified in a similar way, but increasing the refinement cost instead of decreasing the unrefined / initial cost. Mathematically, this is equivalent to decreasing the initial cost, but both ways of writing can be applied.
[0155] In a third embodiment of the present disclosure, the matching cost of TM motion refinement process is modified as the similarity measure cost (e.g., SAD, SSD (sum of squared difference), SATD (sum of absolute transformed difference)) plus a weighted MV difference, where the MV difference is calculated as the difference between the initial MV and the MV of each search point in the TM refinement process. The difference between two MVs is calculated as the absolute difference between the x components (or horizontal components) of the two motion vectors plus the absolute difference between the y components (or vertical components) of the two motion vectors. The weighting factor can be a fixed value, or an adaptive factor based on the coding information (e.g., quantization parameter QP). In addition, the weighting factor can be predetermined, or explicitly indicated in the bitstream.
[0156] In another embodiment, the refined weighting / cost factors can be position dependent. In one embodiment, refinements that are further away from the initial MV have higher cost weight (or lower threshold used when comparing costs and deciding whether to add it to the list), similar to the second stage of MP-DMVR. In this way, the cost of TM-based MV refinements is scaled by how far the refinement is from the initial unrefined MV. In this way, the initial MV is not explicitly prioritized. However, refinements that are closer to the initial MV have lower cost factors and thus are prioritized over refinements that are further away. This approach can be used in combination with one of the approaches to reduce the cost of the initial unrefined MV.
[0157] In another embodiment, when using TM cost to determine the motion compensated inter direction in uni-L0 prediction, uni-L1 prediction, bi-prediction, and / or more than two hypothesis prediction, the cost of a motion vector candidate that has the same inter direction as the initial MV candidate is reduced by an offset value (e.g., 1 / 16 of the TM cost of the initial MV candidate). Note that the inter direction is also referred to as inter prediction index. Specifically, when the initial MV candidate is uni-L0 prediction, the TM cost of the uni-L0 prediction MV candidate is reduced by an offset value. Similarly, when the initial MV candidate is bi-prediction, the TM cost of the bi-prediction MV candidate is reduced by an offset value.
[0158] In one example, when the motion vector candidate is bi-prediction, the TM cost of the bi-prediction motion vector candidate can be calculated by performing bi-directional motion compensation on the template using the L0 and L1 motion vectors. Then, the TM cost is calculated based on the bi-predictor and the template. The TM cost of the uni-L0 / L1 motion vector candidate is calculated based on the uni-L0 / L1 motion vector candidate and the template. After both the bi-directional TM cost and the uni-L0 / L1 TM cost are available, the inter direction used for motion compensation can be determined based on the TM cost of uni-L0 prediction, uni-L1 prediction, and bi-prediction. Specifically, the current block selects the MV candidate that has the minimum TM cost of the inter directions. Assuming the initial MV candidate is bi-directional MV, the TM cost of the bi-prediction MV candidate is reduced by an offset value and then compared with other MV candidates.
[0159] In another embodiment, when using TM cost to determine BCW weight, the cost of MV candidates with the same BCW as the initial MV candidate is reduced by an offset value (e.g., 1 / 16 of the TM cost of the initial MV candidate). Note that BCW supports various weights to average L0 and L1 motion compensation predictors. In particular, when the initial MV candidate uses equal weight BCW, the TM cost of MV candidates using equal weight BCW is reduced by an offset value.
[0160] In one example, when the MV candidate is bi-predicted, the BCW TM cost of the MV candidate can be computed by performing bi-directional motion compensation using L0 and L1 MV pairs templates with different BCW weights. Assuming the initial BCW weights are inherited from neighboring BCW weights, the TM cost of the motion vector candidate with the inherited BCW weights is reduced by an offset value and then compared with other MV candidates. In another example, the initial BCW weights can be equal BCW weights.
[0161] In another embodiment, when using TM cost to determine the motion compensation inter direction between uni-predicted L0, uni-predicted L1, bi-predicted and / or more than two hypothesis predictions, the initial MV candidate is determined by the inherited BCW weights. In particular, when the inherited BCW weights are equal weights, the initial MV candidate is determined to be bi-predicted. Otherwise, when the inherited BCW weights are unequal weights, the initial MV candidate is determined to be uni-predicted with the higher BCW weight. That is, if the BCW weight of L0 is higher, the initial MV candidate is L0 uni-predicted. The TM cost of the initial MV candidate is reduced by an offset value (e.g., 1 / 16 of the TM cost of the initial MV candidate).
[0162] In another embodiment, when using TM cost to determine the signaling order of inter and affine MMVD candidate objects, all inter and affine MMVD candidate objects with corresponding step size and direction are reordered according to TM cost. In one example, the TM cost of candidate objects with smaller step size is reduced by an offset value (e.g., 1 / 16 of the TM cost of the initial MV candidate object). In another example, the reference MV of inter or affine MMVD is refined by TM to derive a TM-refined reference MV. The TM cost of candidate objects with the same direction as the MVD direction from the original reference MV to the TM-refined reference MV is reduced by an offset value (e.g., 1 / 16 of the TM cost of the initial MV candidate object).
[0163] As mentioned before, TM is not only used for optimizing the motion vector predictor and refinement merge candidate of inter mode, but also for many other coding tools to improve their prediction efficiency, such as OBMC, IBC, etc. In the third embodiment of the present application, the method proposed in the present application can be selectively applied to the coding tools that apply TM.
[0164] Any of the above methods can be applied independently or jointly. In addition, any of the above methods can be implemented in an encoder and / or a decoder. For example, any of the above methods can be implemented in an inter prediction module of an encoder and / or a decoder. Alternatively, any of the above methods can be implemented as a circuit coupled to an inter prediction module of an encoder and / or a decoder.
[0165] Reordering motion vector difference candidates by motion vector matching In ECM, when encoding the motion vector difference (MVD), the sign bits are context coded by reordering the MVD candidates based on template matching (TM). In JVET-AC0239(), the suffix bins of the exponential Golomb code of the MVD amplitudes can also be context coded by TM-based reordering. However, in these methods, a large reference frame region needs to be fetched at the decoder side to perform TM, which can be an issue in implementation. Some methods are proposed to solve this problem.
[0166] In one embodiment, the motion vector matching cost is used in the MVD sign / suffix prediction algorithm to replace the TM cost, where the motion vector matching cost is derived by calculating the difference between the candidate motion vector (candMV) and the reference motion vector (refMV). Specifically, candMV is the motion vector predictor plus a possible MVD candidate for reordering.
[0167] Example 1. refMV is the inverse scaled motion vector of the reference block pointed by candMV.
[0168] FIG. 19 shows an illustration of this example. Given an MVD candidate MVD0 1960 to be reordered, it corresponds to a candidate motion vector candMV0 1950. Assume candMV0 1950 points to reference block B0. Consider the motion vector of B0 and scale it to the current frame to obtain the "scaled motion vector of B0 1944" in reference picture 1920. The motion vector match cost of MVD0 is derived by calculating the distance (d0 in the figure) between the current location 1914 of current block 1912 and the location 1916 pointed by the "scaled motion vector of B0 1944". Likewise, the motion vector match cost of MVD1 is derived by calculating the distance (d1 in the figure) between the current location 1914 of current block 1912 and the location 1918 pointed by the "scaled motion vector of B1 1946". When reordering is performed, if the match cost of MVD0 1960 is less than the match cost of MVD1 1962 (as shown, d0 < d1), the ordering of MVD0 1960 will take precedence over MVD1 1962. FIG. 19 shows the reference block 1932 of B0 and the reference block 1934 of B1 in reference picture 1930 of reference picture 1920. The positioning of reference block 1932 of B0 is based on the location of B0 and the MV of B0 1940, while the positioning of reference block 1934 of B1 is based on the location of B1 and the MV of B1 1942. MVP0 1954 corresponds to MV prediction 0.
[0169] Note that this example has another aspect. The match cost of MVD0 can be derived by calculating the difference between candMV0 and the inverse of the "scaled motion vector of B0", resulting in the same distance d0.
[0170] Reference blocks B0 and B1 can cover multiple coding units, resulting in multiple available motion vectors for "MV of B0" and "MV of B1". In this case, some criteria can be used to select the motion vector. For example, always use the motion vector of the top-left coding unit. For another example, use the motion vector of the coding unit that the reference block covers the most. After the motion vector is selected, the top-left location should be moved accordingly when calculating the match cost.
[0171] Example 2. refMV is a scaled motion vector from a neighboring block.
[0172] In this example, assume that if candMV is similar to the motion vector of a neighboring block, it is more reliable. The neighboring block can be the top or left block that has already been coded. Note that if the reference frame of the neighboring block is different from the MV reference frame, the motion vector can be scaled according to the temporal distance.
[0173] Example 3. refMV is scaled motion vector from co-located block.
[0174] Example 4. If bi-prediction is used, MVD sign / suffix prediction is applied only to one side. refMV is scaled motion vector from the other side.
[0175] Example 5. If bi-prediction is used, MVD sign / suffix prediction is applied to both sides. Bi-lateral motion vector matching is used to reorder MVD candidates. That is, candMV of one side is refMV of the other side.
[0176] In this example, it is assumed that there are 2 MVD candidates per side.
[0177] L0: candMV00, candMV01 L1: candMV10, candMV11 A total of 4 motion vector matching costs are calculated: (candMV00, candMV10), (candMV00, candMV11), (candMV01, candMV10) and (candMV01, candMV11), the matching pair with lower matching cost will be considered more reliable.
[0178] In the above examples, if refMV is not available, a default refMV can be used.
[0179] In the above examples, if refMV is not available, a refMV derived by other examples can be used. For example, if a certain candMV has no refMV due to no motion vector in the reference block in example 1, the refMV in example 2 can be used.
[0180] In another example, multiple refMVs are used to determine the motion vector matching cost. For example, refMV1 and refMV2 are reference motion vectors obtained using the methods described in example 1 and example 2 above, the average of these two refMVs is used to calculate the matching cost with candMV.
[0181] As mentioned above, the prioritized initial MVs for MV candidate reordering can be implemented at the encoder side or the decoder side. For example, any of the proposed prioritized initial MV methods can be implemented at the intra / inter coding module (e.g., intra prediction 150 / inter prediction 152) of the decoder or the intra / inter coding module (e.g., intra prediction 150 / inter prediction 152) of the encoder. Figure 1B Figure 1A The units (e.g., units 110 / 112 in FIG. 1A and units 150 / 152 in FIG. IB) can correspond to executable software or firmware code stored on a medium (e.g., a hard disk or flash memory) for a CPU (central processing unit) or programmable device (e.g., a DSP (digital signal processor) or FPGA (field programmable gate array)). The units can also be implemented in hardware that is coupled to the decoder or encoder, such as the intra / inter encoding module. For example, the intra / inter prediction 110 / inter prediction 112 can be implemented as circuitry coupled to the intra / inter encoding module in the encoder. Any proposed shared buffer for storing coding information among multiple coding tools, including the CCM mode, can also be implemented as circuitry coupled to the intra / inter encoding module of the decoder or encoder. However, the decoder or encoder can also use additional processing units to implement the required cross-component prediction processing. While the intra prediction 150 / inter prediction 152 can be coupled to the intra / inter encoding module of the encoder, the inter prediction 152 can be coupled to the intra / inter encoding module of the decoder. Although the units (e.g., units 110 / 112 in FIG. 1A and units 150 / 152 in FIG. IB) are shown as separate processing units, they can correspond to executable software or firmware code stored on a medium (e.g., a hard disk or flash memory) for a CPU (central processing unit) or programmable device (e.g., a DSP (digital signal processor) or FPGA (field programmable gate array)).
[0182] FIG. 20 illustrates a flowchart of an exemplary video coding system using a priority initial motion vector according to an embodiment of the present application. The steps shown in the flowchart can be implemented as program code executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart can also be implemented based on hardware, e.g., one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, input data related to a current block is received in step 2010, where the input data includes pixel data of the current block to be encoded at the encoder side or related data of the current block to be decoded at the decoder side. An initial motion vector (MV) of the current block is determined in step 2020. One or more motion vector candidates are also determined in step 2030. An initial TM cost associated with the initial motion vector (MV) is derived in step 2040, where the initial TM cost includes a first metric associated with a current template of the current block and a first template of a first reference block pointed to by the initial motion vector (MV). A modified initial TM cost is derived by adding an offset value to the initial TM cost in step 2050. One or more candidate TM costs of the one or more motion vector (MV) candidates are determined in step 2060, where the one or more candidate TM costs respectively include one or more second metrics associated with the current template of the current block and one or more second templates, and where each of the one or more second templates is associated with a second reference block pointed to by one of the one or more motion vector (MV) candidates. The initial MV and the one or more MV candidates are reordered according to the modified initial TM cost and the one or more candidate TM costs in step 2070.
[0183] Figure 21 A flowchart of another exemplary video coding system using a priority initial MV according to embodiments of the present application is shown. According to another approach, input data associated with a current block is received in step 2110, wherein the input data comprises pixel data of the current block to be encoded at the encoder side or data associated with the current block to be decoded at the decoder side. An initial MV (motion vector) of the current block is determined in step 2120. One or more MV candidates are also determined in step 2130. In step 2140, an initial TM cost associated with the initial MV is derived based on a current template of the current block and a first template of a first reference block pointed by the initial MV. In step 2150, one or more candidate TM costs of the one or more MV candidates are determined based on the current template of the current block and one or more second templates, respectively, wherein each of the one or more second templates is associated with a second reference block pointed by one of the one or more MV candidates. In step 2160, one or more weighted candidate TM costs are obtained, wherein each of the one or more weighted candidate TM costs is obtained by weighting one of the one or more candidate TM costs with a weighting factor, and the weighting factor is larger for a target MV candidate that is further away from the initial MV; or one or more modified candidate TM costs are obtained by reducing one or more target candidate TM costs with a smaller step size by an offset value. In step 2170, the initial MV and the one or more MV candidates are reordered according to the initial TM cost and the one or more weighted candidate TM costs or according to the initial TM cost and the one or more modified candidate TM costs.
[0184] The flowchart shown is intended to explain one example of video coding according to the present application. A skilled person can modify each step, rearrange the steps, split the steps, or combine the steps to practice the present application without departing from the spirit of the present application. In the present disclosure, specific syntax and semantics are used to explain examples of implementing embodiments of the present application. A skilled person can practice the present application by replacing the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present application.
[0185] The above description is intended to enable those skilled in the art to practice the present application as claimed. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments. Thus, the present application is not intended to be limited to the particular embodiments described. Rather, it is to be given broad scope in accordance with the principles defined herein and / or illustrated in the attached drawings. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present application. However, the skilled artisan will understand that the present application can be practiced without
[0186] Embodiments of the present application as described above can be implemented in various hardware, software code, or combinations thereof. For example, one embodiment of the present application can be one or more circuit circuits integrated into a video compression chip, or program codes integrated into video compression software to perform the processes described herein. One embodiment of the present application can also be program codes to be executed on a Digital Signal Processor (DSP) to perform the processes described herein. The present application can also involve several processes performed by a computer processor, a digital signal processor, a microprocessor, or a field programmable gate array (FPGA). These processors can be configured to perform particular methods by executing software or firmware codes in machine-readable media. The software codes or firmware codes can be developed using different programming languages and different formats or styles. The software codes can also be compiled into different formats or styles for different target platforms. However, different code formats, styles, and languages of software codes and other means of configuring codes to perform the methods in accordance with the present application will not depart from the spirit and scope of the application.
[0187] The present application can be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the application is, therefore, indicated by the appended claims, rather than by the foregoing description. All changes that come within the meaning of the claims are intended to be embraced therein.
Claims
1. A method of video coding, comprising: receiving input data associated with a current block, wherein the input data comprises pixel data of the current block to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side; determining an initial MV (motion vector) of the current block; determining one or more MV candidates; deriving an initial TM cost associated with the initial MV, wherein the initial TM cost comprises a first metric associated with a current template of the current block and a first template of a first reference block pointed by the initial MV; deriving a modified initial TM cost by adding an offset value to the initial TM cost; determining one or more candidate TM costs for the one or more MV candidates, wherein the one or more candidate TM costs comprise one or more second metrics associated with the current template of the current block and one or more second templates, and wherein each of the one or more second templates is associated with a second reference block pointed by one of the one or more MV candidates; and reordering the initial MV and the one or more MV candidates according to the modified initial TM cost and the one or more candidate TM costs.
2. The method of claim 1, wherein the offset value corresponds to a positive value or a negative value.
3. The method of claim 1, wherein the offset value corresponds to a fixed value or the offset value depends on the initial TM cost.
4. The method of claim 1, wherein the offset value is signaled or parsed from a bitstream.
5. The method of claim 1, wherein the offset value is determined according to coding information.
6. The method of claim 5, wherein the coding information comprises a template region, a QP (quantization parameter), or both, and wherein the template region is associated with the current template, the first template, the one or more second templates, or a combination thereof.
7. The method of claim 1, wherein the one or more MV candidates are associated with an inter direction of motion compensation, and the inter direction of motion compensation comprises uni-prediction L0, uni-prediction Ll, bi-prediction, more than two hypotheses of prediction, or a combination thereof, and if one MV candidate has a same inter direction as the initial MV, a respective TM cost of the one MV candidate is reduced by a second offset value, and the second offset value corresponds to a positive value.
8. The method of claim 1, wherein the one or more MV candidates are associated with weights of BCW (bi-prediction with CU level weights), and if one MV candidate has a same BCW weight as the initial MV, a respective TM cost of the one MV candidate is reduced by a second offset value, and the second offset value corresponds to a positive value.
9. The method of claim 1, wherein each of the one or more candidate TM costs further comprises a weighted candidate MV difference between the initial MV and one of the one or more MV candidates.
10. The method of claim 9, wherein a weighting factor of the weighted candidate MV difference corresponds to a fixed value or depends on coding information.
11. The method of claim 9, wherein the weighting factor of the weighted candidate MV difference is predefined, or explicitly signaled or parsed from a bitstream.
12. The method of claim 1, wherein the first metric and the one or more second metrics correspond to SAD (sum of absolute difference), SSD (sum of squared difference), or SATD (sum of absolute transformed difference).
13. A video coding device comprising one or more electronic devices or processors configured to: receive input data related to a current block, wherein the input data comprises pixel data of the current block to be encoded at an encoder side or data related to the current block to be decoded at a decoder side; determine an initial MV (motion vector) of the current block; determine one or more MV candidates; derive an initial TM cost related to the initial MV, wherein the initial TM cost comprises a first metric related to a current template of the current block and a first template of a first reference block pointed by the initial MV; derive a modified initial TM cost by adding an offset value to the initial TM cost; determine one or more candidate TM costs for the one or more MV candidates, wherein the one or more candidate TM costs comprise one or more second metrics and one or more second templates, and wherein each of the one or more second templates is associated with a second reference block pointed by one of the one or more MV candidates; and reorder the initial MV and the one or more MV candidates according to the modified initial TM cost and the one or more candidate TM costs. receive input data related to a current block, wherein the input data comprises pixel data of the current block to be encoded at an encoder side or data related to the current block to be decoded at a decoder side; determine an initial MV (motion vector) of the current block; determine one or more MV candidates; derive an initial TM cost related to the initial MV based on a current template of the current block and a first template of a first reference block pointed by the initial MV; determine one or more candidate TM costs for the one or more MV candidates respectively based on the current template of the current block and one or more second templates, wherein the one or more candidate TM costs comprise one or more second metrics, and wherein each of the one or more second templates is associated with a second reference block pointed by one of the one or more MV candidates; and derive one or more weighted candidate TM costs, wherein each of the one or more weighted candidate TM costs is derived by weighting one of the one or more candidate TM costs with a weighting factor, and the weighting factor is larger for a target MV candidate that is farther away from the initial MV, or derive one or more modified candidate TM costs by reducing one or more target candidate TM costs with a smaller step size by an offset value.
14. A method of video coding, the method comprising: and reordering the initial MV and the one or more MV candidates based on the initial TM cost and the one or more weighted candidate TM costs, or based on the initial TM cost and the one or more modified candidate TM costs.
15. The method of claim 14, wherein the one or more MV candidates comprise inter and affine MMVD (merge mode with motion vector difference) candidates with respective step sizes and directions.
16. A video coding apparatus, comprising one or more electronic devices or processors configured to: receive input data related to a current block, wherein the input data comprises pixel data of the current block to be encoded at an encoder side or data related to the current block to be decoded at a decoder side; determine an initial MV (motion vector) of the current block; determine one or more MV candidates; derive an initial TM cost related to the initial MV based on a current template of the current block and a first template of a first reference block pointed by the initial MV; determine one or more candidate TM costs of the one or more MV candidates respectively based on the current template of the current block and one or more second templates, wherein the one or more candidate TM costs comprise one or more second metrics, and wherein each of the one or more second templates is associated with a second reference block pointed by one of the one or more MV candidates; and derive one or more weighted candidate TM costs, wherein each of the one or more weighted candidate TM costs is derived by weighting one of the one or more candidate TM costs with a weighting factor, and the weighting factor is larger for a target MV candidate that is further away from the initial MV, or derive one or more modified candidate TM costs by reducing one or more target candidate TM costs with a smaller step size by an offset value; and reorder the initial MV and the one or more MV candidates based on the initial TM cost and the one or more weighted candidate TM costs, or based on the initial TM cost and the one or more modified candidate TM costs.