Inter prediction in video coding
By strategically ordering and executing TM and BM processes in video coding and decoding, the method addresses complexity issues in VVC, enhancing compression efficiency and performance.
Patent Information
- Application Number
- CN202380084018.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-05
- Filing Date
- 2023-09-21
- Publication Date
- 2025-07-15
AI Technical Summary
With the increasing demand for efficient video data transmission and storage, the codec complexity and efficiency of inter-frame prediction codec tools still have room for improvement, especially under the framework of the general video codec standard VVC, more efficient codec tools are needed to improve performance.
The strategic sequence of template matching (TM) and bilateral matching (BM) processes is performed in a strategic order, and combined with the early termination mechanism, the inter-frame prediction encoding and decoding process is optimized. By receiving and transmitting the encoding and decoding units in the video bitstream, the encoding and decoding complexity and performance are improved.
By optimizing the sequence of TM and BM processes and the early termination mechanism, the complexity of encoding and decoding is reduced, the encoding and decoding efficiency and performance are improved, and the requirements of different encoding and decoding environments are adapted.
Smart Images

Figure CN120323022A_ABST
Abstract
Description
[0001]
Cross - References
[0002] This application claims the benefit of priority of Provisional Application No. 63 / 378,372, filed on October 5, 2022. The disclosure of the prior application is hereby incorporated by reference in its entirety.
Technical Field
[0003] This disclosure relates to video encoding and decoding.
Background Art
[0004] With the growing demand for efficient video data transmission and storage, the need for powerful video encoding and decoding technologies is increasing. After the finalization of the Versatile Video Coding (VVC) standard, the video encoding and decoding community aims to standardize future video encoding and decoding technologies. As part of the effort, a common software test platform, the Enhanced Compression Model (ECM), has been developed to explore the potential standardization of advanced video encoding and decoding technologies.
[0005] Many inter - frame encoding and decoding tools have been studied on the ECM to evaluate their functions and performance and to decide whether to adopt them in the ECM. For example, to further provide Bjontegaard Delta - Rate savings, the current ECM includes a set of inter - frame prediction encoding and decoding tools, including Template Matching (TM), multi - process decoder - side motion vector refinement (or Bilateral Matching (BM)), Local Illumination Compensation (LIC), Non - Adjacent Spatial Candidate, Overlapped Block Motion Compensation (OBMC), Multi - Hypothesis Prediction (MHP), Bilateral Matching AMVP - Merge mode, etc.
[0006]
Summary of the Invention
Summary of the Invention
[0007] Aspects of the present disclosure provide a method for performing inter-frame prediction in a video decoder. The method includes receiving a coded unit in a bitstream of a video. The coded unit is encoded using a Template Matching (TM) process and a Bilateral Matching (BM) process. The method also includes determining an order of the TM process and the BM process. The method further includes performing inter-frame prediction based on the determined order of the TM process and the BM process to reconstruct the received coded unit.
[0008] Another aspect of the present disclosure provides a method for performing inter-frame prediction in a video encoder. The method includes performing inter-frame prediction based on a determined order of a Template Matching (TM) process and a Bilateral Matching (BM) process to encode a coded unit. The method also includes transmitting the encoded coded unit in a bitstream of the video. **BRIEF DESCRIPTION OF THE DRAWINGS**
[0009] Various embodiments of the present disclosure will be described in detail by way of example with reference to the following drawings, in which like numerals represent like elements, and:
[0010] Figure 1 A block diagram of a video encoder according to an embodiment of the present disclosure is shown;
[0011] Figure 2 A block diagram of a video decoder according to an embodiment of the present disclosure is shown;
[0012] Figures 3A and 3B respectively show flowcharts for performing inter-frame prediction in a video encoder and a video decoder according to embodiments of the present disclosure;
[0013] Figure 4 Various types of tree segmentation patterns are shown;
[0014] Figure 5 An example of a quadtree with a nested multi-type tree coded block structure is shown;
[0015] Figure 6 A search point layout in a Merge Mode with Motion Vector Difference (MMVD) is shown;
[0016] Figure 7A and Figure 7B respectively show 4-parameter and 6-parameter affine motion models based on control points;
[0017] Figure 8 An example of an affine motion vector field (MVF) for each sub-block is shown;
[0018] Figure 9 Shows the positions of the inherited affine motion predictors;
[0019] Figure 10 Shows an example of the inheritance of control point motion vectors;
[0020] Figure 11 Shows the candidate positions for constructing the affine merge mode;
[0021] Figure 12 Illustrates an example of DMVR;
[0022] Figure 13 Shows an example of segmentation by geometric partition mode (GPM) grouped by the same angle;
[0023] Figure 14 Shows the top and left neighboring blocks used in the combined inter-intra prediction (CIIP) weight derivation;
[0024] Figure 15 Shows the template matching (TM) performed on the search area around the initial motion vector (MV);
[0025] Figure 16 Shows five diamond search areas in the search area of the second pass of multi-pass DMVR.
Detailed implementation manners
[0026] The present disclosure provides different embodiments or examples for implementing different features of the provided subject matter. To simplify the present disclosure, specific examples of specific components and configurations are described below. Of course, these are only examples and are not intended to be limiting.
[0027] For example, for clarity, the order of discussion of the different steps described here has been presented. Generally, these steps can be performed in any suitable order. In addition, although different features, techniques, configurations, etc. may be discussed in different places of the present disclosure, the intention is that each concept can be implemented independently of other concepts or in combination with other concepts. Therefore, the present disclosure can be embodied and viewed in many different ways.
[0028] In addition, as used herein, words such as "a", "an", etc. generally have the meaning of "one or more" unless otherwise specified.
[0029] Figure 1Shows a block diagram of a video encoder, which may include or be connected to modules or circuits that implement the methods and techniques described in this disclosure. The video encoder may be implemented based on the Versatile Video Coding (VVC) standard, the High-Efficient Video Coding (HEVC) standard (with the addition of the Adaptive Loop Filter (ALF)), or any other video coding standard.
[0030] When using the inter-frame mode, the intra / inter-frame prediction unit 110 generates an inter-frame prediction based on Motion Estimation (ME) / Motion Compensation (MC).
[0031] When using the intra-frame mode, the intra / inter-frame prediction unit 110 generates an intra-frame prediction. The intra / inter-frame prediction data (i.e., the intra / inter-frame prediction signal) is provided to the subtractor 115 to form a prediction error, also known as the "residual" or "residue", by subtracting the intra / inter-frame prediction signal from the signal related to the input frame. The process of generating the intra / inter-frame prediction data is referred to as the prediction process in this disclosure. Then, the prediction error (i.e., the residual) is processed by Transform (T) followed by Quantization (Q) (T+Q, 120). The transformed and quantized residual is then decoded by the entropy coding / decoding unit 125 to be included in the video bitstream corresponding to the compressed video data.
[0032] The bitstream related to the transform coefficients is then packed together with side information, such as motion, coding / decoding mode, and other information related to the image region. The side information can also be compressed by entropy coding / decoding to reduce the required bandwidth. Since the reconstructed frame may be used as a reference frame for inter-frame prediction, one or more reference frames must also be reconstructed at the encoder side. Therefore, the transformed and quantized residual is processed by Inverse Quantization (IQ) and Inverse Transformation (IT) (IQ+IT, 130) to recover the residual. The reconstructed residual is then added back to the intra / inter-frame prediction data at the Reconstruction unit (REC) 135 to reconstruct the video data. In this disclosure, the process of adding the reconstructed residual to the intra / inter-frame prediction signal is referred to as the reconstruction process. The output frame of the reconstruction process is called the reconstructed frame.
[0033] The bitstream related to the transform coefficients is then packed together with side information, such as motion, coding / decoding mode, and other information related to the image region. The side information can also be compressed by entropy coding / decoding to reduce the required bandwidth. Since the reconstructed frame may be used as a reference frame for inter-frame prediction, one or more reference frames must also be reconstructed at the encoder side. Therefore, the transformed and quantized residual is processed by Inverse Quantization (IQ) and Inverse Transformation (IT) (IQ+IT, 130) to recover the residual. The reconstructed residual is then added back to the intra / inter-frame prediction data at the Reconstruction unit (REC) 135 to reconstruct the video data. In this disclosure, the process of adding the reconstructed residual to the intra / inter-frame prediction signal is referred to as the reconstruction process. The output frame of the reconstruction process is called the reconstructed frame.
[0034] The bitstream related to the transform coefficients is then packed together with side information, such as motion, coding / decoding mode, and other information related to the image region. The side information can also be compressed by entropy coding / decoding to reduce the required bandwidth. Since the reconstructed frame may be used as a reference frame for inter-frame prediction, one or more reference frames must also be reconstructed at the encoder side. Therefore, the transformed and quantized residual is processed by Inverse Quantization (IQ) and Inverse Transformation (IT) (IQ+IT, 130) to recover the residual. The reconstructed residual is then added back to the intra / inter-frame prediction data at the Reconstruction unit (REC) 135 to reconstruct the video data. In this disclosure, the process of adding the reconstructed residual to the intra / inter-frame prediction signal is referred to as the reconstruction process. The output frame of the reconstruction process is called the reconstructed frame.
[0035] To reduce defects in the reconstructed frames, loop filters are used, including but not limited to Deblocking Filter (DF) 140, Sample Adaptive Offset (SAO) 145, and Adaptive Loop Filter (ALF) 150. In the present disclosure, DF, SAO, and ALF are all labeled as filtering processes. The filtered reconstructed frames output by all filtering processes are referred to as decoded frames in the present disclosure. The decoded frames are stored in Frame Buffer 155 and are used for the prediction of other frames.
[0036] Figure 2 A block diagram of a video decoder is shown. The decoder may include or be connected to modules or circuits that implement the methods and techniques described in the present disclosure. The video decoder may be implemented based on the VVC standard, the HEVC standard (with ALF added), or any other video coding and decoding standard. Since the encoder includes a local decoder for reconstructing video data, many decoder components have already been used in the encoder (such as Frame Buffer 255, ALF 250, SAO 245, REC 235, and DF 240) in addition to the entropy decoder. At the decoder side, the entropy decoding unit 226 is used to recover the coded symbols or syntax from the bitstream. The coded residuals generated by the entropy decoding process are processed through inverse quantization (IQ) and inverse transformation (IT) (IQ+IT, 230) to recover the residuals. The process of generating the reconstructed residuals from the input bitstream is referred to as the residual decoding process in the present disclosure. The prediction process for generating intra / inter prediction data is also applied at the decoder side. However, the intra / inter prediction unit 211 at the decoder side is different from the intra / inter prediction unit 110 at the codec side because inter prediction only needs to perform motion compensation using the motion information derived from the bitstream. In addition, the reconstructed residuals are added to the intra / inter prediction data using adder 215.
[0037] The present disclosure generally relates to video coding and decoding. In particular, the disclosure relates to the use of Template Matching (TM) and Bilateral Matching (BM, or Decoder-Side Motion Vector Refinement (DMVR)) in video coding and decoding systems.
[0038] In the ECM, the BM or DMVR process may include multiple passes. In the first pass, a bilateral matching process is applied to the coding / decoding block. In the second pass, the bilateral matching process is applied to each 16x16 sub-block within the coding / decoding block. In the third pass, the motion vectors (MVs) in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for spatial and temporal motion vector prediction.
[0039] Note that the concept of "early termination" can be incorporated into the multi-pass DMVR process. For example, if the sum of absolute differences (SAD) result from the block-level BM pass is lower than a certain threshold, the BM process can end early; there is no need to continue with the subsequent sub-block-level BM and BDOF procedures.
[0040] These three passes are described in detail elsewhere in this disclosure.
[0041] Figure 3A and Figure 3B respectively show the flowcharts for performing inter-frame prediction in a video decoder and a video encoder according to embodiments of this disclosure. The TM and BM processes can be executed strategically for inter-frame prediction, thereby reducing coding / decoding complexity and improving coding / decoding performance.
[0042] As Figure 3A shown in process 300, it can be performed in a video decoder. In step S305, a coding / decoding unit is received from the bitstream of the video. The coding / decoding unit is coded / decoded using the TM process and the BM process. In step S315, the order of the TM and BM processes is determined according to the syntax elements received from the bitstream. In step S325, inter-frame prediction is performed to reconstruct the coding / decoding unit according to the determined order of the TM and BM processes.
[0043] As Figure 3A shown in the embodiment, the order of performing the TM and BM processes is determined according to certain syntax elements indicating the order. However, those skilled in the art can recognize that a predefined order of the TM and BM processes can be used. In this case, the video decoder does not need to parse the syntax elements.
[0044] As Figure 3B shown in process 350, it can be performed in a video encoder. In step S355, inter-frame prediction is performed to code / decode the coding / decoding unit according to the determined order of the TM and BM processes. In step S365, syntax elements indicating the order of the TM and BM processes are signaled in the bitstream of the video. In step S375, the coded / decoded coding / decoding unit is transmitted in the bitstream.
[0045] As described above, a predefined order can be used as the determined order of performing the TM and BM processes. In this case, the video codec does not need to signal the syntax elements.
[0046] Similar to the ECM 6.0 method, the BM process may include block-level MV refinement, sub-block-level MV refinement, and sub-block-level BDOF MV refinement in sequence. In one embodiment, the TM process may be executed immediately after the block-level MV refinement of the BM process. In another alternative embodiment, the BM process may be executed after the TM process. Both embodiments may adopt early termination to improve encoding and decoding efficiency. In other words, whether to execute subsequent procedures may depend on the cost of the previous procedure.
[0047] For example, consider a scenario where the execution of the TM process (if the TM process is executed) is after the block-based (or CU-based) motion vector refinement of the BM process. In this case, if the minimum cost of the block-level MV refinement is less than or equal to a threshold, the TM process may be prohibited.
[0048] Alternatively, in a scenario where the TM process is executed, if the TM process is after the block-based (or CU-based) motion vector refinement of the BM process, when the minimum cost of the block-level MV refinement is less than or equal to a threshold, the TM process may be executed with a smaller search range instead of being prohibited.
[0049] In another example, in a scenario where the TM process is executed, if the TM process is after the block-based (or CU-based) motion vector refinement of the BM process, the sub-block-level MV refinement may always be executed regardless of the cost of the block-level MV refinement.
[0050] For example, since the TM process is executed immediately before the potential sub-block-level MV refinement, if the cost of the TM process is less than or equal to a threshold, the sub-block-level MV refinement may be prohibited.
[0051] Consider a scenario where the BM process is executed. If the BM process is after the TM process, in some embodiments, when the cost of the TM process is less than or equal to a threshold, the BM process may be prohibited.
[0052] Alternatively, in a scenario where the BM process is executed, if the BM process is after the TM process, when the cost of the TM process is less than or equal to a threshold, or the cost of the block-based MV refinement is less than or equal to a threshold, or the potential cost reduction achieved by the best block-level MV refinement cannot exceed the cost reduction achieved by the initial block-level MV refinement (performed on the MV derived from the TM process), the sub-block-level MV refinement may be prohibited.
[0053] As another example, in a scenario where the BM process is executed, if the BM process is after the TM process, when the cost of block-level MV refinement is less than or equal to a threshold, or the potential cost reduction achieved by the best block-level MV refinement cannot exceed the cost reduction achieved by the initial block-level MV refinement (performed on the MV derived from the TM process), the MV modification from the BM process may not be used.
[0054] In the above example, the cost can be calculated using an appropriately selected function without limitation. For example, the cost of block-level MV refinement can be calculated as the sum of the motion vector distance cost (mvDistanceCost) and the SAD cost (sadCost). However, those skilled in the art can recognize that other forms of cost are possible.
[0055] In the above example, the threshold can be a predefined non-negative integer. Alternatively, the threshold can be adaptively determined based on the codec information. For example, the threshold can be determined based on the number of samples in the current block / coding unit, the inter-frame direction of the reference picture and the current picture, the quantization parameter (QP) of the reference picture and the current picture, the number of samples of the template, etc.
[0056] Aspects of the present disclosure can be further described as follows.
[0057] I. Video coding and decoding method
[0058] 1. Segment CTUs using a tree structure
[0059] In the High-Efficient video Coding standard (HEVC), pictures are partitioned into a series of coding tree units (CTUs). For pictures with three sample arrays, a CTU consists of an NxN luma sample block and two corresponding chroma sample blocks, or for pictures coded using three separate color planes, a CTU consists of an NxN monochrome plane sample block. The concept of CTU is roughly similar to the macroblock concept in previous standards such as Advanced video Coding (AVC). In the Main profile, the maximum allowed size of the luma block in a CTU is specified as 64x64. CTUs are split into coding units (CUs) by using a quaternary-tree structure to adapt to various local characteristics. The decision of whether to use inter (temporal) or intra (spatial) prediction to code a picture region is made at the leaf CU level. Each leaf CU can be further split into one, two, or four prediction units (PUs) according to the PU split type. Within a PU, the same prediction process is applied, and the relevant information is transmitted to the decoder based on the PU. After obtaining the residual block by applying the prediction process according to the PU split type, a leaf CU can be split into transform units (TUs) by another quaternary tree structure similar to the CU coding tree. A key feature of the HEVC structure is that it has multiple split concepts, including CUs, PUs, and TUs.
[0060] The Versatile video Coding standard (VVC) is the successor of HEVC. In VVC, a quadtree with a nested multi-type tree structure using binary and ternary splits replaces the concept of multiple split unit types, that is, it removes the separation of the CU, PU, and TU concepts except for CUs with too large a maximum transform length and supports more flexibility in CU split shapes. In the coding tree structure, a CU can have a square or rectangular shape. A coding tree unit (CTU) is first split by a quaternary tree structure. Then, the quadtree leaf nodes can be further split by a multi-type tree structure. Figure 4 Shows various types of multi-type tree split patterns. As Figure 4As shown, there are four splitting types in the multi-type tree structure: vertical binary splitting (SPLIT_BT_VER), horizontal binary splitting (SPLIT_BT_HOR), vertical ternary splitting (SPLIT_TT_VER), and horizontal ternary splitting (SPLIT_TT_HOR). The leaf nodes of the multi-type tree are called coding units (CUs). Unless the CU is too large for the maximum transform length, this splitting is used for prediction and transform processing without any further splitting. This means that, in most cases, the CUs, PUs, and TUs have the same block size in the quadtree coding block structure with nested multi-type trees. An exception occurs when the maximum supported transform length is less than the width or height of the color components of the CU.
[0061] An example of a quadtree with a nested multi-type tree coding block structure is Figure 5 shown. Figure 5 shown. It shows a CTU split into multiple CUs with a quadtree and nested multi-type tree coding block structure, where the thick block edges represent quadtree splitting and the remaining edges represent multi-type tree splitting. The quadtree with nested multi-type tree splitting provides a content-adaptive coding tree structure composed of CUs. The size of the CU can be as large as the CTU or as small as a 4×4 luminance sample unit. For the 4:2:0 chroma format, the maximum chroma CB size is 64×64, and the smallest chroma CB contains 16 chroma samples.
[0062] In VVC, the maximum supported luminance transform size is 64×64, and the maximum supported chroma transform size is 32×32. When the width or height of the CB is greater than the maximum transform width or height, the CB is automatically split in the horizontal and / or vertical directions to meet the transform size limit in that direction.
[0063] In VVC, the coding tree scheme supports the ability to have independent block tree structures for luminance and chrominance. For P and B slices, the luminance and chrominance CTBs in a CTU must share the same coding tree structure. However, for I slices, the luminance and chrominance can have independent block tree structures. When the independent block tree mode is applied, the luminance CTB is split into CUs through one coding tree structure, and the chrominance CTB is split into chrominance CUs through another coding tree structure. This means that the CUs in I slices may consist of coding blocks of the luminance component or coding blocks of both chrominance components, while the CUs in P or B slices always consist of coding blocks of all three color components, unless the video is monochrome.
[0064] For each inter-predicted CU, the motion parameters consist of a motion vector, a reference picture index, a reference picture list use index, and additional information required by the new coding and decoding features of VVC for inter-predicted sample generation. The motion parameters can be flagged in an explicit or implicit manner. When a CU is coded / decoded in skip mode, the CU is associated with a PU and there are no significant residual coefficients, no coded motion vector differences or reference picture indices. A merge mode is specified, in which the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates, and additional schemes introduced in VVC. The merge mode can be applied to any inter-predicted CU, not just the skip mode. The alternative to the merge mode is the explicit transmission of the motion parameters, in which the motion vector of each CU, the reference picture index for each corresponding reference picture list, the reference picture list use flag, and other required information are explicitly flagged.
[0065] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 5) are studying the potential needs for standardization of future video coding technologies, whose compression capabilities significantly exceed the current VVC standard. An Enhanced Compression Model (ECM) reference software is provided to demonstrate the reference implementation and decoding process of the coding technologies explored in the JVET enhanced compression beyond VVC capabilities work. ECM is basically a successor to VVC, so it has many common parts with VVC.
[0066] 2. Overview of Inter-Prediction
[0067] In High Efficiency Video Coding (HEVC), for each inter-prediction unit (PU), one of three prediction modes can be selected, including inter, skip, and merge. Generally, a motion vector competition (MVC) scheme is introduced to select a motion candidate from a given candidate set including spatial and temporal motion candidates. Multiple reference for motion estimation allows finding the best reference in 2 possible reconstructed reference picture lists, namely list 0 and list 1. For the inter mode (informally called the AMVP mode, where AMVP stands for Advanced Motion Vector Prediction), an inter-prediction indicator (list 0, list 1, or bi-prediction), a reference index, a motion candidate index, motion vector differences (MVDs), and a prediction residual are transmitted. As for the skip mode and the merge mode, only the merge index is transmitted, and the current PU inherits the inter-prediction indicator, the reference index, and the motion vector of the neighboring PU referenced by the coded merge index. In the case of a coded unit (CU) coded in skip mode, the residual signal is also omitted.
[0068] In VVC, the AMVP mode is further improved by new modes such as the symmetric motion vector difference (SMVD) mode, adaptive motion vector resolution (AMVR), and affine AMVP mode; the merge / skip mode is further improved by enhanced merge candidates, combined inter-intra prediction (CIIP), affine merge mode, sub-block temporal motion vector prediction (SbTMVP), merge mode with motion vector difference (MMVD), and geometric partition mode (GPM). In VVC, decoder-side motion vector refinement (DMVR), bidirectional optical flow (BDOF), and prediction optical flow refinement (PROF) are used to refine the motion vectors or motion compensation predictors at the decoder side.
[0069] In ECM, several new codec tools are developed to further improve the AMVP, merge, and skip modes, such as bilateral matching AMVP-merge mode, multiple hypothesis prediction (MHP), overlapping block motion compensation (OBMC), etc. In addition, template matching-based decoder-side motion vector refinement is proposed to enhance the codec efficiency of inter prediction.
[0070] In addition to the inter-frame codec features in HEVC, VVC includes many new and improved inter-frame prediction codec tools listed below:
[0071] – Extended merge prediction
[0072] – Merge mode with MVD (MMVD)
[0073] – Symmetric MVD (SMVD) signaling
[0074] – Affine motion compensation prediction
[0075] – Sub-block based temporal motion vector prediction (SbTMVP)
[0076] – Adaptive motion vector resolution (AMVR)
[0077] – Motion field storage: 1 / 16th luminance sample MV storage and 8x8 motion field compression
[0078] – Bi-directional prediction with CU-level weights (BCW)
[0079] – Bidirectional optical flow (BDOF)
[0080] – Decoder-side motion vector refinement (DMVR)
[0081] – Geometric partition mode (GPM)
[0082] – Combined inter and intra prediction (CIIP)
[0083] After the finalization of VVC, an Enhanced Compression Model (ECM) reference software was developed to investigate the potential requirements for future video coding technology standardization. In the current ECM, several inter prediction coding tools are included to provide further BD-rate savings:
[0084] – Local Illumination Compensation (LIC)
[0085] – Non-Contiguous Spatial Candidates
[0086] – Template Matching (TM)
[0087] – Overlapped Block Motion Compensation (OBMC)
[0088] – Multiple Hypothesis Prediction (MHP)
[0089] – Bi-directional Matching AMVP - Merge Mode
[0090] – And some other tools under development
[0091] The following text provides details of some selected inter prediction methods in VVC and ECM.
[0092] 3. Extended Merge Prediction
[0093] In VVC, the merge candidate list is constructed by sequentially including the following five types of candidates:
[0094] 1) Spatial MVP from spatially neighboring CUs
[0095] 2) Temporal MVP from collocated CUs
[0096] 3) History-based MVP from the FIFO table
[0097] 4) Paired-Average MVP
[0098] 5) Zero motion vector.
[0099] The size of the merge list is flagged in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU coded in merge mode, the index of the best merge candidate is coded using truncated unary binarization (TU). The first binary bit of the merge index is coded using context coding, and the other binary bits are coded using bypass coding.
[0100] This section provides the derivation process for each type of merge candidate. Similar to what was done in HEVC, VVC also supports the parallel derivation of the merge candidate list (or merge candidates) for all CUs within a certain size region.
[0101] 4. History-based Merge Candidate Derivation
[0102] After spatial MVP and TMVP, history-based MVP (HMVP) merge candidates are added to the merge list. In this method, the motion information of previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. A table containing multiple HMVP candidates is maintained during the encoding / decoding process. When a new CTU row is encountered, the table is reset (emptied). Whenever there is a CU with non-sub-block inter-frame encoding / decoding, the relevant motion information is added as a new HMVP candidate to the last entry of the table.
[0103] The size S of the HMVP table is set to 6, which means that at most 5 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a restricted first-in-first-out (FIFO) rule is used. First, a redundancy check is applied to find if there is the same HMVP in the table. If found, the same HMVP is deleted from the table, and all subsequent HMVP candidates are moved forward, and then the same HMVP is inserted into the last entry of the table.
[0104] HMVP candidates can be used in the merge candidate list construction process. The last few HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidates. A redundancy check is performed on the HMVP candidates to facilitate spatial or temporal merge candidates.
[0105] To reduce the number of redundancy check operations, the following simplification measures are introduced:
[0106] 1) The last two entries in the table perform redundancy checks on the A1 and B1 spatial candidates respectively.
[0107] 2) Once the total number of available merge candidates reaches the maximum allowed number of merge candidates minus 1, the process of constructing the merge candidate list from HMVP terminates.
[0108] 5. Paired-Average Merge Candidate Derivation
[0109] Pairwise average candidates are generated by averaging a predefined candidate pair in the existing merge candidate list using the first two merge candidates. The first merge candidate is defined as p0Cand, and the second merge candidate is defined as p1Cand. The average motion vector is calculated for each reference list according to the availability of the motion vectors of p0Cand and p1Cand. If both motion vectors in a list are available, even if they point to different reference pictures, the two motion vectors are averaged, and the reference picture is set to the reference picture of p0Cand; if only one motion vector is available, the motion vector is directly used; if no motion vector is available, the list is kept invalid. In addition, if the half-pixel interpolation filter indices of p0Cand and p1Cand are different, they are set to 0.
[0110] When the merge list is not full after adding the pairwise average merge candidate, zero MVPs are inserted at the end until the maximum number of merge candidates is reached
[0111] 6. Merge Mode with Motion Vector Difference (MMVD)
[0112] In addition to the merge mode in which the implicitly derived motion information is directly used for the prediction sample generation of the current CU, VVC introduces a merge mode with motion vector difference (MMVD). After sending the regular merge flag, the MMVD flag is immediately sent to specify whether the MMVD mode is used for the CU.
[0113] In MMVD, after the merge candidate is selected, it is further refined by the signaled MVD information. The further information includes the merge candidate flag, the index specifying the motion magnitude, and the index indicating the motion direction. In the MMVD mode, one of the first two candidates in the merge list is selected as the MV basis. The mmvd candidate flag is signaled to specify which one is used between the first and second merge candidates.
[0114] The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. Figure 6 Shows the search point layout in the merge mode with motion vector difference (MMVD). As Figure 6 shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1.
[0115] Table 1 - Relationship between distance index and predefined offset
[0116]
[0117] The direction index represents the MVD direction relative to the starting point. The direction index can represent the four directions as shown in Table 2. It should be noted that the meaning of the MVD symbol may vary according to the information of the starting MV. When the starting MV is an unpredicted MV or a bi-predicted MV and both lists point to the same side of the current picture (i.e., both reference POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbols in Table 2 specify the sign of the MV offset added to the starting MV. When the starting MV is a bi-predicted MV and the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture and the other reference POC is less than the POC of the current picture), and the POC difference in list 0 is greater than the POC difference in list 1, the symbols in Table 2 specify the sign of the MV offset of the starting MV added to the list 0 MV component, while the sign of the list 1 MV has the opposite value. Otherwise, if the POC difference in list 1 is greater than list 0, the symbols in Table 2 specify the sign of the MV offset of the starting MV added to the list 1 MV component, while the sign of the list 0 MV has the opposite value.
[0118] The MVD is scaled according to the POC difference in each direction. If the POC differences in both lists are the same, no scaling is required. Otherwise, if the POC difference in list 0 is greater than the POC difference in list 1, the MVD of list 1 is scaled, defining the POC difference of L0 as td and the POC difference of L1 as tb. If the POC difference of L1 is greater than L0, the MVD of list 0 is scaled in the same way. If the starting MV is unidirectionally predicted, the MVD is added to the available MV.
[0119] Table 2 - MV offset symbols specified by the direction index
[0120] Direction IDX 00 01 10 11 X-axis + - N / A N / A y-axis N / A N / A + -
[0121] 7. Affine motion compensation prediction
[0122] In HEVC, only the translational motion model is applied for motion compensation prediction (MCP). However, in the real world, there are many kinds of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transform motion compensation prediction is applied. Figure 7A and Figure 7B respectively show the 4-parameter and 6-parameter affine motion models based on control points. As Figure 7A and Figure 7B shown, the affine motion field of a block is described by the motion information of two control points (4-parameter) or three control point motion vectors (6-parameter).
[0123] For a 4-parameter affine motion model, the motion vector for a sample position (x, y) in a block is derived as:
[0124]
[0125] For a 6-parameter affine motion model, the motion vector for a sample position (x, y) in a block is derived as:
[0126]
[0127] where (mv 0x , mv 0y ) is the motion vector of the top-left control point, (mv 1x , mv 1y ) is the motion vector of the top-right control point, and (mv 2x , mv 2y ) is the motion vector of the bottom-left control point.
[0128] To simplify motion compensation prediction, block-based affine transform prediction is applied. Figure 8 Shows an example of the affine motion vector field (MVF) for each sub-block. To derive the motion vector for each 4×4 luminance sub-block, as Figure 8 shown, the motion vector for the center sample of each sub-block is calculated according to the above equations and rounded to 1 / 16 fractional precision. Then a motion compensation interpolation filter is applied to generate the prediction for each sub-block with the derived motion vector. The sub-block size for the chrominance components is also set to 4×4. The MV for a 4×4 chrominance sub-block is calculated as the average of the MVs of the top-left and bottom-right luminance sub-blocks in the co-located 8×8 luminance region.
[0129] Similar to translational motion inter prediction, there are also two affine motion inter prediction modes: affine merge mode and affine AMVP mode.
[0130] 8. Affine Merge Prediction
[0131] The AF_MERGE mode can be applied to coding units (CUs) with a width and height both greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPMV candidates, and an index is sent to indicate which one is used for the current CU. The following three types of CPMV candidates are used to form the affine merge candidate list:
[0132] – Inherited affine merge candidates extrapolated from the CPMVs of neighboring CUs
[0133] – Constructed affine merge candidate CPMV derived using the translational MVs of neighboring CUs
[0134] – Zero MV
[0135] Figure 9 Shows the positions of the inherited affine motion predictors. Figure 10 Shows an example of inheritance of the control point motion vectors. Figure 11 Shows the candidate positions for constructing the affine merge mode.
[0136] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and the other from the upper neighboring CU.
[0137] Figure 9 The candidate blocks are shown in. For the left predictor, the scan order is A0 -> A1, and for the upper predictor, the scan order is B0 -> B1 -> B2. Only the first inherited candidate on each side is selected. No pruning check is performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vectors are used to derive the CPMVP candidates in the affine merge list of the current CU. As Figure 10 shown, if the neighboring lower-left block A is coded in an affine mode, the motion vectors v2, v3, and v4 at the upper-left, upper-right, and lower-left corners of the CU containing block A are obtained. When block A is coded with a 4-parameter affine model, two CPMVs of the current CU are calculated based on v2 and v3. If block A is coded with a 6-parameter affine model, three CPMVs of the current CU are calculated based on v2, v3, and v4.
[0138] Constructing the affine candidates means that the candidates are constructed by combining the neighboring translational motion information of each control point. The motion information of the control points is derived from Figure 11 the specified spatial and temporal neighborhoods shown. The CPMV k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV1, the B2 -> B3 -> A2 blocks are checked and the MV of the first available block is used. For CPMV2, the B1 -> B0 blocks are checked, and for CPMV3, the A1 -> A0 blocks are checked. If available, the TMVP is used as CPMV4.
[0139] After obtaining the MVs of the four control points, the affine merge candidates are constructed based on this motion information. The following combinations of control point MVs are used in order for construction:
[0140] {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}
[0141] A combination of 3 CPMVs constructs a 6-parameter affine merge candidate, and a combination of 2 CPMVs constructs a 4-parameter affine merge candidate. To avoid the motion scaling process, if the reference indices of the control points are different, the associated control point MV combinations are discarded.
[0142] After checking the inherited affine merge candidates and the constructed affine merge candidates, if the list is still not full, zero MVs are inserted at the end of the list.
[0143] 9. Decoder-side Motion Vector Refinement (DMVR)
[0144] To improve the accuracy of the merge mode MVs, decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. In the bi-predictive operation, a refined MV is searched around the initial MVs in the reference picture lists L0 and L1. The BM method calculates the distortion between two candidate blocks in the reference picture lists L0 and L1. Figure 12 An example of decoder-side motion vector refinement is shown. As Figure 12 shown, the SAD between the blocks calculated around the initial MV for each MV candidate is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi-predictive signal.
[0145] In VVC, the application of DMVR is restricted to CUs with the following modes and feature codings:
[0146] – CU-level merge mode with bi-predictive MVs
[0147] – For the current picture, one reference picture is in the past and the other reference picture is in the future
[0148] – The distances of the two reference pictures to the current picture (i.e., POC differences) are the same
[0149] – Both reference pictures are short-term reference pictures
[0150] – The CU has more than 64 luma samples
[0151] – Both the CU height and the CU width are greater than or equal to 8 luma samples
[0152] – The BCW weight index indicates equal weights
[0153] – WP is not enabled for the current block
[0154] – The CIIP mode is not used for the current block
[0155] The refined MVs derived from the DMVR process are used to generate inter - prediction samples and also for temporal motion - vector prediction in future - picture coding / decoding. The original MVs are used for the de - blocking process and also for spatial motion - vector prediction in future - CU coding / decoding.
[0156] Additional features of DMVR are mentioned in the following sub - clauses.
[0157] 10. Geometric Partitioning Mode (GPM)
[0158] In VVC, geometric partitioning mode is supported for inter - prediction. The geometric partitioning mode is represented using a CU - level flag as a merge mode, and other merge modes include the regular merge mode, MMVD mode, CIIP mode, and sub - block merge mode. The geometric partitioning mode supports a total of 64 partitions for each possible CU size w×h = 2 m ×2 n excluding 8x64 and 64x8, where m,n∈{3…6}.
[0159] Figure 13 shows an example of geometric - partitioning - mode (GPM) partitions grouped by the same angle. When using this mode, a CU is divided into two parts by a geometrically - positioned line ( Figure 13 ). The position of the dividing line is mathematically derived from the angle and offset parameters of a specific partition. Each geometric - partition part in the CU uses its own motion for inter - prediction; each partition allows only uni - directional prediction, i.e., each part has one motion vector and one reference index. A uni - directional - prediction motion constraint is applied to ensure that, like traditional bi - directional prediction, each CU requires only two motion - compensated predictions. The uni - directional prediction motion for each partition is derived using the process described in 3.4.11.1.
[0160] If the current coding / decoding unit uses the geometric partitioning mode, then a geometric - partition index indicating the geometric - partitioning (angle and offset) mode, and two merge indices (one for each partition) are further sent. The number of maximum GPM candidate sizes is explicitly signaled in the SPS, and syntax binarization is specified for the GPM merge indices. After predicting each part of the geometric partition, the sample values along the geometric - partition edge are adjusted using a blending process with adaptive weights, as shown in 3.4.11.2. This is the prediction signal for the entire coding / decoding unit, and the transform and quantization processes are applied to the entire coding / decoding unit as shown in other prediction modes. Finally, the motion field of the coding / decoding unit predicted using the geometric partitioning mode is stored, as shown in 3.4.11.3.
[0161] 11. Combined Inter - and Intra - Prediction (CIIP)
[0162] Figure 14Displays the top and left neighboring blocks used in the combined inter-frame intra prediction (CIIP) weight derivation.
[0163] In VVC, when a coding unit is coded in merge mode, if the coding unit contains at least 64 luma samples (i.e., the coding unit width multiplied by the coding unit height is equal to or greater than 64), and both the width and height of the coding unit are less than 128 luma samples, an additional flag is sent to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current coding unit. As its name implies, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P inter in the CIIP mode is derived using the same inter prediction process applied to the regular merge mode; the intra prediction signal P intra is derived according to the planar mode of the regular intra prediction process. Then, the intra and inter prediction signals are combined using weighted averaging, where the weight values are calculated based on the coding modes of the top and left neighboring blocks (as Figure 16 shown):
[0164] – If the top neighbor is available and coded intra, set isIntraTop to 1, otherwise set isIntraTop to 0;
[0165] – If the left neighbor is available and coded intra, set isIntraLeft to 1, otherwise set isIntraLeft to 0;
[0166] – If (isIntraLeft + isIntraTop) equals 2, set wt to 3;
[0167] – Otherwise, if (isIntraLeft + isIntraTop) equals 1, set wt to 2;
[0168] – Otherwise, set wt to 1.
[0169] The CIIP prediction is formed as follows:
[0170] P CIIP = ((4 - wt) * P inter + wt * P intra + 2) >> 2 (3 - 43)
[0171] 12. Template matching (TM)
[0172] Figure 15 Displays the template matching performed in the search area around the initial motion vector (MV).
[0173] Template matching (TM) is a decoder-side MV derivation method used to refine the motion information of the current coding unit by finding the closest match between a template of the current coding unit (i.e., the top and / or left neighboring blocks of the current coding unit) in the current image and a block in the reference image (i.e., of the same size as the template). As Figure 15 shown, a better MV is searched within the [–8, +8]-pel search range around the initial motion of the current coding unit. The template matching method in JVET-J0021 uses the following modifications: the search step size is determined based on the AMVR mode, and TM can be cascaded with the bilateral matching process in the merge mode.
[0174] In the AMVP mode, the MVP candidate is determined based on the template matching error, and the one with the smallest difference between the current block template and the reference block template is selected. Then, only this specific MVP candidate is subjected to TM for MV refinement. TM starts from the full-pel MVD accuracy (or 4 pixels for the 4-pel AMVR mode) and refines this MVP candidate using iterative diamond search within the [–8, +8] pixel search range. Depending on the AMVR mode, the MVP candidate may be further refined by cross search using the full-pel MVD accuracy (or 4 pixels for the 4-pel AMVR mode), and then followed by half-pixel and quarter-pixel refinements in sequence, as shown in Table 3. This search process ensures that the MVP candidate still maintains the same MV accuracy indicated by the AMVR mode after the TM process. During the search process, if the difference between the previous minimum cost and the current minimum cost within an iteration is less than the threshold (equal to the block area), the search process terminates.
[0175] Table 3. Search patterns for AMVR and merge mode with AMVR
[0176]
[0177]
[0178] In the merge mode, a similar search method is applied to the merge candidates indicated by the merge index. As shown in the above table, the template matching (TM) process can be performed to 1 / 8 pixel MVD accuracy, or skip the part beyond half-pixel MVD accuracy, depending on whether an alternative interpolation filter is used based on the merge motion information (in the case where AMVR operates in half-pixel mode). Additionally, when the TM mode is enabled, template matching can be an independent process or an additional MV refinement process between the block-based and sub-block-based bilateral matching procedures, depending on whether BM can be enabled according to its enabling conditions.
[0179] 13. Multi-pass decoder-side motion vector refinement
[0180] Apply multi-stage decoder-side motion vector refinement. In the first stage, bilateral matching (BM) is applied to the coded block. In the second stage, BM is applied to each 16x16 sub-block within the coded block. In the third stage, the MV in each 8x8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for spatial and temporal motion vector prediction.
[0181] (1) First stage - Block-based bilateral matching MV refinement
[0182] In this first stage, the refined MV is derived by applying the BM process to the coded block. Similar to the decoder-side motion vector refinement (DMVR) in VVC, in the bidirectional prediction operation, the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) around the initial MVs are derived based on the minimum bilateral matching cost between the two reference blocks in L0 and L1.
[0183] The BM process performs a local search to derive the integer sample precision intDeltaMV. The local search applies a 3×3 square search pattern, cycling the search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction. The values of sHor and sVer are determined by the size of the block, and their maximum value is 8.
[0184] The bilateral matching cost is calculated as
[0185] bilCost = mvDistanceCost + sadCost.
[0186] When the block size (cbW*cbH) is greater than 64, the Mean-Removed SAD (MRSAD) cost function is applied to eliminate the DC influence of the distortion between the reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the local search for intDeltaMV terminates. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern and continues to search for the minimum cost until the end of the search range is reached.
[0187] Existing fractional sample refinement is further applied to derive the final deltaMV. The refined MVs from the first stage are then represented as:
[0188] ● MV0_pass1 = MV0 + deltaMV
[0189] · MV1_pass1 = MV1 – deltaMV
[0190] (2) Second stage - sub - block - based bilateral matching MV refinement
[0191] In the second stage, refined MVs are derived by applying BM to 16×16 grid sub - blocks. For each sub - block, a refined MV is searched around two MVs (MV0_pass1 and MV1_pass1) in the reference picture lists L0 and L1, and these two MVs come from the first stage. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between two reference sub - blocks in L0 and L1.
[0192] For each sub - block, the BM process performs a full search to derive the integer - sample accuracy intDeltaMV. The full search has a search range of [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction. The values of sHor and sVer are determined by the size of the block, and their maximum value is 8.
[0193] The bilateral matching cost is calculated by applying a cost factor to the sum of absolute transform differences (SATD) cost between two reference sub - blocks: bilCost = satdCost * costFactor. Figure 16 A diamond - shaped area in the search region is shown. The search region (2*sHor + 1)*(2*sVer + 1) is divided into Figure 16 the 5 diamond - shaped search regions as shown. Each search region is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed in the order starting from the center of the search region. In each region, starting from the upper - left corner, the search points are processed in raster - scan order until the lower - right corner of the region. When the minimum bilCost within the current search region is less than or equal to the threshold of sbW * sbH, the integer - pixel (int - pel) full search terminates; otherwise, the integer - pixel full search continues to the next search region until all search points are checked. Additionally, if the difference between the previous minimum cost and the current minimum cost in an iteration is less than the threshold (the threshold is equal to the block size), the search process terminates.
[0194] Existing VVC DMVR fractional - sample refinement is further applied to derive the final deltaMV(sbIdx2). The refined MVs in the second stage are as follows:
[0195] ·MV0_pass2(sbIdx2)=MV0_pass1 + deltaMV(sbIdx2)
[0196] ● MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2)
[0197] (3) Third stage - sub - block based bidirectional optical flow MV refinement
[0198] In the third stage, refined MVs are derived by applying BDOF to 8×8 grid sub - blocks. For each 8×8 sub - block, starting from the refined MV of the parent sub - block in the second stage, BDOF refinement is applied to derive scaled Vx and Vy without clipping. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped to be between - 32 and 32.
[0199] The refined MVs (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) generated by the third stage are as follows:
[0200] ● MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2)+bioMv
[0201] ● MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2)–bioMv
[0202] 14. Bidirectional matching AMVP - merge mode
[0203] The bidirectional predictor consists of an AMVP predictor in one direction and a merge predictor in the other direction. When the selected merge predictor and AMVP predictor satisfy the DMVR condition, this mode can be enabled to encode and decode the block, where there is at least one reference picture from the past and one reference picture from the future relative to the current picture, and the distances from the two reference pictures to the current picture are the same. Then, bilateral matching MV refinement is applied to the merge MV candidate and AMVP MVP as a starting point. Otherwise, if the template matching function is enabled, template matching MV refinement is applied to the merge predictor or AMVP predictor with a higher template matching cost.
[0204] The AMVP part of this mode is signaled in the form of a regular unidirectional AMVP, that is, the reference index and MVD are signaled, and it has a derived MVP index if template matching is used, or the MVP index is signaled when template matching is disabled.
[0205] For the AMVP direction LX, X can be 0 or 1, and the merged part in the other direction (1–LX) is implicitly derived by minimizing the bilateral matching cost between the AMVP predictor and the merged predictor, i.e., for a pair of AMVP and merged motion vectors. For each merge candidate in the merge candidate list with a motion vector in the other direction (1-LX), the bilateral matching cost is calculated using the merge candidate motion vector and the AMVP motion vector. The merge candidate with the minimum cost is selected. Starting from the selected merge candidate motion vector and the AMVP motion vector, bilateral matching refinement is applied to the coded block.
[0206] The third stage of the multi-stage DMVR, i.e., the 8x8 sub-PU BDOF refinement, is enabled for AMVP merge mode coded blocks.
[0207] This mode is indicated by a flag, and if this mode is enabled, the AMVP direction LX is further indicated by a flag.
[0208] When using the bilateral matching (BM) AMVP merge mode for the current block and template matching is enabled, the MVD is not signaled. A pair of additional AMVP merge MVPs is introduced. The merge candidate list is sorted in ascending order based on the BM cost. An index (0 or 1) is signaled to indicate which merge candidate in the sorted merge candidate list to use. When there is only one candidate in the merge candidate list, the pair of AMVP MVP and the merge MVP without bilateral matching MV refinement is filled.
[0209] II. Proposed Methods
[0210] In ECM 6.0, when the TM mode is enabled, template matching can be an independent process or an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be checked for its enabling conditions.
[0211] In the present invention, several methods are proposed to improve TM to reduce complexity or increase codec performance.
[0212] In the first embodiment of the present invention, when bilateral matching (BM) is enabled for a block / CU decoded in the TM mode and the minimum cost of block-based (or CU-based) BM is less than or equal to a threshold TH, the TM process is prohibited.
[0213] In the second embodiment of the present invention, when bilateral matching (BM) is enabled for a block / CU decoded in the TM mode, if the minimum cost of block-based (or CU-based) BM is less than or equal to the threshold TH, the search range of the TM process is set to a smaller range.
[0214] In the third embodiment of the present invention, when bilateral matching (BM) is enabled for blocks / CUs encoded / decoded in TM mode, sub-block-based BM is always performed regardless of the cost of block-based BM.
[0215] In the fourth embodiment of the present invention, when bilateral matching (BM) is enabled for blocks / CUs encoded / decoded in TM mode, if the cost of TM is less than or equal to a threshold TH, sub-block-based BM is prohibited.
[0216] In the fifth embodiment of the present invention, when bilateral matching (BM) is enabled for blocks / CUs encoded / decoded in TM mode, TM is first performed, and then block-based BM and sub-block-based BM are performed.
[0217] In the sixth embodiment of the present invention, when bilateral matching (BM) is enabled for blocks / CUs encoded / decoded in TM mode, TM is first performed, and then block-based BM and sub-block-based BM are performed. When the cost of TM is less than or equal to the threshold TH, the block-based and sub-block-based BM processes are prohibited.
[0218] In the seventh embodiment of the present invention, when bilateral matching (BM) is enabled for encoded blocks / CUs using TM mode, TM is first performed, and then block-based BM and sub-block-based BM are performed. When the cost of TM is less than or equal to the threshold TH, or the cost of block-based BM is less than or equal to the threshold, or the cost of the best block-based BM cannot be reduced to a threshold higher or lower than the cost of the initial block-based BM (MV inherited from TM), sub-block-based BM is prohibited.
[0219] In the eighth embodiment of the present invention, when bilateral matching (BM) is enabled for TM mode encoded blocks / CUs, TM is first performed, and then block-based BM and sub-block-based BM are performed. When the cost of block-based BM is less than or equal to the threshold, or the cost of the best block-based BM cannot be reduced to a threshold higher than or equal to the cost of the initial block-based BM (MV inherited from TM), MV modification from BM is not used.
[0220] In the above method, TH can be any non-negative integer. In another scheme, TH can be adaptively adjusted according to encoding / decoding information. For example, TH can be related to the number of samples in the current block / CU, the inter-dir, the reference image and the PoC of the current image, the QP of the reference image and the current image, and the number of samples of the template.
[0221] Any of the proposed methods described above can be implemented in an encoder and / or a decoder. For example, any of the proposed methods can be implemented in the inter-frame prediction module of an encoder and / or a decoder. Alternatively, any of the proposed methods can be implemented as a circuit connected to the inter-frame prediction module of an encoder and / or a decoder.
[0222] Those skilled in the art will also understand that many variations can be made while implementing the above technical operations, while still achieving the same objectives of the present disclosure. Such variations are intended to be covered by the scope of the present disclosure. Therefore, the above description of the embodiments of the present disclosure is not intended to be restrictive. Instead, any limitations on the embodiments of the present disclosure are set forth in the following claims.
Claims
1. A method for performing inter - frame prediction in a video decoder, comprising: Receiving a coding - decoding unit in a video bit - stream, the coding - decoding unit being encoded using a template - matching process and a bilateral - matching process; Determining an order of the template - matching process and the bilateral - matching process; And Performing inter - frame prediction based on the determined order of the template - matching process and the bilateral - matching process to reconstruct the received coding - decoding unit.
2. The method according to claim 1, wherein the determining step further comprises: Determining a predefined order of the template - matching process and the bilateral - matching process as the determined order of the template - matching process and the bilateral - matching process, or Determining the order of the template - matching process and the bilateral - matching process based on a syntax element included in the bit - stream as the determined order of the template - matching process and the bilateral - matching process.
3. The method according to claim 1, wherein the bilateral - matching process comprises, in sequence, block - level motion - vector refinement, sub - block - level motion - vector refinement, and sub - block - level bidirectional optical - flow motion - vector refinement, wherein whether to perform subsequent refinement depends on the cost of the previous refinement, and When the template - matching process is performed, the template - matching process is performed immediately after the block - level motion - vector refinement of the bilateral - matching process.
4. The method according to claim 3, wherein the template - matching process is prohibited in response to the minimum cost of the block - level motion - vector refinement being less than or equal to a threshold.
5. The method according to claim 4, wherein the threshold is a non - negative integer, and the value of the threshold is predefined or adjustable based on the coding - decoding unit.
6. The method according to claim 3, wherein the template - matching process is performed in response to the minimum cost of the block - level motion - vector refinement being less than or equal to a first threshold, and its search range is less than a second threshold.
7. The method according to claim 3, wherein the sub - block - level motion - vector refinement is performed after the template - matching process, regardless of the cost of the block - level motion - vector refinement.
8. The method according to claim 3, wherein the sub - block - level motion - vector refinement is prohibited in response to the cost of the template - matching process being less than or equal to a threshold.
9. The method according to claim 2, wherein the bilateral - matching process comprises, in sequence, block - level motion - vector refinement, sub - block - level motion - vector refinement, and sub - block - level bidirectional optical - flow motion - vector refinement, wherein whether to perform subsequent refinement depends on the cost of the previous refinement, and When the bilateral - matching process is performed, it is performed after the template - matching process.
10. The method according to claim 9, wherein the bilateral - matching process is prohibited in response to the cost of the template - matching process being less than or equal to a threshold.
11. The method according to claim 9, wherein the sub - block - level motion - vector refinement is prohibited in response to the cost of the template - matching process being less than or equal to a third threshold, or the cost of the block - level motion - vector refinement being less than or equal to a fourth threshold, or the cost reduction of the best block - level motion - vector refinement not exceeding the cost reduction of the initial block - level motion - vector refinement.
12. The method according to claim 9, wherein in response to the cost of the block-level motion vector refinement being less than or equal to a sixth threshold, or the cost reduction of the optimal block-level motion vector refinement not exceeding the cost reduction of the initial block-level motion vector refinement, the use of the motion vector modification from the bilateral matching process is prohibited.
13. A method for performing inter prediction in a video encoder, comprising: performing inter prediction to encode a coding unit based on a determined order of a template matching process and a bilateral matching process; and transmitting the encoded coding unit in a video bitstream.
14. The method according to claim 13, wherein the performing step further comprises determining a predefined order of the template matching process and the bilateral matching process as the determined order of the template matching process and the bilateral matching process, or the performing step further comprises determining an order of the template matching process and the bilateral matching process, and the transmitting step further comprises transmitting a syntax element in the bitstream for indicating the determined order of the template matching process and the bilateral matching process.
15. The method according to claim 13, wherein the bilateral matching process comprises block-level motion vector refinement, sub-block-level motion vector refinement, and sub-block-level bidirectional optical flow motion vector refinement in sequence, and whether to perform subsequent refinement depends on the cost of the previous refinement, and when the template matching process is performed, the template matching process is performed immediately after the block-level motion vector refinement of the bilateral matching process.
16. The method according to claim 15, wherein in response to the minimum cost of the block-level motion vector refinement being less than or equal to a threshold, the template matching process is prohibited.
17. The method according to claim 16, wherein the threshold is a non-negative integer, and the value of the threshold is predefined or adjustable based on the coding unit.
18. The method according to claim 15, wherein in response to the minimum cost of the block-level motion vector refinement being less than or equal to a first threshold, the template matching process is performed, and its search range is less than a second threshold.
19. The method according to claim 15, wherein the sub-block-level motion vector refinement is performed after the template matching process, regardless of the cost of the block-level motion vector refinement.
20. The method according to claim 15, wherein in response to the cost of the template matching process being less than or equal to a threshold, the sub-block-level motion vector refinement is prohibited.
21. The method according to claim 13, wherein the bilateral matching process comprises block-level motion vector refinement, sub-block-level motion vector refinement, and sub-block-level bidirectional optical flow motion vector refinement in sequence, and whether to perform subsequent refinement depends on the cost of the previous refinement, and when the bilateral matching process is performed, it is performed after the template matching process.
22. The method according to claim 21, wherein in response to the cost of the template matching process being less than or equal to a threshold, the bilateral matching process is prohibited.
23. The method according to claim 21, wherein the sub-block-level motion vector refinement is prohibited when one of the following conditions is met: the cost of the template matching process is less than or equal to a third threshold, The cost of the block-level motion vector refinement is less than or equal to a fourth threshold, or The cost reduction of the optimal block-level motion vector refinement cannot exceed the cost reduction of the initial block-level motion vector refinement.
24. The method according to claim 21, wherein in response to the cost of the block-level motion vector refinement being less than or equal to a sixth threshold or the cost reduction of the optimal block-level motion vector refinement not exceeding the cost reduction of the initial block-level motion vector refinement, the use of motion vector modification from the bilateral matching process is prohibited.