Improved temporal merge candidates in a merge candidate list in video coding - Patents.com
Patent Information
- Application Number
- JP2024515043
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-29
- Filing Date
- 2022-09-28
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2042-09-28
AI Technical Summary
The existing video coding standards, such as VVC and ECM, do not adequately improve the performance of temporal motion vector prediction (TMVP) techniques, leading to redundancy and inefficiencies in motion estimation.
Implement enhanced temporal motion vector prediction candidate selection methods in VVC standard encoders and decoders, utilizing colocated CTU repositioning, expanded selection ranges, unconditional derivation of scaled motion vectors, and refined motion information processing to optimize TMVP performance.
Improves the efficiency and competitiveness of TMVP by reducing redundancy and enhancing motion estimation accuracy, leading to better compression performance and reduced bit rates in video coding.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Related Applications
[0001] This PCT application claims priority to U.S. patent application Ser. No. 63 / 250,208, filed Sep. 29, 2021, which is incorporated by reference herein. [Background technology]
[0002]
[0002] In 2020, the Joint Video Experts Team ("JVET") of the ITU-T Video Coding Experts Group ("ITU-T VCEG") and the ISO / IEC Moving Picture Experts Group ("ISO / IEC MPEG") published the final draft of the next-generation video codec specification, Versatile Video Coding ("VVC"). The specification further improves video coding performance over previous standards, such as H.264 / AVC (Advanced Video Coding) and H.265 / HEVC (High Efficiency Video Coding). JVET continues to propose additional techniques beyond the scope of the VVC standard itself, collected under the name of the Enhanced Compression Model ("ECM").
[0003]
[0003] In each of the AVC, HEVC, and VVC standards, motion compensated prediction ("MCP") is implemented as another central image compression technique, alongside discrete cosine transform ("DCT"). Images are partitioned into coding blocks, and MCP improves image compression efficiency based on the principle that motion in a block of a picture tends to repeat in adjacent blocks, as well as in blocks of temporally preceding and succeeding pictures. MCP is implemented by searching for such "motion candidates" in motion prediction and deriving motion information therefrom to reconstruct a block.
[0004]
[0004] Previous standards such as HEVC implemented MCP based only on translational motion, but VVC also implements affine motion compensation prediction ("affine MCP"). In general, the motion information of spatially and temporally neighboring blocks is reconstructed based on different formats of motion vectors, and in particular, the temporal motion vector is derived according to a technique based on temporal motion vector prediction ("TMVP"). In these manners, redundant motion information is reduced in the coded image, which reduces the bit rate required to transmit the video stream, thus achieving a rate gain. Summary of the Invention
[0005]
[0005] The first draft of the ECM (presented as "Exploration experiment on enhanced compression beyond VVC capability" at the 133rd Meeting of the Moving Picture Experts Group ("MPEG") in January 2021) includes a proposal to further expand the range of motion candidates explored according to the MCP technique of VVC. However, according to both the VVC implementation of MCP and the ECM implementation of MCP, the TMVP technique remains substantially unchanged from the HEVC implementation of MCP. Because TMVP remains an integral component of MCP, it is desirable to further improve the performance of TMVP so that TMVP is not redundant with respect to other motion vector prediction techniques.
[0006]
[0006] The detailed description is set forth with reference to the accompanying drawings, in which the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. Use of the same reference number in different figures indicates similar or equivalent items or features. [Brief description of the drawings]
[0007] [Figure 1A] 1 is an exemplary block diagram of a video encoding process according to an exemplary embodiment of the present disclosure. [Figure 1B] 4 is an example block diagram of a video decoding process according to an example embodiment of the present disclosure. [Diagram 2]
[0008] A diagram showing multiple spatially neighboring blocks of a current CU of a picture. [Diagram 3]
[0009] A diagram showing an exemplary selection of motion candidates for a CU of a picture according to motion predictive coding according to the VVC standard. [Figure 4]
[0010] FIG. 2 illustrates obtaining scaled motion vectors for temporal merge candidates according to the VVC standard. [Diagram 5]
[0011] FIG. 2 illustrates the selection of a position for a time candidate between candidates C0 and C1 according to the VVC standard. [Figure 6]
[0012] FIG. 1 illustrates possible spatial neighboring blocks from which not only adjacent but also non-adjacent spatial merging candidates may be derived according to ECM3. [Figure 7A]
[0013] FIG. 13 is a diagram showing a method of selecting a temporal motion vector prediction candidate by ECM. [Figure 7B] FIG. 1 illustrates a temporal motion vector prediction candidate selection method utilizing repositioning of a collocated CTU according to an exemplary embodiment of the present disclosure. [Figure 8]
[0014] FIG. 13 illustrates adding temporal merge candidates to a merge candidate list according to motion information of neighboring blocks, according to an exemplary embodiment of the present disclosure. [Figure 9]
[0015] FIG. 2 illustrates an example system for implementing the above-described processes and methods for implementing improved temporal motion candidate behavior. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008]
[0016] According to the VVC video coding standard ("VVC standard") and the motion prediction described therein, computer-readable instructions stored on a computer-readable storage medium are executable by one or more processors of a computing system to configure the one or more processors to perform the encoder operations described by the VVC standard and the decoder operations described by the VVC standard. Some of these encoder and decoder operations according to the VVC standard are described in further detail below, but these subsequent descriptions should not be understood as exhaustive of the encoder and decoder operations according to the VVC standard. Thereafter, a "VVC standard encoder" and a "VVC standard decoder" shall describe respective computer-readable instructions stored on a computer-readable storage medium that configure one or more processors to perform these respective operations (which may be referred to, by way of example, as a "reference implementation" of the encoder or decoder).
[0009]
[0017] Moreover, according to an exemplary embodiment of the present disclosure, the VVC standard encoder and the VVC standard decoder further include computer-readable instructions stored on a computer-readable storage medium that are executable by one or more processors of a computing system to configure the one or more processors to perform operations not specified by the VVC standard. The VVC standard encoder should not be understood as being limited to the operation of the reference implementation of the encoder, but should be understood as including further computer-readable instructions that configure one or more processors of a computing system to perform further operations described herein. The VVC standard decoder should not be understood as being limited to the operation of the reference implementation of the decoder, but should be understood as including further computer-readable instructions that configure one or more processors of a computing system to perform further operations described herein.
[0010]
[0018] 1A and 1B show example block diagrams of an encoding process 100 and a decoding process 150, respectively, according to an example embodiment of the present disclosure.
[0011]
[0019] In the encoding process 100, a VVC standard encoder configures one or more processors of a computing system to receive as input one or more input pictures from an image source 102. The input picture includes a number of pixels sampled by an image capture device, such as a photosensor array, and includes an uncompressed stream of multiple color channels (such as RGB color channels) that store color data at the original resolution of the picture, with each channel using a number of bits to store the color data for each pixel of the picture. The VVC standard encoder configures one or more processors of the computing system to store this uncompressed color data in a compressed format, where the color data is stored at a resolution lower than the original resolution of the picture and is encoded as a luma ("Y") channel and two chroma ("U" and "V") channels at a lower resolution than the luma channel.
[0012]
[0020] A VVC standard encoder encodes a picture (a picture being encoded, called a “current picture”, which is distinguished from other pictures received from the image source 102) by configuring one or more processors of the computing system to partition the original picture into units and sub-units according to a partition structure. The VVC standard encoder configures one or more processors of the computing system to subdivide the picture into macroblocks (“MBs”), each having dimensions of 16×16 pixels, and the MBs may be further subdivided into partitions. The VVC standard encoder configures one or more processors of the computing system to subdivide the picture into coding tree units (“CTUs”), and the luma and chroma components of the CTUs may be further subdivided into coding tree blocks (“CTBs”), which are further subdivided into coding units (“CUs”). Alternatively, the VVC standard encoder configures one or more processors of the computing system to subdivide the picture into units of N×N pixels, which may then be further subdivided into sub-units. Each of these largest subdivided units of a picture may be generally referred to as a "block" in this disclosure.
[0013]
[0021] A CU is coded using one block of luma samples and two corresponding blocks of chroma samples, and a picture is coded using one coding tree, rather than monochrome.
[0014]
[0022] A VVC standard encoder configures one or more processors of a computing system to subdivide a block into partitions having dimensions in multiples of 4 x 4 pixels. For example, a partition of a block may have dimensions of 8 x 4 pixels, 4 x 8 pixels, 8 x 8 pixels, 16 x 8 pixels, or 8 x 16 pixels.
[0015]
[0023] By encoding color information of blocks and block subdivisions of a picture rather than color information of the pixels of the original picture at full resolution, a VVC standard encoder configures one or more processors of a computing system to encode color information of a picture at a lower resolution than the input picture and to store the color information in fewer bits than the input picture.
[0016]
[0024] Furthermore, a VVC standard encoder encodes a picture by configuring one or more processors of a computing system to perform motion prediction on blocks of a current picture. Motion predictive coding refers to storing image data of blocks of a current picture (wherein blocks of the original picture before coding are called "input blocks") using motion information (not pixel data) and a prediction unit ("PU") by intra prediction 104 or inter prediction 106.
[0017]
[0025] Motion information refers to data describing the motion of a block structure or unit of a picture or a subunit thereof, such as motion vectors and references to blocks of the current picture or of a reference picture. A PU may refer to a unit or subunits corresponding to one of a plurality of block structures of a picture, such as an MB or a CTU, and blocks are partitioned based on picture data and coded according to the VVC standard. The motion information corresponding to a PU may describe motion prediction encoded by a VVC standard encoder described herein.
[0018]
[0026] A VVC standard encoder configures one or more processors of a computing system to code motion prediction information across each block of a picture in a coding order among the blocks, such as a raster scan order in which the first block decoded is the topmost and leftmost block of the picture. The block being coded is called the "current block" to distinguish it from other blocks of the same picture.
[0019]
[0027] According to intra prediction 104, one or more processors of the computing system are configured to encode a block with reference to motion information of one or more other blocks of the same picture and to the PU. According to intra predictive coding, one or more processors of the computing system perform intra prediction 104 calculation (also called spatial prediction) by coding the motion information of the current block based on spatially neighboring samples from the spatially neighboring blocks of the current block.
[0020]
[0028] According to inter prediction 106, one or more processors of the computing system are configured to encode a block with reference to motion information and PUs of one or more other pictures. In inter predictive coding, one or more processors of the computing system are configured to store one or more previously coded and decoded pictures in a reference picture buffer, where these stored pictures are called reference pictures.
[0021]
[0029] The one or more processors are configured to perform inter prediction 106 calculation (also called temporal prediction or motion compensated prediction) by coding the motion information of the current block based on samples from one or more reference pictures. Inter prediction can further be calculated according to uni-prediction or bi-prediction, i.e., in uni-prediction, only one motion vector, which points to one reference picture, is used to generate a prediction signal for the current block. In bi-prediction, two motion vectors, each of which points to a respective reference picture, are used to generate a prediction signal for the current block.
[0022]
[0030] The VVC standard encoder configures one or more processors of the computing system to code the CU to include a reference index for identifying the prediction signal(s) of the current block for reference to the VVC standard decoder. The one or more processors of the computing system can code the CU to include an inter-prediction indicator. The inter-prediction indicator indicates list 0 prediction with respect to a first reference picture list called list 0, list 1 prediction with respect to a second reference picture list called list 1, or bi-prediction with respect to both reference picture lists called list 0 and list 1, respectively.
[0023]
[0031] If the inter prediction indicator indicates list 0 prediction or list 1 prediction, the one or more processors of the computing system are configured to code the CU including a reference index pointing to a reference picture of a reference picture buffer referenced by list 0 or list 1, respectively. If the inter prediction indicator indicates bi-prediction, the one or more processors of the computing system are configured to code the CU including a first reference index pointing to a first reference picture of a reference picture buffer referenced by list 0 and a second reference index pointing to a second reference picture of the reference picture referenced by list 1.
[0024]
[0032] The VVC standard encoder configures one or more processors of the computing system to code current blocks of a picture individually and output a prediction block for each. According to the VVC standard, a CTU can be as large as 128x128 luma samples (and corresponding chroma samples depending on the chroma format). The CTU can be further partitioned into CUs according to a quad-tree, a binary tree, or a ternary tree. The one or more processors of the computing system are configured to finally record coding parameter sets, such as coding mode (intra mode or inter mode), motion information (reference index, motion vector, etc.) for inter-coded blocks, and quantized residual coefficients, in a syntax structure of a leaf node of the partitioning structure.
[0025]
[0033] After the predictive block is output, the VVC standard encoder configures one or more processors of the computing system to send a set of coding parameters, such as the coding mode (i.e., intra or inter prediction), the mode of intra prediction or the mode of inter prediction, and motion information, to an entropy coder 124 (described later).
[0026]
[0034] The VVC standard provides semantics for recording coding parameter sets for CUs. For example, with respect to the coding parameter sets described above, pred_mode_flag for the CU is set to 0 for inter-coded blocks and 1 for intra-coded blocks, general_merge_flag for the CU is set to indicate whether merge mode is used in inter prediction of the CU, inter_affine_flag and cu_affine_type_flag for the CU are set to indicate whether affine motion compensation is used in inter prediction of the CU, mvp_l0_flag and mvp_l1_flag are set to indicate motion vector index in list 0 or list 1, respectively, and ref_idx_l0 and ref_idx_l1 are set to indicate reference picture index in list 0 or list 1, respectively. It should be understood that the VVC standard includes semantics for recording various other information, flags, and options beyond the scope of this disclosure.
[0027]
[0035] The VVC standard encoder further implements one or more mode decisions and encoder control settings 108, including rate control settings. One or more processors of the computing system are configured to perform the mode decision by selecting an optimized prediction mode for the current block based on a rate-distortion optimization method after intra or inter prediction.
[0028]
[0036] The rate control settings configure one or more processors of a computing system to assign different quantization parameters ("QP") to different pictures. The magnitude of the QP determines the scale at which picture information is quantized during encoding by one or more processors (as described below), and thus the degree to which encoding process 100 discards picture information from MBs of a sequence during coding (by information falling between steps of the scale).
[0029]
[0037] The VVC standard encoder further implements a subtractor 110. One or more processors of the computing system are configured to perform the subtraction operation by calculating a difference between an input block and a predicted block. Based on the optimized prediction mode, the predicted block is subtracted from the input block. The difference between the input block and the predicted block is called a prediction residual, or "residual" for brevity.
[0030]
[0038] Based on the prediction residual, the VVC standard encoder further implements a transform 112. One or more processors of the computing system are configured to perform a transform operation on the residual by matrix arithmetic operations to derive an array of coefficients (which may be referred to as "residual coefficients", "transform coefficients", etc.), thereby encoding the current block as a transform block ("TB"). The transform coefficients may refer to coefficients that represent one of several spatial transformations, such as a diagonal flip, a vertical flip, or a rotation, which may be applied to the sub-blocks.
[0031]
[0039] It should be appreciated that the coefficients may be stored as two components, a magnitude and a sign, as will be explained in more detail below.
[0032]
[0040] The sub-blocks of a CU, such as PUs and TBs, may be arranged in any combination of sub-block dimensions as described above. A VVC standard encoder configures one or more processors of a computing system to subdivide a CU into a residual quadtree ("RQT"), which is a hierarchical structure of TBs. The RQT provides an order for motion prediction and residual coding across the sub-blocks of each level of the RQT, recursively descending each level of the RQT.
[0033]
[0041] The VVC standard encoder further implements quantization 114. One or more processors of the computing system are configured to perform a quantization operation on the residual coefficients by matrix arithmetic operations based on the quantization matrix and the QP assigned above. Residual coefficients that fall within an interval are kept, and residual coefficients that fall outside the interval step are discarded.
[0034]
[0042] The VVC standard encoder further implements inverse quantization 116 and inverse transform 118. One or more processors of the computing system are configured to perform inverse quantization and inverse transform operations on the quantized residual coefficients by matrix arithmetic operations that are the inverse of the quantization and transform operations described above. The inverse quantization and inverse transform operations result in a reconstructed residual.
[0035]
[0043] The VVC standard encoder further implements an adder 120. One or more processors of the computing system are configured to perform an addition operation by adding the prediction block and the reconstructed residual, and output a reconstructed block.
[0036]
[0044] The VVC standard encoder further implements a loop filter 122. One or more processors of the computing system are configured to apply loop filters, such as a deblocking filter, a sample adaptive offset ("SAO") filter, and an adaptive loop filter ("ALF"), to the reconstructed blocks and output filtered reconstructed blocks.
[0037]
[0045] The VVC standard encoder further configures one or more processors of the computing system to output the filtered reconstructed blocks to a decoded picture buffer ("DPB") 200. The DPB 200 stores reconstructed pictures used by the one or more processors of the computing system as reference pictures in coding pictures other than the current picture, as described above with respect to inter prediction.
[0038]
[0046] The VVC standard encoder further implements an entropy coder 124. The one or more processors of the computing system are configured to perform entropy coding, in which symbols constituting the quantized residual coefficients are coded by mapping to binary strings (hereinafter "bins"), which can be transmitted in an output bitstream at a compressed bit rate, according to a context-dependent binary arithmetic codec ("CABAC"). The coded symbols of the quantized residual coefficients include absolute values of the residual coefficients (these absolute values are hereinafter referred to as "residual coefficient levels").
[0039]
[0047] However, while the residual coefficient levels are predicted and coded, the residual coefficient signs are signaled using bins indicating equal probability (hereinafter "EP") states (it should be understood that coefficients with value 0 do not have signs and therefore do not need to be signaled). VVC standard encoders do not configure one or more processors of a computing system to predict the residual coefficient signs due to computational challenges (which will be understood by those skilled in the art, but need not be repeated here to understand the exemplary embodiments of this disclosure). For these reasons, CABAC configures one or more processors of a computing system to bypass the coding of the residual coefficient signs and to additionally transmit one bit per sign in the output bitstream.
[0040]
[0048] Thus, the entropy coder configures one or more processors of a computing system to code a residual coefficient level of a block, bypass coding of the residual coefficient code and record the residual coefficient code together with the coded block, record coding parameter sets, such as, for example, a coding mode, a mode of intra-prediction or inter-prediction, and motion information, that are coded within a syntax structure of the coded block (e.g., a picture parameter set ("PPS") included in a picture header and a sequence parameter set ("SPS") included in a sequence of multiple pictures), and output the coded block.
[0041]
[0049] The VVC standard encoder configures one or more processors of the computing system to output a coded picture that is made up of the coded blocks from the entropy coder 124. The coded picture is output to a transmission buffer and ultimately packed into a bitstream for output from the VVC standard encoder.
[0042]
[0050] In the decoding process 150, the VVC standard decoder configures one or more processors of a computing system to receive as input one or more coded pictures from a bitstream.
[0043]
[0051] The VVC standard decoder implements an entropy decoder 152. One or more processors of the computing system are configured to perform entropy decoding, and the bins are decoded by reversing the symbol-to-bin mapping according to CABAC, thereby recovering the entropy-coded quantized residual coefficients. The entropy decoder 152 outputs the quantized residual coefficients, outputs coding-bypassed residual coefficient codes, and outputs syntax structures, such as PPS and SPS.
[0044]
[0052] The VVC standard decoder further implements inverse quantization 154 and inverse transform 156. One or more processors of the computing system are configured to perform inverse quantization and inverse transform operations on the decoded quantized residual coefficients by matrix arithmetic operations that are the inverse of the quantization and transform operations described above. The inverse quantization and inverse transform operations result in a reconstructed residual.
[0045]
[0053] Furthermore, based on the coding parameter set recorded by the entropy coder 124 in syntax structures such as PPS and SPS (or received by out-of-band transmission or coded within the decoder) and the coding mode contained in the coding parameter set, the VVC standard decoder determines whether to apply intra prediction 156 (i.e., spatial prediction) or motion compensated prediction 158 (i.e., temporal prediction) to the reconstructed residual.
[0046]
[0054] If the coding parameter set specifies intra prediction, the VVC standard decoder configures one or more processors of the computing system to perform intra prediction 156 using the prediction information specified in the coding parameter set, thereby generating a prediction signal.
[0047]
[0055] If the coding parameter set specifies inter prediction, the VVC standard decoder configures one or more processors of the computing system to perform motion compensated prediction 158 using reference pictures from the DPB 200. The motion compensated prediction 158 thereby generates a prediction signal.
[0048]
[0056] The VVC standard decoder further implements an adder 160. The adder 160 configures one or more processors of the computing system to perform an addition operation on the reconstructed residual and the prediction signal, thereby outputting a reconstructed block.
[0049]
[0057] The VVC standard decoder further implements a loop filter 162. One or more processors of the computing system are configured to apply loop filters, such as a deblocking filter, an SAO filter, and an ALF, to the reconstructed blocks and output filtered reconstructed blocks.
[0050]
[0058] The VVC standard decoder further configures one or more processors of the computing system to output the filtered reconstructed blocks to the DBP 200. As described above, the DPB 200 stores reconstructed pictures used by the one or more processors of the computing system as reference pictures in coding pictures other than the current picture, as described above with respect to motion compensated prediction.
[0051]
[0059] The VVC standard decoder further configures one or more processors of the computing system to output the reconstructed picture from the DPB to a user-viewable display of the computing system, such as a television display, a personal computing monitor, a smartphone display, or a tablet display.
[0052]
[0060] Thus, as illustrated by the encoding process 100 and the decoding process 150 described above, the VVC standard encoder and the VVC standard decoder, respectively, implement motion predictive coding according to the VVC specification. The VVC standard encoder and the VVC standard decoder, respectively, configure one or more processors of a computing system to generate a reconstructed picture based on a previous reconstructed picture of the DPB according to the motion compensated prediction described by the VVC standard, where the previous reconstructed picture serves as a reference picture in the motion compensated prediction as described herein.
[0053]
[0061] As described above with respect to the coding parameter set, for a reconstructed picture coded by inter-predictive coding, the VVC standard encoder and the VVC standard decoder implement merge modes and affine motion compensation for inter-prediction of the reconstructed block. The VVC standard encoder and the VVC standard decoder implement multiple merge modes for inter-prediction of the motion information of the CU of the reconstructed picture, including motion compensation prediction ("MCP"), affine motion compensation prediction ("affine MCP"), and other merge modes, as specified by the VVC standard. The motion information may include multiple motion vectors.
[0054]
[0062] The motion information of the CU of the reconstructed picture may further include a motion candidate list. According to the VVC standard, the motion candidate list may be a data structure that contains references to multiple motion candidates. The motion candidates may be block structures, or subunits of block structures, such as pixels, or any other suitable subdivision of the block structure of the current picture, or may be references to motion candidates of another picture. The motion candidates may be spatial motion candidates or temporal motion candidates. By applying motion vector compensation ("MVC"), the VVC standard decoder may select a motion candidate from the motion candidate list and derive the motion vector of the motion candidate as the motion vector of the CU of the reconstructed picture.
[0055]
[0063] FIG. 3 shows an exemplary selection of motion candidates for a CU of a picture according to merge mode coding according to the VVC standard.
[0056]
[0064] According to the VVC standard, the motion candidate list may be a merge candidate list and may include up to five types of merge candidates (six according to the ECM, as described later). A VVC standard encoder may implement coding of the syntax structure of a CU to include a merge index.
[0057]
[0065] For each CU coded in merge mode, the index of the best merge candidate is coded using truncated unary binarization (TU).
[0058]
[0066] A merge candidate list for a CU of a picture coded according to the merge mode may include, in order, the following merge candidates:
[0059]
[0067] Spatial MVP candidates from CUs spatially neighboring to the current CU,
[0060]
[0068] Temporal MVP candidates ("TMVP candidates") from co-located CUs for the current CU;
[0061]
[0069] History-based MVP candidates from FIFO tables,
[0062]
[0070] Pairwise average MVP candidates, and
[0063]
[0071] Zero motion vector.
[0064]
[0072] As shown in Figure 2, there are multiple spatially neighboring blocks of a current CU of a picture. The spatially neighboring blocks of the current CU include a block near the left edge of the current CU and a block near the top edge of the current CU. The spatially neighboring blocks have left-right and top-bottom relationships to the current CU as shown in Figure 2. By way of example in Figure 2, a merge candidate list for a picture coded according to a merge mode may include up to the following merge candidates:
[0065]
[0073] A spatially adjacent block on the left (A0),
[0066]
[0074] A spatially adjacent block (B0) on the upper side,
[0067]
[0075] A spatially adjacent block (B1) in the upper right corner,
[0068]
[0076] A block (A1) that is spatially adjacent to the lower left side, and
[0069]
[0077] A spatially adjacent block (B2) on the upper left side.
[0070]
[0078] Of the spatially neighboring blocks shown herein, block A0 is the block to the left of the current CU, block A1 is the block to the left of the current CU, block B0 is the block above the current CU, block B1 is the block above the current CU, and block B2 is the block above the current CU. The relative positioning of the spatially neighboring blocks with respect to the current CU or with respect to each other shall not be limited beyond these relationships, and there shall be no limitations on the relative size of the spatially neighboring blocks with respect to the current CU or with respect to each other.
[0071]
[0079] The VVC standard encoder and decoder implement deriving at most four merge candidates from searching for spatially neighboring blocks on the left side of the current CU and searching for spatially neighboring blocks on the top side of the current CU. These spatially neighboring blocks may be searched in the order of B0, A0, B1, A1, and B2. Any of these spatially neighboring blocks may be available for the merge candidate list as long as they do not belong to another slice or tile. Therefore, B2 is only added to the merge candidate list if none of the other four spatially neighboring blocks are available or are intra-coded.
[0072]
[0080] For each spatially neighboring block found to be available, a merge candidate is derived from the motion of its spatially neighboring blocks and added to the merge candidate list. If further candidates are found after the A1 candidate is added in this way, the VVC standard encoder and the VVC standard decoder further implement performing a redundancy check. A candidate that contains the same motion information as another candidate should not be added to the list. However, in order to reduce the computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, as shown in Figure 3, only pairs linked by an arrow are considered, and a candidate is added to the list only if the corresponding candidate used for the redundancy check does not have the same motion information.
[0073]
[0081] Then, only one temporal merge candidate is added to the list. In particular, in the derivation of this temporal merge candidate, a scaled motion vector is derived based on the collocated CU belonging to the collocated reference picture. The VVC standard encoder implements explicit signaling of the reference picture list and reference index to be used for the derivation of the collocated CU in the slice header.
[0074]
[0082] It should be understood that the VVC standard defines a "co-located picture" as a picture that has the same spatial resolution, the same scaling window offset, the same number of sub-pictures, and the same CTU size as the current picture.
[0075]
[0083] 4 illustrates obtaining a scaled motion vector for a temporal merge candidate according to the VVC standard by a dotted line, where the scaled motion vector is scaled from the motion vector of the co-located CU using a picture order count ("POC") distance, tb and td, where tb indicates the POC difference between the reference picture of the current picture and the current picture, and td indicates the POC difference between the co-located picture and the reference picture of the co-located picture. The reference picture index of the temporal merge candidate is set equal to 0.
[0076]
[0084] It should be understood that when deriving a temporal merge candidate, the VVC standard encoder and VVC standard decoder implement deriving a scaled motion vector from one of the L0 motion vector and the L1 motion vector of the co-located CU, and one of the L0 motion vector and the L1 motion vector of the co-located CU is determined according to the following steps.
[0077]
[0085] If the motion vector of the co-located CU is a bi-predictive motion vector and the current picture is a low latency picture, the L0 motion vector of the TMVP candidate is scaled from the L0 motion vector of the co-located CU, and the L1 motion vector of the TMVP candidate is scaled from the L1 motion vector of the co-located CU.
[0078]
[0086] Otherwise, if the motion vector of the co-located CU is a bi-predictive motion vector and the current picture is a non-low latency picture, the VVC standard encoder and the VVC standard decoder implement determining one of the two motion vectors of the co-located CU as a basis for scaling according to the reference picture list of the co-located CU. More specifically, if the co-located CU is from the L0 reference picture list, the L0 motion vector and the L1 motion vector of the TMVP candidate are both scaled from the L1 motion vector of the co-located CU. Similarly, if the co-located CU is from the L1 reference picture list, the L0 motion vector and the L1 motion vector of the TMVP candidate are both scaled from the L0 motion vector of the co-located CU.
[0079]
[0087] Otherwise, if the motion vector of the co-located CU is an L0 predicted motion vector, then both the L0 and L1 motion vectors of the TMVP candidate are scaled from the L0 motion vector of the co-located CU, regardless of whether the current picture is a low latency picture. Similarly, if the motion vector of the co-located CU is an L1 predicted motion vector, then both the L0 and L1 motion vectors of the TMVP candidate are scaled from the L1 motion vector of the co-located CU.
[0080]
[0088] 5 illustrates the selection of a position for a temporal candidate between candidates C0 and C1, where the solid outlined block indicates the location of the current CU according to the VVC standard. If the co-located CU at position C0 is not available, is intra-coded, or is outside the current row of the CTU, the VVC standard encoder and decoder implement deriving a temporal merge candidate using the co-located CU at position C1. In other cases, the VVC standard encoder and decoder implement deriving a temporal merge candidate using the co-located CU at position C0.
[0081]
[0089] Therefore, it should be understood that according to the VVC standard, the temporal candidates are derived from either a co-located CU positioned relative to the lower right corner of the current CU, or a co-located CU positioned relative to the center of the current CU.
[0082]
[0090] Next, the VVC standard encoder and VVC standard decoder implement adding a history-based MVP ("HMVP") merge candidate after the spatial MVP candidate and the TMVP candidate in the merge candidate list. In this specification, the motion information of a previously coded block is stored in a table and used as an MVP candidate for the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (emptied) when a new CTU row is encountered. Whenever there is an inter-coded CU that is not a subunit, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0083]
[0091] The HMVP table size S is set to be 6, which indicates that up to five HMVP candidates can be added to the table. When inserting a new motion candidate into the table, the VVC standard encoder and VVC standard decoder implement a constrained first-in-first-out ("FIFO") process, and a redundancy check is first applied to find whether there is an identical HMVP in the table. If found, the identical HMVP is removed from the table, all subsequent HMVP candidates are moved forward, and the identical HMVP is inserted into the last entry of the table.
[0084]
[0092] HMVP candidates can be used in the merge candidate list construction process. The most recent few HMVP candidates in the table are checked in order and inserted in the candidate list after the TMVP candidate. Redundancy checks are applied to HMVP candidates for spatial or temporal merge candidates.
[0085]
[0093] To reduce the number of redundant check operations, the following simplifications are introduced.
[0086]
[0094] For the A1 and B1 space candidates, respectively, the last two entries in the table are redundancy checked.
[0087]
[0095] When the total number of available merge candidates reaches one less than the maximum allowed merge candidates, the merge candidate list construction process from HMVP is terminated.
[0088]
[0096] Then, the VVC standard encoder and the VVC standard decoder implement generating pair-wise average candidates by averaging predefined pairs of candidates in the existing merge candidate list using the first two merge candidates. The first merge candidate may be defined as p0Cand and the second merge candidate may be defined as p1Cand, respectively. The averaged motion vector is calculated according to the availability of the motion vectors of p0Cand and p1Cand separately for each reference list. If both motion vectors are available in one list, these two motion vectors are averaged even when they point to different reference pictures, and its reference picture is set to the motion vector of p0Cand. If only one motion vector is available, it uses that motion vector directly. If the motion vector is not available, it keeps this list invalid. Also, if the half-pel interpolation filter index of p0Cand and p1Cand is different, it is set to 0.
[0089]
[0097] Finally, if the merge list is not full after the pairwise average merge candidates are added, zero MVPs are inserted at the end positions until the maximum number of merge candidates is reached. A zero motion vector may have a motion shift of (0,0).
[0090]
[0098] While the VVC standard provides a merge candidate list of at most six candidates, JVET's continuing work in this area (presented as "Exploration experiment on enhanced compression beyond VVC capability" at the 133rd Meeting of the Moving Picture Experts Group ("MPEG") in January 2021, and as "Algorithm description of Enhanced Compression Model 3 (ECM3)" at the 136th Meeting of MPEG in October 2021) goes beyond the scope of the VVC standard and proposes an expanded merge candidate list of at most 15 candidates, including, in order:
[0091]
[0099] Spatial MVP candidates from CUs spatially neighboring to the current CU,
[0092]
[0100] Time MVP candidates from the current CU's co-located CUs,
[0093]
[0101] Non-adjacent spatial candidates,
[0094]
[0102] History-based MVP candidates from FIFO tables,
[0095]
[0103] Pairwise average MVP candidates, and
[0096]
[0104] Zero motion vector.
[0097]
[0105] Figure 6 shows possible spatial neighboring blocks according to ECM3, from which adjacent as well as non-adjacent spatial merge candidates can be derived. Non-adjacent spatial merge candidates are usually inserted after the TMVP candidate in the merge candidate list. The distance between the non-adjacent spatial candidate and the current coding block is based on the width and height of the current coding block. No line buffer limitations are applied.
[0098]
[0106] Furthermore, after the merge candidate list is constructed, the merge candidates are sorted (according to adaptive sorting of merge candidates, hereinafter referred to as "ARMC"). The merge candidates are first divided into several subgroups. The subgroup size is set to 5 for normal merge mode and TM merge mode. The subgroup size is set to 3 for affine merge mode. The merge candidates in each subgroup are sorted in ascending order according to the cost value based on template matching. For simplicity, the merge candidates in the last subgroup but not the first subgroup are not sorted. The template matching cost of the merge candidates is measured by the sum of absolute differences ("SAD") between the samples of the template of the current block and their corresponding reference samples. The template includes a set of reconstructed samples in the neighborhood of the current block. The reference samples of the template are located by the motion information of the merge candidates.
[0099]
[0107] Although the merge candidate list search technique of ECM3 focuses on expanding the scope of merge candidate search, neither the VVC standard nor the ECM proposal has the improved performance of the TMVP technique, since the TMVP candidate continues to occupy only one position in the merge candidate list. Therefore, the TMVP candidate is increasingly likely to underperform compared to other merge candidates. It is desirable to improve the TMVP performance so that it remains competitive with merge candidates based on other motion prediction techniques in the merge candidate list.
[0100]
[0108] Thus, the exemplary embodiments of this disclosure provide a temporal motion vector prediction candidate selection method that offers improvements over VVC and ECM in several respects.
[0101]
[0109] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a temporal motion vector prediction candidate selection method that utilizes repositioning of a collocated CTU.
[0102]
[0110] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a temporal motion vector prediction candidate selection method that takes advantage of an expanded selection range.
[0103]
[0111] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a temporal motion vector prediction candidate selection method that utilizes an unconditional derivation of a scaled motion vector.
[0104]
[0112] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a temporal motion vector prediction candidate selection method that omits scaling uni-predictive motion vectors to bi-predictive motion vectors.
[0105]
[0113] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a temporal motion vector prediction candidate selection method that utilizes multiple options in setting reference picture indexes.
[0106]
[0114] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a temporal motion vector prediction candidate selection method that utilizes scaling factor offsetting.
[0107]
[0115] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a merge candidate list creation method that omits temporal motion vector prediction candidates.
[0108]
[0116] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a picture reconstruction method that utilizes motion information refinement.
[0109]
[0117] Each of the above aspects of exemplary embodiments of the present disclosure are described in further detail below.
[0110]
[0118] According to the ECM design, in order to minimize the on-chip buffer size of temporal motion, the temporal motion vector may be obtained only from the co-located CTU and one column located to the right of the co-located CTU, where the co-located CTU is a CTU in a co-located reference picture whose position is the same as that of the current CTU. However, this design is not suitable for sequences with fast motion or for pictures whose co-located reference pictures are far away (i.e., the POC distance between a picture and its co-located reference picture is large). Therefore, the exemplary embodiment of the present disclosure provides a VVC standard encoder and a VVC standard decoder that allow the temporal motion vector to be derived from a position other than the co-located CTU.
[0111]
[0119] 7A and 7B respectively show a temporal motion vector prediction candidate selection method by ECM and a temporal motion vector prediction candidate selection method utilizing repositioning of a collocated CTU according to an exemplary embodiment of the present disclosure. First, block partitioning of a picture effectively divides the picture into a grid of blocks, where some grids of the picture may have larger block sizes and other grids have smaller block sizes. For each grid of a picture, a motion vector is signaled to indicate where the temporal motion of the grid comes from, and thus such a grid is referred to as a "motion grid" for brevity, since the grid granularity determines the distribution of the motion vectors.
[0112]
[0120] 7A and 7B show an example where the motion grid size of the current block and the motion grid size of the co-located block are equal to the size of the CTU, the left diagram shows the ECM proposal, and the right diagram shows the present disclosure. According to an exemplary embodiment of the present disclosure, the VVC standard encoder and the VVC standard decoder implement changing the position of the co-located CTU according to the signaled motion vector 702. In other words, the VVC standard encoder and the VVC standard decoder derive the TMVP candidate of the current CU of the current CTU 704 of the current picture 706 from the repositioned co-located CTU 708 of the co-located picture 710 according to the motion vector 702 (which may be signaled, but not necessarily, as will be described later), and the repositioned co-located CTU 708 is positioned in the co-located picture 710 by the motion vector 702 relative to the current CTU 704 in the current picture 706.
[0113]
[0121] It should be understood that "changing the position" of the co-located CTU 708 or "repositioning" the co-located CTU 708 according to an example embodiment of the present disclosure does not mean that the picture data of the co-located picture 710 is moved. Instead, "changing the position" or "repositioning" the co-located CTU 708 should be understood as a VVC standard encoder and a VVC standard decoder performing the operations described herein when using the co-located CTU 708. The co-located CTU 708 does not need to be positioned in the co-located picture 710 in the same way as the current CTU 704 in the current picture 706; instead, another CTU of the co-located picture 710 (which need not be, but may be positioned differently with respect to the current CTU 704) is used in place of the co-located CTU 708 when performing the above operations.
[0114]
[0122] Alternatively, the motion grid size of the current block and the motion grid size of the co-located block may differ from the size of the CTU. In one example, the motion grid size may be N×N, where N is equal to 256, 128, 64, 32, or 16 luma samples. In another example, the motion grid size may be N×M, where N is not equal to M, and where N and M are both integer powers of two.
[0115]
[0123] The motion grid size may vary depending on the sequence, temporal layer or picture type difference, resulting in different block partitions. According to some exemplary embodiments, the VVC standard encoder and VVC standard decoder implement signaling the motion grid size of the current block in a sequence level, picture level or slice level syntax structure. According to other exemplary embodiments, the motion grid size is adjusted according to the temporal layer.
[0116]
[0124] To understand this disclosure, it should be understood that when a VVC standard encoder and a VVC standard decoder implement signaling parameters in a syntax structure, the encoder implements recording the parameters in a syntax structure such as a block, picture, sequence, slice, etc., and transmitting the coded syntax structure on the bitstream, and the encoder implements parsing the coded syntax structure from the bitstream.
[0117]
[0125] The VVC standard encoder and decoder may implement signaling the motion grid size of the highest temporal layer in the sequence level syntax structure, or may implement signaling a fixed motion grid size of the highest temporal layer, such as the same size as the CTU. In that case, the VVC standard encoder and decoder may implement reducing the motion grid size for any lower temporal layer, because in the lower temporal layer, the POC distance between a picture and its collocated reference picture is large. Thus, in the lower temporal layers, the motion is more complex and finer than in the higher temporal layers, and thus the accuracy of temporal motion prediction can be improved with smaller granularity.
[0118]
[0126] Furthermore, VVC standard encoders and decoders can implement signaling the motion grid size of the lowest temporal layer in a syntax structure and increasing the motion grid size for every higher temporal layer.
[0119]
[0127] To signal the motion vector for each motion grid, the motion vector may be directly signaled or predicted. According to some exemplary embodiments, the motion vector of the current motion grid may be merged from any of its neighboring motion grids (e.g., the motion grid on the left or top neighbor according to raster scan order coding, or any other neighboring motion grid of the block previously coded according to other scan orders, as described below). VVC standard encoders and VVC standard decoders implement signaling a parameter (e.g., a flag or index) in the syntax structure to indicate whether the motion vector of the current motion grid is the same as the motion vector of the neighboring motion grid. If the signaled parameter indicates the same, the motion vector of the current motion grid is not signaled and directly inherits the motion grid of the neighboring motion grid. Otherwise, the motion vector of the current motion grid is signaled in the syntax structure.
[0120]
[0128] Alternatively, the motion vector of the current motion grid can be predicted from its neighboring motion grids. The VVC standard encoder and VVC standard decoder implement using the motion vector of the neighboring motion grid as the motion vector predictor of the current motion grid. Instead of signaling the parameters described above, only the motion vector differential is signaled in the syntax structure.
[0121]
[0129] Furthermore, if a motion vector for each motion grid is signaled in the syntax structure, the VVC standard encoder and VVC standard decoder implement coding of the motion grid according to a default order, where the default order may be one of a raster scan order, a z-order scan order, a horizontal scan order, a vertical scan order, and a diagonal scan order.
[0122]
[0130] It should be understood that the co-located CTU may be in the best position for the temporal motion vector so that the signaled motion vector is equal to zero motion. Therefore, to reduce signaling overhead, VVC standard encoders and VVC standard decoders implement signaling in a syntax structure a control parameter that determines whether the position of the co-located CTU is repositioned from the current CTU.
[0123]
[0131] According to another exemplary embodiment, a VVC standard encoder and a VVC standard decoder implement signaling a sequence level, picture level, or slice level syntax structure flag to indicate whether the position of a co-located CTU is repositioned from a current CTU.
[0124]
[0132] According to another example embodiment, the position of the co-located CTU of the higher temporal layer remains unchanged from the current CTU and no other parameters need to be signaled.
[0125]
[0133] In one or more aspects, example embodiments of the present disclosure provide a temporal motion vector prediction candidate selection method that takes advantage of an expanded selection range.
[0126]
[0134] According to the VVC and ECM described above with reference to Figure 5, the locations for the time candidates may be selected only from the candidates C0 and C1 shown in Figure 5. Thus, the exemplary embodiment of the present disclosure provides for the selection of the time candidates from additional locations.
[0127]
[0135] According to some exemplary embodiments, a VVC standard encoder and a VVC standard decoder implement selection of temporal candidates from a set combining C0 and C1 ("temporal candidates") shown in FIG. 5 with A0, A1, B0, B1, and B2 ("spatial candidates") shown in FIG. 2 (i.e., blocks that are spatially nearby as described with reference to FIG. 2, relative to the position of the current CU shown in FIG. 5).
[0128]
[0136] According to some exemplary embodiments, the VVC standard encoder and the VVC standard decoder implement selecting temporal candidates according to a default order. The default order is C0, then C1, then B0, then A0, then A1, then B1, then B2. If the CU at position C0 is not available, is intra-coded, or is outside the current row of the CTU, then position C1 is examined. Otherwise, position C0 is used in deriving the TMVP candidate, and the search ends. Similarly, if the CU at position C1 is not available, is intra-coded, or is outside the current row of the CTU, then position B0 is examined, and so on. The default order according to the exemplary embodiments of the present disclosure is not limited and may be any combination of C0, C1, A0, A1, B0, B1, B2.
[0129]
[0137] According to another exemplary embodiment, a VVC standard encoder and a VVC standard decoder implement deriving TMVP candidates by averaging all the temporal motion vectors of temporal candidates, which are normalized by scaling to a fixed reference picture, and then the normalized temporal motion vectors are averaged.
[0130]
[0138] According to another exemplary embodiment, a VVC standard encoder and a VVC standard decoder implement deriving TMVP candidates by comparing the temporal motion vectors of temporal candidates with spatial merge candidates, and the temporal motion vector that results in the largest motion vector difference is used in deriving the TMVP candidate.
[0131]
[0139] According to another exemplary embodiment, the VVC standard encoder and the VVC standard decoder implement selecting TMVP candidates from the temporal motion vectors of the temporal candidates according to the respective cost values of template matching. The template matching cost of the temporal candidates is measured by the SAD between the samples of the template of the current block and their corresponding reference samples. The template includes a set of reconstructed samples in the neighborhood of the current block. The reference samples of the template are located by the motion information of the temporal candidates.
[0132]
[0140] According to another exemplary embodiment, instead of indirectly determining from where the TMVP candidates are selected, the derivation of the TMVP candidates may be explicitly signaled. The VVC standard encoder and decoder implement obtaining the temporal candidates from a set including {C0, C1, A0, A1, B0, B1, B2}. The temporal candidates are {C0, C1, A0, A1, B0, B1, B2}. 1, B2}. The time candidates may also be taken from any position within the collocated CTU.
[0133]
[0141] VVC standard encoders and decoders implement signaling an index in the syntax structure of each CTU to identify the derivation of a TMVP candidate. For example, a signaled index of 0 identifies that the TMVP candidate is derived from the C0 position, a signaled index of 1 identifies that the TMVP candidate is derived from the C1 position, and so on. The index may be signaled in syntax structures of different granularity, such as at the sequence level, picture level, slice level, 64x64 grid level, 32x32 grid level, 16x16 grid level, etc.
[0134]
[0142] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a temporal motion vector prediction candidate selection method that utilizes an unconditional derivation of a scaled motion vector.
[0135]
[0143] According to the ECM, the temporal motion vector is scaled from either the L0 temporal motion or the L1 temporal motion, conditional on whether the current picture is a low-latency picture. According to an exemplary embodiment of the present disclosure, this condition may be omitted, and thus, when deriving a temporal merge candidate, the VVC standard encoder and the VVC standard decoder implement deriving a scaled motion vector from one of the motion vectors of the co-located CU. The one of the motion vectors of the co-located CU is determined according to the following steps.
[0136]
[0144] If the motion vector of the co-located CU is a bi-predictive motion vector, the L0 motion vector of the TMVP candidate is scaled from the L0 motion vector of the co-located CU, and the L1 motion vector of the TMVP candidate is scaled from the L1 motion vector of the co-located CU, regardless of whether the current picture is a low latency picture.
[0137]
[0145] Otherwise, if the motion vector of the co-located CU is an L0 predicted motion vector, then both the L0 and L1 motion vectors of the TMVP candidate are scaled from the L0 motion vector of the co-located CU, regardless of whether the current picture is a low latency picture. Similarly, if the motion vector of the co-located CU is an L1 predicted motion vector, then both the L0 and L1 motion vectors of the TMVP candidate are scaled from the L1 motion vector of the co-located CU.
[0138]
[0146] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a temporal motion vector prediction candidate selection method that omits scaling uni-predictive motion vectors to bi-predictive motion vectors.
[0139]
[0147] According to ECM, the motion vectors of TMVP candidates are always derived by bi-predictive motion, regardless of whether the motion of the co-located CU is uni-predictive or bi-predictive. Scaling uni-predictive motion vectors to bi-predictive motion vectors is not preferred because the scaling process is not accurate.
[0140]
[0148] According to some exemplary embodiments, it is proposed to omit the scaling process of converting uni-predictive motion vectors into bi-predictive motion vectors. When deriving a temporal merge candidate, the VVC standard encoder and the VVC standard decoder implement deriving a scaled motion vector from one of the motion vectors of the co-located CU. The one of the motion vectors of the co-located CU is determined according to the following steps.
[0141]
[0149] If the motion vector of the co-located CU is a bi-predictive motion vector and the current picture is a low latency picture, the L0 motion vector of the TMVP candidate is scaled from the L0 motion vector of the co-located CU, and the L1 motion vector of the TMVP candidate is scaled from the L1 motion vector of the co-located CU.
[0142]
[0150] Otherwise, if the motion of the co-located CU is a bi-predictive motion vector and the current picture is a non-low latency picture, which of the two motion vectors of the co-located CU is used to perform scaling is determined according to the reference picture list of the co-located CU. That is, if the co-located CU is from the L0 reference picture list, the L0 motion vector and the L1 motion vector of the TMVP candidate are both scaled from the L1 motion vector of the co-located CU. Similarly, if the co-located CU is from the L1 reference picture list, the L0 motion vector and the L1 motion vector of the TMVP candidate are both scaled from the L0 motion vector of the co-located CU.
[0143]
[0151] Otherwise, if the motion vector of the co-located CU is an L0 predicted motion vector, the L0 motion vector of the TMVP candidate is scaled from the L0 motion vector of the co-located CU, regardless of whether the current picture is a low latency picture, while the L1 motion vector of the TMVP candidate is set to be unavailable. Similarly, if the motion of the co-located CU is an L1 predicted motion, the L1 motion vector of the TMVP candidate is scaled from the L1 motion vector of the co-located CU, while the L0 motion vector of the TMVP candidate is set to be unavailable.
[0144]
[0152] According to another example embodiment, the scaling process of converting uni-predictive motion to bi-predictive motion may be skipped only for the lowest temporal layer, rather than for all temporal layers, e.g., when the temporal layer is lower than layer 3, scaling is skipped.
[0145]
[0153] According to another example embodiment, the scaling process of converting uni-predictive motion to bi-predictive motion may be omitted only for the lowest temporal layer and only for non-low latency pictures.
[0146]
[0154] According to other exemplary embodiments, the scaling process of converting uni-predictive motion to bi-predictive motion is omitted only for some merge modes, which may be any, some, or all of the following: normal merge mode, merge with MVD, geometric partitioning mode, combined inter and intra mode, sub-block-based temporal motion vector prediction, affine merge mode, and template matching mode.
[0147]
[0155] According to another exemplary embodiment, whether the motion vector of a TMVP candidate is derived by uni-predictive motion or bi-predictive motion is determined according to the cost value of template matching. The VVC standard encoder and VVC standard decoder implement measuring the template matching cost by the absolute difference sum between the template samples of the current block and their corresponding reference samples. The template includes a set of reconstructed samples in the neighborhood of the current block. The template reference samples are located by the L0 predictive motion information, the L1 predictive motion information, and the bi-predictive motion information of the temporal candidate.
[0148]
[0156] It should be understood that template matching to determine the uni-predictive or bi-predictive motion vector of a TMVP candidate is performed only when ARMC is enabled. Furthermore, to simplify implementation, when constructing a merge candidate list, the VVC standard encoder and VVC standard decoder first implement scaling the TMVP candidate to a bi-predictive motion vector. Then, when ARMC is applied, the TMVP candidate can be converted to a uni-predictive motion vector based on the cost value of template matching.
[0149]
[0157] According to another example embodiment, a VVC standard encoder and a VVC standard decoder implement adding an additional uni-predictive TMVP candidate when the motion of a co-located CU is uni-predictive.
[0150]
[0158] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a temporal motion vector prediction candidate selection method that utilizes multiple options in setting reference picture indexes.
[0151]
[0159] According to the ECM, the reference picture index of the temporal merge candidate is set equal to 0. According to an example embodiment of this disclosure, a different reference picture index may be selected.
[0152]
[0160] According to some example embodiments, the selected reference picture index is the reference picture index of the co-located picture whose scaling factor (i.e., tb / td shown in FIG. 4) is closest to one.
[0153]
[0161] According to another example embodiment, the selected reference picture index is the most frequently selected reference picture index for spatially neighboring blocks. The spatial neighboring blocks may be spatial candidates, HMVP candidates or non-neighboring candidates, as described above with respect to the VVC standard and ECM.
[0154]
[0162] According to other example embodiments, a VVC standard encoder and a VVC standard decoder implement signaling of reference picture indexes in a sequence level, a picture level, a slice level or a CTU level syntax structure.
[0155]
[0163] According to another exemplary embodiment, a VVC standard encoder and a VVC standard decoder implement selecting a different reference picture index for each sub-block when a block is coded using a sub-block-based temporal motion vector prediction ("SbTMVP") mode.
[0156]
[0164] According to another example embodiment, a VVC standard encoder and a VVC standard decoder implement determining a consensus reference picture index for each sub-block based on a per-sub-block reference picture index selection when the block is coded using SbTMVP mode. For each sub-block, a reference picture index is first selected, and the selected reference picture index is the reference index of the co-located picture whose scaling factor is closest to 1. Then, the reference picture index for the whole block is the most frequently selected reference picture index among each sub-block.
[0157]
[0165] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a temporal motion vector prediction candidate selection method that utilizes scaling factor offsetting.
[0158]
[0166] According to the ECM, the scaling factor is calculated using the POC distance as described above with respect to the VVC standard and the ECM, but the scaling factor calculation is observed to be inaccurate. Exemplary embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a scaling factor offset to improve accuracy.
[0159]
[0167] According to some example embodiments, the scaling factor may be offset as follows:
[0160]
number
[0161]
[0168] In this specification, tb denotes the POC difference between the reference picture of the current picture and the current picture, td denotes the POC difference between the reference picture of the co-located picture and the co-located picture, and N denotes a non-zero integer (e.g., N is equal to ±8, ±16). Assuming a negative N, the scaling factor is adjusted to be smaller. Assuming a positive N, the scaling factor is adjusted to be larger.
[0162]
[0169] According to another exemplary embodiment, the above offset scaling factor is influenced to be closer to 1. Assuming a scaling factor smaller than 1, N defined above is set to a positive number. Assuming a scaling factor larger than 1, N is set to a negative number.
[0163]
[0170] According to other example embodiments, a VVC standard encoder and a VVC standard decoder implement signaling the offset (ie, the number N) in a sequence level, picture level, slice level, or CTU level syntax structure.
[0164]
[0171] As an example, when signaling an offset, both the magnitude and the sign of the offset are signaled.
[0165]
[0172] As another example, when signaling an offset, only the sign of the offset is signaled, and the absolute value is fixed to a default number.
[0166]
[0173] As another example, the absolute value and sign of the offset may be signaled at different levels, with the absolute value being signaled in a sequence level syntax structure and the sign being signaled in a CTU level syntax structure.
[0167]
[0174] According to another example embodiment, the scaling factor is offset or not offset for each CU individually. For each CU, the scaling factor is
[0168]
number
[0169]
[0175] According to other example embodiments, the scaling factors may be offset for each temporal layer, i.e., each temporal layer may have a different scaling factor offset.
[0170]
[0176] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a merge candidate list creation method that omits temporal motion vector prediction candidates.
[0171]
[0177] According to ECM, temporal merge candidates are added to normal merge mode, geometric partition mode ("GPM"), merge mode with MVD ("MMVD"), combined inter and intra prediction ("CIIP"), SbTMVP and affine mode. It is observed that temporal merge candidates are not always ultimately used in coding.
[0172]
[0178] According to some example embodiments, a VVC standard encoder and a VVC standard decoder implement conditionally adding temporal merge candidates to a merge candidate list according to motion information of neighboring blocks. When the temporal motion of a neighboring block is similar to that of a current block and the motion of a neighboring block is not obtained from the temporal motion, the TMVP candidate of the current block is treated as unavailable for adding to the merge candidate list.
[0173]
[0179] 8 illustrates conditionally adding a temporal merge candidate to a merge candidate list according to motion information of neighboring blocks, according to an exemplary embodiment of the present disclosure. In FIG. 8, the motion vector
[0174]
number
[0175]
number
[0176]
number
[0177]
number
[0178]
number
[0179]
number
[0180]
number
[0181]
[0180] The neighboring blocks can be any subset of {A0, A1, B0, B1, B2}. The neighboring blocks can also be non-adjacent spatial or HMVP merging candidates.
[0182]
[0181] The similarity between the temporal motion vector of the current block and the temporal motion vector of the neighboring block is compared with a default threshold value. When the motion vector difference is smaller than the default threshold value, the two temporal motion vectors are treated as similar.
[0183]
[0182] The default threshold is an integer greater than 0. The default threshold may be set to different values depending on the coding mode of the current block or the size of the current block. For example, the default threshold is set to 1 for normal merging mode and 16 for template matching mode.
[0184]
[0183] According to another example embodiment, the VVC standard encoder and the VVC standard decoder set the adaptive merge list construction order according to the temporal layer, the picture type (e.g., low latency or non-low latency picture), or the coding mode of the current CU. In one example, for higher temporal layers, the priority of TMVP candidates is higher, which causes TMVP candidates to be preferentially added before spatial merge candidates, thereby overriding the merge candidate order described above.
[0185]
[0184] In one or more aspects, example embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a picture reconstruction method that utilizes motion information refinement.
[0186] According to some exemplary embodiments, after a current block coded in inter mode is reconstructed, the motion information, including inter prediction direction (i.e., L0 prediction, L1 prediction, or bi-prediction), reference picture index, and motion vector, is refined. The reconstructed sample of the current block is used as a template to perform motion estimation.
[0187]
[0186] When performing motion estimation, only the distortion is considered. The refined motion information is then used as the temporal motion for future coding pictures. The reconstructed samples used in the motion estimation process can be samples before or after the loop filter process.
[0188]
[0187] According to some exemplary embodiments, when constructing a merge candidate list for normal merge mode, CIIP, GPM, MMVD and template matching mode, the VVC standard encoder and VVC standard decoder implement treating the TMVP candidate of the current block as unavailable if the temporal motion of the neighboring block is similar to that of the current block and the motion of the neighboring block is not obtained from the temporal motion. When a TMVP candidate is added to the merge candidate list, the TMVP candidate is derived as follows:
[0189]
[0188] If the motion of the co-located CU is bi-predictive motion, regardless of whether the current picture is a low latency picture or not, the L0 motion vector of the TMVP candidate is scaled from the L0 motion vector of the co-located CU, and the L1 motion vector of the TMVP candidate is scaled from the L1 motion vector of the co-located CU.
[0190]
[0189] Otherwise, if the motion vector of the co-located CU is an L0 predicted motion vector, then both the L0 and L1 motion vectors of the TMVP candidate are scaled from the L0 motion vector of the co-located CU, regardless of whether the current picture is a low latency picture or not. Similarly, if the motion vector of the co-located CU is an L1 predicted motion vector, then both the L0 and L1 motion vectors of the TMVP candidate are scaled from the L1 motion vector of the co-located CU.
[0191]
[0190] Furthermore, the VVC standard encoder and the VVC standard decoder implement setting the reference picture index of a TMVP candidate to the reference picture index of the co-located picture whose scaling factor (i.e., tb / td shown in FIG. 4) is closest to 1. Furthermore, when ARMC is enabled and the current picture is a non-low latency picture, the VVC standard encoder and the VVC standard decoder implement applying a template matching method to determine whether the TMVP candidate is uni-predictive or bi-predictive and to select the optimal offset or non-offset scaling factor.
[0192]
[0191] Those skilled in the art will recognize that all of the above aspects of the present disclosure may be implemented simultaneously in any combination thereof, and that all aspects of the present disclosure may be implemented in combination as another embodiment of the present disclosure.
[0193]
[0192] FIG. 9 illustrates an example system 900 for implementing the above-described processes and methods for implementing refined temporal motion candidate behavior.
[0194]
[0193] The techniques and mechanisms described herein may be implemented by multiple instances of the system 900, as well as by any other computing device, system, and / or environment. The system 900 shown in FIG. 9 is only one example of a system and is not intended to suggest any limitation on the scope of use or functionality of any computing device utilized to implement the processes and / or procedures described above. Other well-known computing devices, systems, environments, and / or configurations that may be suitable for use with the embodiments include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, gaming consoles, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, implementations using field programmable gate arrays ("FPGAs") and application specific integrated circuits ("ASICs"), and the like.
[0195] The system 900 may include one or more processors 902 and a system memory 904 communicatively coupled to the one or more processors 902. The one or more processors 902 may execute one or more modules and / or processes to cause the one or more processors 902 to perform various functions. In some embodiments, the one or more processors 902 may include a central processing unit ("CPU"), a graphics processing unit ("GPU"), both a CPU and a GPU, or other processing units or components known in the art. Additionally, each of the one or more processors 902 may possess its own local memory, which may also store program modules, program data, and / or one or more operating systems.
[0196] Depending on the exact configuration and type of system 900, the system memory 904 may be volatile, such as RAM, non-volatile, such as ROM, flash memory, a mini hard drive, a memory card, etc., or some combination thereof. The system memory 904 may include one or more computer-executable modules 906 that are executable by the one or more processors 902.
[0197]
[0196] The module 906 may include, but is not limited to, one or more of an encoder 908 and a decoder 910.
[0198]
[0197] The encoder 908 may be a VVC standard encoder that implements any, some, or all aspects of the exemplary embodiments of the present disclosure described above and is executable by one or more processors 902 to configure the one or more processors 902 to perform the operations described above.
[0199]
[0198] The decoder 910 may be a VVC standard encoder that implements any, some, or all aspects of the exemplary embodiments of the present disclosure described above, and is executable by one or more processors 902 to configure the one or more processors 902 to perform the operations described above.
[0200]
[0199] The system 900 may further include an input / output (I / O) interface 940 for receiving image source data and bitstream data, and for outputting reconstructed pictures to a reference picture buffer or DBP and / or a display buffer. The system 900 may also include a communication module 950 that enables the system 900 to communicate with other devices (not shown) over a network (not shown). The network may include the Internet, wired media (such as a wired network or a direct wired connection), wireless media (such as acoustic, radio frequency ("RF"), infrared, and other wireless media).
[0201]
[0200] Some or all of the operations of the methods described above may be performed by execution of computer-readable instructions stored on a computer-readable storage medium, as defined below. The term "computer-readable instructions" as used in this specification and claims includes routines, applications, application modules, program modules, programs, components, data structures, algorithms, etc. The computer-readable instructions may be implemented on a variety of system configurations, including single-processor or multi-processor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based, programmable consumer electronics, combinations thereof, and the like.
[0202]
[0201] The computer-readable storage medium may include volatile memory (such as random access memory ("RAM")) and / or non-volatile memory (such as read-only memory ("ROM"), flash memory, etc.). The computer-readable storage medium may also include additional removable and / or non-removable storage including, but not limited to, flash memory, magnetic storage, optical storage, and / or tape storage that may provide non-volatile storage of computer-readable instructions, data structures, program modules, and the like.
[0203]
[0202] A non-transitory computer-readable storage medium is an example of a computer-readable medium. A computer-readable medium includes at least two types of computer-readable media, namely, a computer-readable storage medium and a communication medium. A computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented in any process or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, phase-change memory ("PRAM"), static random access memory ("SRAM"), dynamic random access memory ("DRAM"), other types of random access memory ("RAM"), read-only memory ("ROM"), electrically erasable programmable read-only memory ("EEPROM"), flash memory or other memory technology, compact disc read-only memory ("CD-ROM"), digital versatile disk ("DVD") or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that may be used to store information for access by a computing device. In contrast, communication media may embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, or other transmission mechanism. Computer-readable storage media as employed herein are not to be construed as transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating in a waveguide or other transmission medium (such as light pulses passing through a fiber optic cable), or electrical signals propagating in wires.
[0204]
[0203] Computer-readable instructions stored on one or more non-transitory computer-readable storage media, which when executed by one or more processors, may perform the operations described above with reference to Figures 1A-8. Generally, computer-readable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular abstract data types. The order in which the operations are described is not to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement a process.
[0205]
[0204] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.
[0206]
[0205] Exemplary embodiments of the present disclosure are further described by at least the following clauses:
[0207]
[0206] A. A method comprising selecting multiple motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein the one or more processors (902) derive temporal motion vector prediction candidates ("TMVP candidates") for a current CU of a current CTU (704) of the current picture (706) from a repositioned collocated CTU (708) of the collocated picture (710) in accordance with a motion vector (702) of a motion grid of the current picture (706), and the repositioned collocated CTU (708) is positioned in the collocated picture (710) by the motion vector (702) relative to the current CTU (704) in the current picture (706).
[0208]
[0207] B. The method of paragraph A, further comprising signaling the motion grid size of the current picture (706) in a sequence level, picture level or slice level syntax structure, wherein the motion grid size of the current CU and the co-located CU is different from the size of the current CTU (704) and the co-located CTU (708).
[0209]
[0208] C. The method of paragraph B, wherein signaling the motion grid size of the current CU includes signaling the motion grid size of the highest temporal layer of the current picture (706) in a sequence level syntax structure.
[0210]
[0209] D. The method of paragraph B, wherein signaling the grid size of the current CU includes decreasing the grid size for every lower temporal layer of the current picture (706), starting from the highest temporal layer of the current picture (706).
[0211]
[0210] E. The method of paragraph B, wherein signaling the grid size of the current CU includes increasing the grid size for every higher temporal layer of the current picture (706), starting from the lowest temporal layer of the current picture (706).
[0212]
[0211] F. The method of paragraph A, in which the motion vector (702) is not signaled in the syntax structure, and a parameter indicating the identity of the motion vector (702) to another motion vector of a neighboring motion grid is signaled in the syntax structure.
[0213]
[0212] G. The method of paragraph A, wherein the motion vector (702) is not signaled in the syntax structure, and a parameter indicating the difference of the motion vector (702) from another motion vector of a neighboring motion grid is signaled in the syntax structure.
[0214]
[0213] H. The method of paragraph A, in which the motion vector (702) is signaled in a syntax structure and the blocks of the motion grid are coded according to one of a default order: a raster scan order, a z-order scan order, a horizontal scan order, a vertical scan order, and a diagonal scan order.
[0215]
[0214] I. The method of paragraph A, further comprising signaling a parameter in a sequence level, picture level, or slice level syntax structure flag indicating that the position of the co-located CTU (708) is to be repositioned from the current CTU (704).
[0216]
[0215] J. The method of paragraph A, in which no parameter is signaled in a sequence level, picture level, or slice level syntax structure flag indicating that the position of the co-located CTU (708) is to be repositioned from the current CTU (704).
[0217]
[0216] K. A method comprising selecting multiple motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein the one or more processors (902) select a temporal motion vector prediction candidate ("TMVP candidate") from a set of motion candidates including spatially proximate blocks of a co-located picture (710) for the current CU.
[0218]
[0217] L. The method of paragraph K, in which one or more processors (902) derive TMVP candidates by averaging temporal motion vectors of multiple temporal candidates.
[0219]
[0218] M. The method of paragraph K, in which one or more processors (902) derive a TMVP candidate by one of a plurality of temporal motion vectors of the temporal candidate that produces the largest motion vector difference compared to the spatial merge candidate.
[0220]
[0219] N. The method of paragraph K, wherein one or more processors (902) derive TMVP candidates by the motion vector of the temporal candidate having the lowest cost value of template matching, and the cost value of template matching includes an absolute difference sum between the template samples of the current block and corresponding reference samples in the vicinity of the current block.
[0221]
[0220] O. The method of paragraph K, further comprising signaling, by one or more processors (902), an index identifying the derived TMVP candidate in a syntax structure of the CTU, wherein the syntax structure is at the sequence level, the picture level, the slice level, the 64x64 grid level, the 32x32 grid level, or the 16x16 grid level.
[0222]
[0221] P. A method for predicting a CU (706) from a current coding unit ("CU") of a current picture (706), comprising: selecting, by one or more processors (902) of a computing system (900), a plurality of motion candidates for the current coding unit ("CU") of the current picture (706), wherein the one or more processors (902) select temporal motion vector prediction candidates ("TMVP candidates") by deriving scaled motion vectors from motion vectors of the co-located CUs, the scaled motion vectors including an L0 motion vector and an L1 motion vector, and for bi-predictive motion vectors of the co-located CUs, regardless of whether the current picture (706) is a low latency picture, the L0 motion vector being a L1 motion vector. a co-located CU, wherein the L0 motion vector is scaled from the L0 motion vector of the co-located CU, and the L1 motion vector is scaled from the L1 motion vector of the co-located CU, and for the L0 predicted motion vector of the co-located CU, both the L0 motion vector and the L1 motion vector are scaled from the L0 motion vector of the co-located CU regardless of whether the current picture (706) is a low latency picture, and for the L1 predicted motion vector of the co-located CU, both the L0 motion vector and the L1 motion vector are scaled from the L1 motion vector of the co-located CU regardless of whether the current picture (706) is a low latency picture.
[0223] Q. A method including selecting, by one or more processors (902) of a computing system (900), a plurality of motion candidates for a current coding unit ("CU") of a current picture (706), where the one or more processors (902) select temporal motion vector prediction candidates ("TMVP candidates") by deriving scaled motion vectors from motion vectors of the co-located CUs, the scaled motion vectors including an L0 motion vector and an L1 motion vector, and for the bi-predictive motion vectors of the co-located CU, if the current picture (706) is a low-latency picture, the L0 motion vector is scaled from the L0 motion vector of the co-located CU, and the L1 motion vector is scaled from the L1 motion vector of the co-located CU. Or, if the co-located CU is from an L0 reference picture list and the current picture (706) is a low-latency picture, both the L0 motion vector and the L1 motion vector are scaled from the L0 motion vector of the co-located CU. Or if the co-located CU is from the L1 reference picture list and the current picture (706) is a low latency picture, then both the L0 and L1 motion vectors are scaled from the L1 motion vector of the co-located CU. For the L0 predicted motion vector of the co-located CU, regardless of whether the current picture (706) is a low latency picture, the L0 motion vector is scaled from the L0 motion vector of the co-located CU and the L1 motion vector is set to unavailable, and for the L1 predicted motion vector of the co-located CU, regardless of whether the current picture (706) is a low latency picture, the L0 motion vector is set to unavailable and the L1 motion vector is scaled from the L1 motion vector of the co-located CU.
[0224]
[0223] R. The method of paragraph Q, wherein the temporal layer of the current CU is lower than layer 3.
[0225] S. A method including selecting, by one or more processors (902) of a computing system (900), a plurality of motion candidates for a current coding unit ("CU") of a current picture (706), where the one or more processors (902) select temporal motion vector prediction candidates ("TMVP candidates") by deriving scaled motion vectors from motion vectors of a co-located CU, the scaled motion vectors including an L0 motion vector and an L1 motion vector, and for a bi-predictive motion vector of the co-located CU, if the current picture (706) is a non-low latency picture, the L0 motion vector is scaled from the L0 motion vector of the co-located CU, and the L1 motion vector is scaled from the L1 motion vector of the co-located CU; or if the co-located CU is from an L0 reference picture list and the current picture (706) is a low latency picture, the L0 motion vector and the L1 motion vector are both scaled from the L0 motion vector of the co-located CU. Or if the co-located CU is from the L1 reference picture list and the current picture (706) is a non-low latency picture, both the L0 motion vector and the L1 motion vector are scaled from the L1 motion vector of the co-located CU. For the L0 predicted motion vector of the co-located CU, if the current picture (706) is a non-low latency picture, the L0 motion vector is scaled from the L0 motion vector of the co-located CU and the L1 motion vector is set to unavailable, and for the L1 predicted motion vector of the co-located CU, if the current picture (706) is a non-low latency picture, the L0 motion vector is set to unavailable and the L1 motion vector is scaled from the L1 motion vector of the co-located CU.
[0226]
[0225] T. The method of paragraph S, wherein the temporal layer of the current CU is lower than layer 3.
[0227]
[0226] U. A method according to any one of paragraphs Q, R, S, or T, wherein multiple motion candidates are selected for one of a normal merge mode, a merge with MVD, a geometric partitioning mode, a combined inter and intra mode, a sub-block based temporal motion vector prediction, an affine merge mode, and a template matching mode.
[0228]
[0227] V. A method comprising selecting multiple motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein the one or more processors (902) select a temporal motion vector prediction candidate ("TMVP candidate") by deriving a scaled motion vector from a motion vector of a co-located CU, either by uni-predictive motion or bi-predictive motion, in response to a lowest cost value of template matching, wherein the cost value of template matching comprises an absolute sum of differences between a template sample of the current block and a corresponding reference sample in a neighborhood of the current block.
[0229]
[0228] W. The method of paragraph V, in which the scaled motion vectors are derived by scaling TMVP candidates into bi-predictive motion vectors and then converting the bi-predictive motion vectors into uni-predictive motion vectors.
[0230]
[0229] X. The method of paragraph V, wherein the TMVP candidates include two TMVP candidates each derived by uni-predictive motion.
[0231]
[0230] Y. A method comprising selecting, by one or more processors (902) of a computing system (900), multiple motion candidates for a current coding unit ("CU") of a current picture (706), wherein the one or more processors (902) select a temporal motion vector prediction candidate ("TMVP candidate") and set a reference picture index of the TVMP candidate to a value other than 0.
[0232]
[0231] Z. The method of paragraph Y, wherein the reference picture index includes a reference picture index of a co-located picture having a scaling factor closest to 1.
[0233]
[0232] AA. The method of paragraph Y, wherein the reference picture index includes one of a plurality of reference picture indexes that is most frequently selected for blocks that are spatially nearby.
[0234]
[0233] AB. The method of paragraph Y, further comprising signaling, by one or more processors (902), the reference picture index in a sequence level, picture level, slice level or CTU level syntax structure.
[0235]
[0234] AC. The method of paragraph Y, wherein the current block is coded using a sub-block-based temporal motion vector prediction ("SbTMVP") mode, and the method further includes selecting at least some different reference picture indexes for different sub-blocks of the current block.
[0236]
[0235] A.D. The method of paragraph Y, wherein the current block is coded using a sub-block-based temporal motion vector prediction ("SbTMVP") mode, and the method further includes selecting, for each sub-block of the current block, a reference picture index of a co-located picture having a scaling factor closest to 1, and determining the most frequently selected reference picture index for the sub-blocks of the current block as a consensus reference picture index for the current block.
[0237]
[0236] AE. A method comprising selecting, by one or more processors (902) of a computing system (900), multiple motion candidates for a current coding unit ("CU") of a current picture (706), wherein the one or more processors (902) select temporal motion vector prediction candidates ("TMVP candidates") by deriving a scaled motion vector from a motion vector of a co-located CU by a scaling factor, the scaling factor being offset by adding to a scaling factor multiplied by an offset.
[0238]
[0237] AF. The method of paragraph AE, further including signaling the offset in a sequence level, picture level, slice level, or CTU level syntax structure.
[0239]
[0238] AG. The method of paragraph AF, wherein both the absolute value of the offset and the sign of the offset are signaled.
[0240]
[0239] AH. The method of paragraph AF, in which the magnitude and sign are signaled in different levels of a syntax structure.
[0241]
[0240] AI. The method of paragraph AF, wherein the sign of the offset is signaled and the absolute value of the offset is not signaled.
[0242]
[0241] AJ. The method of paragraph AE, in which the scaling factor is offset or not offset individually for each CU of the current picture (706).
[0243]
[0242] AK. The method of paragraph AE, in which the scaling factor is different for each temporal layer of the current picture (706).
[0244]
[0243] AL. A method including selecting, by one or more processors (902) of a computing system (900), multiple motion candidates for a current coding unit ("CU") of a current picture (706), wherein the one or more processors (902) do not select a temporal motion vector prediction candidate ("TMVP candidate") if the motion vector of the current CU is similar to the motion vector of a neighboring block of the co-located CU and the motion vector of the neighboring block of the current CU is not a scaled motion vector derived from the neighboring block of the co-located CU.
[0245]
[0244] AM. The method of paragraph AL, wherein the similarity is determined according to a similarity threshold based on at least one of a coding mode of the current block and a size of the current block.
[0246]
[0245] AN. A method including selecting multiple motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein the one or more processors (902) adaptively vary an order in which the multiple merge candidates are selected based on at least one of a temporal layer of the current picture (706) and a coding mode of the current block, regardless of whether the current picture (706) is low latency or non-low latency.
[0247]
[0246] AO. The method of paragraph AN, in which one or more processors (902) preferentially select temporal motion vector prediction candidates ("TMVP candidates") before spatial motion vector prediction candidates for a current picture (706) having a high temporal layer.
[0248]
[0247] AP. A method comprising: reconstructing, by one or more processors (902) of a computing system (900), an inter-predictively coded current block; performing, by the one or more processors (902), motion estimation using samples of the reconstructed block as templates to generate refined motion information for the reconstructed block; and selecting, by the one or more processors (902), temporal motion vector prediction candidates ("TMVP candidates") derived from the refined motion information for the reconstructed block.
Claims
1. A method for encoding a video sequence, comprising: receiving a video sequence; The video sequence Selecting multiple merging candidates for a current coding unit ("CU") of a current picture By encoding it, wherein a temporal motion vector prediction candidate ("TMVP candidate") is selected by setting the reference picture index of the TMVP candidate to the reference picture index of the co-located picture whose scaling factor is closest to 1.
2. The scaling factor is a picture order count ("POC") difference between the current picture's reference picture and the current picture; a POC difference between the co-located picture and a reference picture for the co-located picture; The method of claim 1 , comprising a ratio between:
3. The encoding step comprises: The method of claim 1 , further comprising sorting the plurality of merge candidates according to adaptive sorting of merge candidates ("ARMC").
4. 2. The method of claim 1, wherein the TMVP candidates of the plurality of merging candidates are sorted by cost values based on template matching, and the template matching cost values of the merging candidates comprise sums of absolute differences ("SAD") between the template samples of the current block and the respective reference samples.
5. A method for signaling a bitstream, comprising: receiving a video sequence; The video sequence Selecting multiple merging candidates for a current coding unit ("CU") of a current picture and encoding the signaling a bitstream generated based on said encoding. wherein a temporal motion vector prediction candidate ("TMVP candidate") is selected by setting the reference picture index of the TMVP candidate to the reference picture index of a co-located picture whose scaling factor is closest to 1.
6. The scaling factor is a picture order count ("POC") difference between the current picture's reference picture and the current picture; a POC difference between the co-located picture and a reference picture for the co-located picture; The method of claim 5 , comprising a ratio between:
7. The encoding of the video sequence, The method of claim 5 , further comprising sorting the plurality of merge candidates according to adaptive sorting of merge candidates ("ARMC").
8. 6. The method of claim 5, wherein the TMVP candidates of the plurality of merging candidates are sorted by cost values based on template matching, and the template matching cost values of the merging candidates comprise sums of absolute differences ("SAD") between the template samples of the current block and the respective reference samples.
9. A method for decoding a bitstream, comprising: receiving a bitstream; decoding the bitstream to output a video sequence; wherein said decoding comprises: Parsing a merge index signaled in a syntax structure of a current coding unit ("CU") of a current picture; selecting multiple merge candidates for the CU, wherein a temporal motion vector prediction candidate ("TMVP candidate") is selected by setting a reference picture index of the TMVP candidate to a reference picture index of a co-located picture whose scaling factor is closest to 1; performing inter prediction for the CU by applying a merge mode based on the merge index and the plurality of merge candidates; A method comprising:
10. The scaling factor is a picture order count ("POC") difference between a reference picture of a current picture and the current picture; and a POC difference between the co-located picture and a reference picture for the co-located picture; 10. The method of claim 9, comprising a ratio between:
11. 10. The method of claim 9, wherein the decoding further comprises reordering a plurality of merging candidates according to adaptive reordering of merging candidates ("ARMC").
12. 10. The method of claim 9, wherein the TMVP candidates of the plurality of merging candidates are sorted by cost values based on template matching, and the template matching cost values of the merging candidates comprise sums of absolute differences ("SAD") between the template samples of the current block and the respective reference samples.