Improved time-based merge candidates in merge candidate lists in video coding.
The proposed time-motion vector prediction candidate selection method addresses the underperformance of TMVP in VVC and ECM by repositioning the collated CTU, expanding the selection range, and applying a scaling factor offset, resulting in improved competitiveness and reduced redundancy in video coding.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ALIBABA DAMO (HANGZHOU) TECH CO LTD
- Filing Date
- 2022-09-28
- Publication Date
- 2026-07-22
AI Technical Summary
The existing video coding standards, such as VVC and ECM, do not effectively improve the performance of Time Motion Vector Prediction (TMVP) beyond the HEVC implementation, leading to redundancy and underperformance compared to other motion prediction techniques.
Implement a time-motion vector prediction candidate selection method that includes repositioning of a collated CTU, expands the selection range, omits scaling single prediction motion vectors to biprediction motion vectors, utilizes multiple reference picture index options, and applies a scaling factor offset, while omitting time motion vector prediction candidates when necessary, to enhance the TMVP process.
Enhances the performance of TMVP by improving its competitiveness with other motion prediction techniques, reducing redundancy and optimizing computational efficiency in video coding processes.
Smart Images

Figure 0007893862000010 
Figure 0007893862000011 
Figure 0007893862000012
Abstract
Description
Technical Field
[0001] Related Applications
[0001] This PCT application claims priority to U.S. Patent Application No. 63 / 250,208, filed Sep. 29, 2021, which is incorporated herein by reference.
Background Art
[0002]
[0002] In 2020, the Joint Video Expert Team (「JVET」) of the ITU-T Video Coding Experts Group (「ITU-T VCEG」) and the ISO / IEC Moving Picture Experts Group (「ISO / IEC MPEG」) published the final draft of the next-generation video codec specification, Versatile Video Coding (「VVC」). This specification further improves video coding performance compared to previous standards such as H.264 / AVC (Advanced Video Coding) and H.265 / HEVC (High Efficiency Video Coding). JVET has continued to propose additional techniques beyond the scope of the VVC standard itself, which are collected under the name of Extended Compression Model (「ECM」).
[0003]
[0003] In each of the AVC, HEVC, and VVC standards, along with the Discrete Cosine Transform (「DCT」), Motion Compensation Prediction (「MCP」) is implemented as another central image compression technique. The image is divided into coding blocks, and MCP improves the compression efficiency of the image based on the principle that the motion in a block of a picture tends to repeat in adjacent blocks, as well as in blocks of temporally preceding and subsequent pictures. MCP is implemented by searching for such 「motion candidates」 in motion prediction and deriving motion information therefrom to reconstruct the block.
[0004]
[0004] While previous standards such as HEVC implemented MCP based only on translational motion, VVC also implements affine motion compensation prediction ("affine MCP"). Generally, motion information between spatially neighboring blocks and temporally neighboring blocks is reconstructed based on different formats of motion vectors, in particular, the time motion vector is derived according to a technique based on time motion vector prediction ("TMVP"). In these forms, redundant motion information is reduced in the coded image, which reduces the bitrate required to transmit the video stream and thus achieves rate gain. [Overview of the project]
[0005]
[0005] The initial draft of ECM (presented as “Exploration experiment on enhanced compression beyond VVC capability” at the 133rd meeting of the Moving Picture Expert Group (“MPEG”) in January 2021) includes proposals to further expand the range of motion candidates explored according to the MCP technique of VVC. However, according to both the VVC implementation and the ECM implementation of MCP, the TMVP technique remains substantially unchanged from the HEVC implementation of MCP. Since TMVP remains an integral component of MCP, it is desirable to further improve the performance of TMVP so that it does not become redundant with respect to other motion vector prediction techniques.
[0006]
[0006] A detailed explanation is provided with reference to the attached figures. In the figures, the leftmost (one or more) digit of the reference number identifies the figure in which the reference number first appears. The use of the same reference number in different figures indicates similar or equivalent items or features. [Brief explanation of the drawing]
[0007] [Figure 1A]
[0007] This is an exemplary block diagram of a video encoding process according to an exemplary embodiment of the present disclosure. [Figure 1B] This is an exemplary block diagram of a video decoding process according to an exemplary embodiment of the present disclosure. [Figure 2]
[0008] This diagram shows multiple spatially neighboring blocks of the current CU of the picture. [Figure 3]
[0009] This figure shows an example selection of motion candidates for a picture's CU using motion prediction coding according to the VVC standard. [Figure 4]
[0010] This figure shows how to obtain scaled motion vectors for time merge candidates according to the VVC standard. [Figure 5]
[0011] This figure shows the selection of positions for time candidates between candidate C0 and candidate C1 according to the VVC standard. [Figure 6]
[0012] This figure shows possible spatial neighborhood blocks that can be derived using ECM3, including not only adjacent spatial merge candidates but also non-adjacent spatial merge candidates. [Figure 7A]
[0013] This figure shows the method for selecting candidate time motion vector predictions using ECM. [Figure 7B] This figure shows a method for selecting time-motion vector prediction candidates that utilizes the repositioning of a collated CTU, according to an exemplary embodiment of the present disclosure. [Figure 8]
[0014] This figure shows how, according to an exemplary embodiment of the present disclosure, time merge candidates are added to a merge candidate list according to the movement information of neighboring blocks. [Figure 9]
[0015] This figure shows an exemplary system for implementing the processes and methods described above for implementing improved time-motion candidate behavior. [Modes for carrying out the invention]
[0008]
[0016] According to the VVC Video Coding Standard ("VVC Standard") and the motion predictions described herein, computer-readable instructions stored in a computer-readable storage medium are executable by one or more processors of a computing system to perform the encoder operations and decoder operations described in the VVC Standard. Some of these encoder and decoder operations according to the VVC Standard will be described in more detail later, but these subsequent descriptions should not be understood as exhaustive of the encoder and decoder operations according to the VVC Standard. Thereafter, the "VVC Standard encoder" and "VVC Standard decoder" will be described as computer-readable instructions stored in a computer-readable storage medium that configure one or more processors to perform these respective operations (sometimes referred to as "reference implementations" of the encoder or decoder, for example).
[0009]
[0017] Furthermore, according to exemplary embodiments of this disclosure, the VVC standard encoder and VVC standard decoder further include computer-readable instructions stored in a computer-readable storage medium, which are executable by one or more processors of a computing system to perform operations not specified by the VVC standard. The VVC standard encoder should not be understood as being limited to the operation of the reference implementation of the encoder, but rather as including further computer-readable instructions to configure one or more processors of a computing system to perform further operations described herein. The VVC standard decoder should not be understood as being limited to the operation of the reference implementation of the decoder, but rather as including further computer-readable instructions to configure one or more processors of a computing system to perform further operations described herein.
[0010]
[0018] Figures 1A and 1B show exemplary block diagrams of an encoding process 100 and a decoding process 150, respectively, according to an exemplary embodiment of the present disclosure.
[0011]
[0019] In the encoding process 100, the VVC standard encoder configures one or more processors of the computing system to receive one or more input pictures from an image source 102 as input. The input picture contains a number of pixels sampled by an image capture device, such as a photosensor array, and contains an uncompressed stream of multiple color channels (such as RGB color channels) that store color data at the original resolution of the picture, with each channel using a number of bits to store the color data for each pixel of the picture. The VVC standard encoder configures one or more processors of the computing system to store this uncompressed color data in a compressed format, where the color data is stored at a resolution lower than the original resolution of the picture and encoded as a luma ("Y") channel and two chroma ("U" and "V") channels at a lower resolution than the luma channel.
[0012]
[0020] A VVC standard encoder encodes a picture (the picture being encoded, called the “current picture,” which is distinguishable from other pictures received from the image source 102) by configuring one or more processors of the computing system to divide the original picture into units and subunits according to a partitioning structure. The VVC standard encoder configures one or more processors of the computing system to subdivide the picture into macroblocks ("MBs"), each having dimensions of 16 × 16 pixels, and the MBs may be further subdivided into partitions. The VVC standard encoder configures one or more processors of the computing system to subdivide the picture into coding tree units ("CTUs"), and the luma and chroma components of the CTUs may be further subdivided into coding tree blocks ("CTBs"), and the CTBs may be further subdivided into coding units ("CUs"). Alternatively, the VVC standard encoder configures one or more processors of the computing system to subdivide the picture into N × N pixel units, and those units may then be further subdivided into subunits. Each of these largest subdivided units of the picture may, in general, be referred to as a “block” in this disclosure.
[0013]
[0021] CU is coded using one block of luma sample and two corresponding blocks of chroma sample, while picture is coded using a single coding tree, rather than monochrome.
[0014]
[0022] A VVC standard encoder configures one or more processors in a computing system to subdivide a block into sections having dimensions that are multiples of 4x4 pixels. For example, the block sections may have dimensions of 8x4 pixels, 4x8 pixels, 8x8 pixels, 16x8 pixels, or 8x16 pixels.
[0015]
[0023] Rather than encoding the pixel color information of the original picture at full resolution, the VVC standard encoder configures one or more processors of the computing system to encode the color information of the picture blocks and block subdivisions at a resolution lower than that of the input picture and store the color information in fewer bits than the input picture.
[0016]
[0024] Furthermore, the VVC standard encoder encodes a picture by configuring one or more processors of the computing system to perform motion prediction on the blocks of the current picture. Motion prediction coding refers to storing the image data of the blocks of the current picture (where the blocks of the original picture before coding are called "input blocks") using motion information and prediction units ("PUs") (which are not pixel data) by intra prediction 104 or inter prediction 106.
[0017]
[0025] Motion information refers to data that describes the motion of the block structure of a picture or a unit or its sub-units, such as motion vectors and references to blocks of the current picture or reference pictures. A PU may refer to one unit or multiple sub-units corresponding to one block structure among multiple block structures of a picture, such as an MB or a CTU. The blocks are divided based on picture data and coded according to the VVC standard. The motion information corresponding to a PU may describe the motion prediction encoded by the VVC standard encoder described herein.
[0018]
[0026] The VVC standard encoder configures one or more processors of the computing system to code the motion prediction information across each block of the picture in a coding order between blocks, such as a raster scan order where the first block to be decoded is the topmost and leftmost block of the picture. The block being coded is called the "current block" to be distinguished from other blocks of the same picture.
[0019]
[0027] According to the intra prediction 104, one or more processors of a computing system are configured to encode a block by using motion information of one or more other blocks of the same picture and a reference to a PU. According to the intra prediction coding, one or more processors of a computing system perform an intra prediction 104 calculation (also called spatial prediction) by coding motion information of a current block based on spatially neighboring samples from spatially neighboring blocks of the current block.
[0020]
[0028] According to the inter prediction 106, one or more processors of a computing system are configured to encode a block by using motion information of one or more other pictures and a reference to a PU. One or more processors of a computing system are configured to store one or more previously coded and decoded pictures in a reference picture buffer in inter prediction coding, and these stored pictures are called reference pictures.
[0021]
[0029] One or more processors are configured to perform an inter prediction 106 calculation (also called temporal prediction or motion compensation prediction) by coding motion information of a current block based on samples from one or more reference pictures. The inter prediction can be further calculated according to single prediction or dual prediction, that is, in single prediction, only one motion vector pointing to one reference picture is used to generate a prediction signal for the current block. In dual prediction, two motion vectors each pointing to a respective reference picture are used to generate a prediction signal for the current block.
[0022]
[0030] A VVC standard encoder configures one or more processors in a computing system to code a CU to include a reference index for identifying one or more prediction signals of the current block for reference by a VVC standard decoder. One or more processors in a computing system may code a CU to include an interprediction indicator. The interprediction indicator shows a List 0 prediction for a first reference picture list called List 0, a List 1 prediction for a second reference picture list called List 1, or a biprediction for both reference picture lists, called List 0 and List 1, respectively.
[0023]
[0031] If the interprediction indicator indicates a List 0 prediction or a List 1 prediction, one or more processors in the computing system are configured to code the CU including a reference index pointing to the reference picture in the reference picture buffer referenced by List 0 or List 1, respectively. If the interprediction indicator indicates a dual prediction, one or more processors in the computing system are configured to code the CU including a first reference index pointing to the first reference picture in the reference picture buffer referenced by List 0, and a second reference index pointing to the second reference picture in the reference picture referenced by List 1.
[0024]
[0032] A VVC standard encoder configures one or more processors in a computing system to code the current block of a picture individually and output a predicted block for each. According to the VVC standard, a CTU can be the same size as 128 × 128 chroma samples (and, depending on the chroma format, the corresponding chroma samples). A CTU can further be partitioned into CUs according to a quadtree, binary tree, or ternary tree. One or more processors in a computing system are configured to ultimately record a set of coding parameters, such as the coding mode (intra-mode or inter-mode), motion information for the intercoded block (reference index, motion vector, etc.), and quantized residual coefficients, in the syntactic structure of the leaf nodes of the partitioned structure.
[0025]
[0033] After the prediction block is output, the VVC standard encoder configures one or more processors of the computing system to send a set of coding parameters to the entropy coder 124 (described later), including the coding mode (i.e., intra or inter-prediction), the mode of intra-prediction or inter-prediction, and motion information.
[0026]
[0034] The VVC standard provides semantics for recording coding parameter sets for CUs. For example, with respect to the coding parameter sets described above, the pred_mode_flag for the CU is set to 0 for interconnected blocks and to 1 for intra-interconnected blocks; the general_merge_flag for the CU is set to indicate whether merge mode is used in the CU's inter-prediction; the inter_affine_flag and cu_affine_type_flag for the CU are set to indicate whether affine motion compensation is used in the CU's inter-prediction; the mvp_l0_flag and mvp_l1_flag are set to indicate motion vector indices in List 0 or List 1, respectively; and the ref_idx_l0 and ref_idx_l1 are set to indicate reference picture indices in List 0 or List 1, respectively. It should be understood that the VVC standard includes semantics for recording various other information, flags, and options beyond the scope of this disclosure.
[0027]
[0035] The VVC standard encoder further implements one or more mode determination and encoder control settings 108, including rate control settings. One or more processors of the computing system are configured to perform mode determination after intra or inter prediction by selecting an optimized prediction mode for the current block based on a rate strain optimization method.
[0028]
[0036] The rate control setting configures one or more processors in the computing system to assign different quantization parameters ("QP") to different pictures. The magnitude of the QP determines the scale to which the picture information is quantized during encoding by one or more processors (as will be described later), and thus determines the extent to which the encoding process 100 discards picture information from MB of the sequence during coding (by having the information enter between steps of the scale).
[0029]
[0037] The VVC standard encoder further implements a subtractor 110. One or more processors in the computing system are configured to perform the subtraction operation by calculating the difference between the input block and the prediction block. Based on the optimized prediction mode, the prediction block is subtracted from the input block. The difference between the input block and the prediction block is called the prediction residual, or for brevity, the "residual".
[0030]
[0038] Based on the predicted residuals, the VVC standard encoder further implements transformation 112. One or more processors in the computing system are configured to perform a transformation operation on the residuals by matrix arithmetic operations to derive an array of coefficients (sometimes called “residual coefficients,” “transformation coefficients,” etc.), thereby encoding the current block as a transformed block (“TB”). Transformation coefficients can refer to coefficients representing one of several spatial transformations, such as diagonal inversion, vertical inversion, or rotation, which may be applied to subblocks.
[0031]
[0039] It should be understood that coefficients can be remembered as having two components: their absolute value and their sign, as will be explained in more detail later.
[0032]
[0040] The subblocks of the CU, such as the PU and TB, can be arranged in any combination of subblock dimensions, as described above. The VVC standard encoder configures one or more processors in the computing system to subdivide the CU into a residual quadtree ("RQT"), which is a hierarchical structure of TBs. The RQT provides the motion prediction and residual coding sequence across the subblocks at each level of the RQT, recursively descending each level of the RQT.
[0033]
[0041] The VVC standard encoder further implements quantization 114. One or more processors in a computing system are configured to perform quantization operations on residual coefficients by matrix arithmetic operations based on the quantization matrix and the QP assigned above. Residual coefficients within a certain interval are retained, and residual coefficients outside that interval step are discarded.
[0034]
[0042] The VVC standard encoder further implements inverse quantization 116 and inverse transformation 118. One or more processors of a computing system are configured to perform inverse quantization and inverse transformation operations on the quantized residual coefficients by matrix arithmetic operations, which are the inverses of the quantization and transformation operations described above. The inverse quantization and inverse transformation operations yield the reconstructed residuals.
[0035]
[0043] The VVC standard encoder further implements an adder 120. One or more processors in the computing system are configured to perform an addition operation by adding the predicted block and the reconstructed residual, and to output the reconstructed block.
[0036]
[0044] The VVC standard encoder further implements a loop filter 122. One or more processors in the computing system are configured to apply loop filters, such as a deblocking filter, a sample-adaptive offset ("SAO") filter, and an adaptive loop filter ("ALF"), to the reconstructed block and output a filtered reconstructed block.
[0037]
[0045] The VVC standard encoder further configures one or more processors in the computing system to output filtered, reconstructed blocks to a decoded picture buffer ("DPB") 200. The DPB 200 stores the reconstructed pictures that are used by one or more processors in the computing system as reference pictures when coding pictures other than the current picture, as described above with respect to interpretation.
[0038]
[0046] The VVC standard encoder further implements the entropy coder 124. One or more processors in a computing system are configured to perform entropy coding, and according to the context-dependent binary arithmetic codec ("CABAC"), the symbols constituting the quantized residual coefficients are coded by mapping to binary sequences (hereinafter "bins"), which can be transmitted in a compressed bitrate output bitstream. The symbols of the quantized residual coefficients to be coded include the absolute values of the residual coefficients (these absolute values are hereinafter referred to as "residual coefficient levels").
[0039]
[0047] However, while residual coefficient levels are predicted and coded, residual coefficient codes are signaled using bins representing equally probable (hereinafter "EP") states (it should be understood that coefficients with a value of 0 do not have a sign and therefore do not need to be signaled). Due to computational challenges (which should be understood by those skilled in the art, but do not need to be repeated here to understand the exemplary embodiments of this disclosure), the VVC standard encoder does not configure one or more processors of the computing system to predict residual coefficient codes. For these reasons, the CABAC is configured to bypass coding of residual coefficient codes and transmit the output bitstream with one bit added for each code.
[0040]
[0048] Therefore, the entropy coder codes the residual coefficient level of a block, bypasses coding the residual coefficient code and records the residual coefficient code along with the coded block, records coding parameters such as coding mode, intra-prediction mode or inter-prediction mode, and motion information (e.g., picture parameter set ("PPS") contained in the picture header, and sequence parameter set ("SPS") contained in a sequence of multiple pictures) which are coded within the syntax structure of the coded block, and configures one or more processors of the computing system to output the coded block.
[0041]
[0049] The VVC standard encoder configures one or more processors in a computing system to output a coded picture consisting of coded blocks from the entropy coder 124. The coded picture is output to a transmit buffer and finally packed into a bitstream for output from the VVC standard encoder.
[0042]
[0050] In the decoding process 150, the VVC standard decoder configures one or more processors of the computing system to receive one or more coded pictures from the bitstream as input.
[0043]
[0051] The VVC standard decoder implements the entropy decoder 152. One or more processors in the computing system are configured to perform entropy decoding, and the bins are decoded by reversing the symbol-to-bin mapping according to CABAC, thereby recovering the entropy-coded quantized residual coefficients. The entropy decoder 152 outputs the quantized residual coefficients, the coding-bypassed residual coefficient codes, and syntax structures such as PPS and SPS.
[0044]
[0052] The VVC standard decoder further implements inverse quantization 154 and inverse transformation 156. One or more processors in a computing system are configured to perform inverse quantization and inverse transformation operations on the decoded quantized residual coefficients by matrix arithmetic operations, which are the inverses of the quantization and transformation operations described above. The inverse quantization and inverse transformation operations yield the reconstructed residuals.
[0045]
[0053] Furthermore, based on the coding parameter set recorded by the entropy coder 124 in syntax structures such as PPS and SPS (or received by out-of-band transmission or coded within the decoder), and the coding modes included in the coding parameter set, the VVC standard decoder determines whether to apply intra-prediction 156 (i.e., spatial prediction) or motion-compensated prediction 158 (i.e., temporal prediction) to the reconstructed residual.
[0046]
[0054] If the coding parameter set specifies intra-prediction, the VVC standard decoder configures one or more processors of the computing system to perform intra-prediction 156 using the prediction information specified in the coding parameter set. Intra-prediction 156 thereby generates a prediction signal.
[0047]
[0055] If the coding parameter set specifies interpretation, the VVC standard decoder configures one or more processors in the computing system to perform motion-compensated prediction 158 using a reference picture from DPB200. The motion-compensated prediction 158 thereby generates a prediction signal.
[0048]
[0056] The VVC standard decoder further implements an adder 160. The adder 160 performs an addition operation on the reconstructed residual and the predicted signal, thereby configuring one or more processors of the computing system to output the reconstructed block.
[0049]
[0057] The VVC standard decoder further implements the loop filter 162. One or more processors in the computing system are configured to apply deblocking filters, SAO filters, and loop filters such as ALF to the reconstructed block and output the filtered reconstructed block.
[0050]
[0058] The VVC standard decoder further configures one or more processors of the computing system to output filtered and reconstructed blocks to the DBP200. As described above, the DPB200 stores the reconstructed picture, which is used by one or more processors of the computing system as a reference picture when coding pictures other than the current picture, as described above with respect to motion compensation prediction.
[0051]
[0059] The VVC standard decoder further configures one or more processors in the computing system to output the reconstructed picture from the DPB to a user-viewable display of the computing system, such as a television display, personal computing monitor, smartphone display, or tablet display.
[0052]
[0060] Therefore, as shown by the encoding process 100 and decoding process 150 described above, the VVC standard encoder and VVC standard decoder each implement motion prediction coding according to the VVC specification. The VVC standard encoder and VVC standard decoder each configure one or more processors of the computing system to generate a reconstructed picture based on a previously reconstructed picture of the DPB, according to the motion compensation prediction described by the VVC standard, the previously reconstructed picture serves as a reference picture in the motion compensation prediction as described herein.
[0053]
[0061] As described above regarding the coding parameter set, for reconstructed pictures coded by interpredictive coding, VVC standard encoders and VVC standard decoders implement merge modes and affine motion compensation for interpredictive of the reconstructed blocks. VVC standard encoders and VVC standard decoders implement multiple merge modes for interpredictive of motion information of the reconstructed picture's CUs, including motion-compensated prediction ("MCP"), affine motion-compensated prediction ("affine MCP"), and other merge modes, as specified by the VVC standard. Motion information may include multiple motion vectors.
[0054]
[0062] The motion information of the reconstructed picture's CU may further include a motion candidate list. According to the VVC standard, the motion candidate list may be a data structure containing references to multiple motion candidates. A motion candidate may be a block structure, or a subunit of a block structure such as a pixel, or any other suitable subdivision of the current picture's block structure, or a reference to a motion candidate of another picture. A motion candidate may be a spatial motion candidate or a temporal motion candidate. By applying motion vector compensation ("MVC"), the VVC standard decoder may select a motion candidate from the motion candidate list and derive the motion vector of that motion candidate as the motion vector of the reconstructed picture's CU.
[0055]
[0063] Figure 3 shows an exemplary selection of motion candidates for a picture CU using merge mode coding according to the VVC standard.
[0056]
[0064] According to the VVC standard, a motion candidate list can be a merge candidate list and may contain up to five types of merge candidates (six according to the ECM, as will be explained later). VVC standard encoders can implement coding of the CU syntax structure to include merge indices.
[0057]
[0065] For each CU coded in merge mode, the index of the best merge candidate is encoded using truncated unary binauralization (TU).
[0058]
[0066] The list of merge candidates for a CU of a picture, coded according to the merge mode, may include the following merge candidates in order:
[0059]
[0067] Candidates for the spatially nearest CU to the current CU,
[0060]
[0068] Current Time MVP candidates ("TMVP candidates") from collated CUs for CUs,
[0061]
[0069] History-based MVP candidates from FIFO tables
[0062]
[0070] Pairwise average MVP candidates, and
[0063]
[0071] Zero motion vector.
[0064]
[0072] As shown in Figure 2, there are multiple spatially neighboring blocks to the current CU of the picture. These spatially neighboring blocks include those near the left edge of the current CU and those near the top edge of the current CU. The spatially neighboring blocks have left-right and top-down relationships with respect to the current CU, as shown in Figure 2. As illustrated in the example in Figure 2, the list of merge candidates for a picture coded according to the merge mode may include the following merge candidates:
[0065]
[0073] Block (A0) is spatially adjacent to the left side.
[0066]
[0074] Block (B0) that is spatially close to the upper side,
[0067]
[0075] Block (B1) located in the spatial vicinity on the upper right side,
[0068]
[0076] Block (A1) is spatially adjacent to the lower left side, and
[0069]
[0077] Block (B2) is spatially adjacent to the upper left side.
[0070]
[0078] Of the spatially adjacent blocks shown herein, block A0 is the block to the left of the current CU, block A1 is the block to the left of the current CU, block B0 is the block above the current CU, block B1 is the block above the current CU, and block B2 is the block above the current CU. The relative positioning of each spatially adjacent block with respect to the current CU or to each other is not limited beyond these relationships, and there are no limitations on the relative size of each spatially adjacent block with respect to the current CU or to each other.
[0071]
[0079] VVC standard encoders and decoders implement the deriving of at most four merge candidates by searching for spatially neighboring blocks to the left of the current CU and by searching for spatially neighboring blocks above the current CU. These spatially neighboring blocks may be searched in the order B0, A0, B1, A1, and B2. Any of these spatially neighboring blocks may be available for the merge candidate list, provided they do not belong to another slice or tile. Thus, B2 is only added to the merge candidate list if none of the other four spatially neighboring blocks are available or are intracoded.
[0072]
[0080] For each spatially adjacent block found to be available, a merge candidate is derived from the movement of that spatially adjacent block and added to the merge candidate list. If further candidates are found after candidate A1 has been added in this manner, the VVC standard encoder and VVC standard decoder further implement redundancy checking. Candidates that contain the same movement information as another candidate should not be added to the list. However, to reduce computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, as shown in Figure 3, only pairs linked by arrows are considered, and a candidate is added to the list only if the corresponding candidate used for the redundancy check does not have the same movement information.
[0073]
[0081] Next, only one time merge candidate is added to the list. In detail, the derivation of this time merge candidate derives the scaled motion vector based on the collated CU belonging to the collated reference picture. The VVC standard encoder implements explicit signaling of the reference picture list and reference index that should be used for the derivation of the collated CU in the slice header.
[0074]
[0082] Please understand that the VVC standard defines a "collocated picture" as a picture that has the same spatial resolution, the same scaling window offset, the same number of subpictures, and the same CTU size as the current picture.
[0075]
[0083] Figure 4 shows, by dotted lines, how to obtain a scaled motion vector for a time merge candidate according to the VVC standard. The scaled motion vector is scaled from the motion vector of the collated CU using picture order count ("POC") distances, tb and td, where tb represents the POC difference between the current picture and its reference picture, and td represents the POC difference between the collated picture and its reference picture. The reference picture index of the time merge candidate is set to equal to 0.
[0076]
[0084] When deriving time merge candidates, VVC standard encoders and VVC standard decoders implement the deriving of a scaled motion vector from either the L0 motion vector or the L1 motion vector of the collated CU, and it should be understood that either the L0 motion vector or the L1 motion vector of the collated CU is determined according to the following process.
[0077]
[0085] If the motion vector of the collated CU is a bipredictive motion vector and the current picture is a low-latency picture, the L0 motion vector of the TMVP candidate is scaled from the L0 motion vector of the collated CU, and the L1 motion vector of the TMVP candidate is scaled from the L1 motion vector of the collated CU.
[0078]
[0086] Instead, if the motion vector of the collated CU is a bipredictive motion vector and the current picture is a non-low-latency picture, the VVC standard encoder and VVC standard decoder implement determining one of the two motion vectors of the collated CU as the basis for scaling, according to the reference picture list of the collated CU. More specifically, if the collated CU is from the L0 reference picture list, both the L0 and L1 motion vectors of the TMVP candidate are scaled from the L1 motion vector of the collated CU. Similarly, if the collated CU is from the L1 reference picture list, both the L0 and L1 motion vectors of the TMVP candidate are scaled from the L0 motion vector of the collated CU.
[0079]
[0087] Instead, if the motion vector of the collated CU is the L0 predicted motion vector, both the L0 motion vector and L1 motion vector of the TMVP candidate are scaled from the L0 motion vector of the collated CU, regardless of whether the current picture is a low-latency picture or not. Similarly, if the motion vector of the collated CU is the L1 predicted motion vector, both the L0 motion vector and L1 motion vector of the TMVP candidate are scaled from the L1 motion vector of the collated CU.
[0080]
[0088] Figure 5 shows the selection of a position for a time candidate between candidate C0 and candidate C1, where the blocks outlined with solid lines indicate the location of the current CU according to the VVC standard. If the collated CU at position C0 is unavailable, intracoded, or outside the current row of CTUs, the VVC standard encoder and VVC standard decoder implement the use of the collated CU at position C1 to derive a time merge candidate. Otherwise, the VVC standard encoder and VVC standard decoder implement the use of the collated CU at position C0 to derive a time merge candidate.
[0081]
[0089] Therefore, it should be understood that, according to the VVC standard, the time candidate is derived from either a collated CU located relative to the lower right corner of the current CU, or a collated CU located relative to the center of the current CU.
[0082]
[0090] Next, VVC standard encoders and VVC standard decoders implement the addition of history-based MVP ("HMVP") merge candidates after spatial MVP candidates and TMVP candidates in the merge candidate list. Hereinafter, motion information for previously coded blocks is stored in a table and used as an MVP candidate for the current CU. A table with multiple HMVP candidates is maintained during the coding / decoding process. The table is reset (emptied) when a new CTU row is encountered. Whenever there is an intercoded CU that is not a subunit, the associated motion information is added to the last entry in the table as a new HMVP candidate.
[0083]
[0091] The HMVP table size S is set to 6, which indicates that a maximum of 5 HMVP candidates can be added to the table. When inserting a new move candidate into the table, the VVC standard encoder and VVC standard decoder implement restricted first-in, first-out ("FIFO") processing, and redundancy checking is first applied to find if there are any identical HMVPs in the table. If found, the identical HMVP is removed from the table, all subsequent HMVP candidates are moved forward, and the identical HMVP is inserted as the last entry in the table.
[0084]
[0092] HMVP candidates can be used in the merge candidate list construction process. The most recent HMVP candidates in the table are examined in order and inserted after TMVP candidates in the candidate list. Redundancy checks are applied to HMVP candidates for spatial or temporal merge candidates.
[0085]
[0093] To reduce the number of redundant check operations, the following simplifications are introduced.
[0086]
[0094] For the A1 and B1 space candidates, the last two entries in the table are checked for redundancy.
[0087]
[0095] The process of building the merge candidate list from HMVP is terminated when the total number of available merge candidates reaches one less than the maximum allowed number of merge candidates.
[0088]
[0096] Next, the VVC standard encoder and VVC standard decoder implement the generation of a pairwise average candidate by averaging predefined pairs of candidates in an existing merge candidate list using the first two merge candidates. The first merge candidate may be defined as p0Cand, and the second merge candidate as p1Cand. The averaged motion vector is calculated separately for each reference list, according to the availability of the motion vectors for p0Cand and p1Cand. If both motion vectors are available in a single list, these two motion vectors are averaged even when they point to different reference pictures, and its reference picture is set to the motion vector for p0Cand. If only one motion vector is available, that motion vector is used directly. If no motion vector is available, this list is kept invalid. Also, if the half-per interpolation filter indices for p0Cand and p1Cand are different, it is set to 0.
[0089]
[0097] Finally, if the merge list is not full after the pairwise average merge candidates are added, zero MVPs are inserted at the end positions until the maximum number of merge candidates is reached. A zero motion vector may have a motion shift of (0,0).
[0090]
[0098] While the VVC standard provides a merge candidate list of at most six candidates, JVET's ongoing work in this area (presented at the 133rd meeting of the Moving Picture Expert Group (“MPEG”) in January 2021 as “Exploration experiment on enhanced compression beyond VVC capability” and at the 136th meeting of MPEG (“Algorithm description of Enhanced Compression Model 3 (ECM3)”) goes beyond the scope of the VVC standard and proposes an expanded merge candidate list of at least 15 candidates, including, in order:
[0091]
[0099] Candidates for the spatially nearest CU to the current CU,
[0092]
[0100] Current MVP candidates from the collocated CUs in CU,
[0093]
[0101] Non-adjacent space candidates,
[0094]
[0102] History-based MVP candidates from FIFO tables
[0095]
[0103] Pairwise average MVP candidates, and
[0096]
[0104] Zero motion vector.
[0097]
[0105] Figure 6 shows possible spatial neighbor blocks according to ECM3, from which not only adjacent spatial merge candidates but also non-adjacent spatial merge candidates can be derived. Non-adjacent spatial merge candidates are usually inserted after TMVP candidates in the merge candidate list. The distance between a non-adjacent spatial candidate and the current coding block is based on the width and height of the current coding block. Line buffer limitations do not apply.
[0098]
[0106] Furthermore, after the merge candidate list is constructed, the merge candidates are sorted (according to an adaptive sort of merge candidates, hereafter referred to as "ARMC"). The merge candidates are initially divided into several subgroups. The subgroup size is set to 5 in normal merge mode and TM merge mode. The subgroup size is set to 3 in affine merge mode. The merge candidates within each subgroup are sorted in ascending order according to their cost values based on template matching. For simplicity, merge candidates in the last subgroup but not in the first subgroup are not sorted. The template matching cost of a merge candidate is measured by the absolute difference sum ("SAD") between the template samples in the current block and their corresponding reference samples. A template contains a set of reconstructed samples in the neighborhood of the current block. The reference samples of a template are positioned by the motion information of the merge candidates.
[0099]
[0107] While the ECM3 merge candidate list search technique focuses on expanding the scope of merge candidate search, neither the VVC standard nor the ECM proposal exhibits improved performance from the TMVP technique, as TMVP candidates continue to occupy only one position in the merge candidate list. Consequently, TMVP candidates are increasingly likely to underperform compared to other merge candidates. It is desirable to improve TMVP performance so that it remains competitive with merge candidates based on other motion prediction techniques in the merge candidate list.
[0100]
[0108] Accordingly, exemplary embodiments of this disclosure provide a time-motion vector prediction candidate selection method that offers improvements to VVC and ECM in several respects.
[0101]
[0109] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a time motion vector prediction candidate selection method utilizing the repositioning of a collated CTU.
[0102]
[0110] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a time-motion vector prediction candidate selection method that utilizes an expanded selection range.
[0103]
[0111] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a time motion vector prediction candidate selection method utilizing the unconditional derivation of scaled motion vectors.
[0104]
[0112] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a time motion vector prediction candidate selection method that omits scaling single prediction motion vectors to biprediction motion vectors.
[0105]
[0113] In one or more embodiments of the present disclosure, exemplary embodiments of the present disclosure provide VVC standard encoders and VVC standard decoders that implement a time motion vector prediction candidate selection method that utilizes multiple options when setting a reference picture index.
[0106]
[0114] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a time-motion vector prediction candidate selection method utilizing a scaling factor offsetting.
[0107]
[0115] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a merge candidate list creation method for omitting time motion vector prediction candidates.
[0108]
[0116] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a picture reconstruction method utilizing motion information refinement.
[0109]
[0117] Each of the above-described aspects of the exemplary embodiments of this disclosure will be described in further detail below.
[0110]
[0118] According to the ECM design, in order to minimize the on-chip buffer size of time motion, the time motion vector may be obtained from only a collated CTU and one column positioned to the right of the collated CTU, where the collated CTU is a CTU in a collated reference picture whose position is the same as the position of the current CTU. However, this design is not suitable for sequences with high-speed motion or for pictures where the collated reference picture is far away (i.e., the POC distance between the picture and its collated reference picture is large). Accordingly, exemplary embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that allow the time motion vector to be derived from a position other than the collated CTU.
[0111]
[0119] Figures 7A and 7B show a time motion vector prediction candidate selection method using ECM and a time motion vector prediction candidate selection method utilizing the repositioning of a collated CTU, respectively, according to an exemplary embodiment of the present disclosure. First, the block division of the picture effectively divides the picture into a grid of multiple blocks, some grids of the picture may have a larger block size, and others may have a smaller block size. For each grid of the picture, motion vectors are signaled to indicate where the time motion of the grid comes from, and such grids are therefore called “motion grids” for brevity, as the grid granularity determines the distribution of motion vectors.
[0112]
[0120] Figures 7A and 7B illustrate an example where the motion grid size of the current block and the motion grid size of the collated block are equal to the size of the CTU, with the left figure showing the ECM proposal and the right figure showing the present disclosure. According to exemplary embodiments of the present disclosure, VVC standard encoders and VVC standard decoders implement the repositioning of a collated CTU according to a signaled motion vector 702. In other words, the VVC standard encoders and VVC standard decoders derive a TMVP candidate for the current CU of the current CTU 704 in the current picture 706 from a repositioned collated CTU 708 in the collated picture 710, according to a motion vector 702 (which may not necessarily be signaled, as will be described later), and the repositioned collated CTU 708 is positioned in the collated picture 710 by the motion vector 702 for the current CTU 704 in the current picture 706.
[0113]
[0121] It should be understood that “repositioning” or “repositioning” the collated CTU 708 in the exemplary embodiments of this disclosure does not mean that the picture data of the collated picture 710 is moved. Rather, “repositioning” or “repositioning” the collated CTU 708 should be understood as the VVC standard encoder and VVC standard decoder performing the operations described herein when using the collated CTU 708. The collated CTU 708 does not need to be positioned in the collated picture 710 in the same way as the current CTU 704 in the current picture 706; instead, another CTU in the collated picture 710 (which does not need to be, but may be positioned differently relative to the current CTU 704) is used in place of the collated CTU 708 when performing the operations described above.
[0114]
[0122] Alternatively, the motion grid size of the current block and the motion grid size of the collated block may differ from the size of the CTU. In one example, the motion grid size could be N×N, where N is equal to 256, 128, 64, 32, or 16 luma samples. In another example, the motion grid size could be N×M, where N is not equal to M, and both N and M are integer powers of 2.
[0115]
[0123] The motion grid size varies depending on the sequence, time layer, or picture type difference, which can result in different block divisions. According to some exemplary embodiments, VVC standard encoders and VVC standard decoders implement signaling the motion grid size of the current block in a sequence-level, picture-level, or slice-level syntax structure. According to other exemplary embodiments, the motion grid size is adjusted according to the time layer.
[0116]
[0124] To understand this disclosure, it should be understood that when VVC standard encoders and VVC standard decoders implement signaling parameters in a syntax structure, the encoder implements recording parameters in a syntax structure such as a block, picture, sequence, or slice, transmitting the coded syntax structure over a bitstream, and the encoder implements parsing the coded syntax structure from the bitstream.
[0117]
[0125] VVC standard encoders and decoders can implement signaling for the motion grid size of the highest time layer in the sequence-level syntax structure, or they can implement signaling for a fixed motion grid size of the highest time layer, such as the same size as the CTU. In the latter case, the VVC standard encoders and decoders implement reducing the motion grid size for all lower time layers because the POC distance between a picture and its collated reference picture is larger in lower time layers. Thus, motion is more complex and finer in lower time layers than in higher time layers, and therefore the accuracy of time motion prediction can be improved with finer granularity.
[0118]
[0126] Furthermore, VVC standard encoders and decoders can implement signaling for the motion grid size of the lowest time layer in the syntax structure and increase the motion grid size for any higher time layer.
[0119]
[0127] To signal the motion vector for each motion grid, the motion vector can be directly signaled or predicted. According to some exemplary embodiments, the motion vector of the current motion grid can be merged from any of its neighboring motion grids (for example, a motion grid to the left or above, based on raster scan order coding, or any other neighboring motion grid in the block that was previously coded according to other scan orders, as described below). VVC standard encoders and VVC standard decoders implement signaling a parameter (e.g., a flag or index) in the syntax structure to indicate whether the motion vector of the current motion grid is the same as the motion vector of its neighboring motion grids. If the signaled parameter indicates identity, the motion vector of the current motion grid is not signaled and directly inherits from its neighboring motion grids. Otherwise, the motion vector of the current motion grid is signaled in the syntax structure.
[0120]
[0128] Alternatively, the motion vector of the current motion grid can be predicted from neighboring motion grids. VVC standard encoders and decoders implement the use of motion vectors from neighboring motion grids as motion vector predictors for the current motion grid. Instead of signaling the parameters described above, only the motion vector difference is signaled in the syntax structure.
[0121]
[0129] Furthermore, if motion vectors for each motion grid are signaled in the syntax structure, VVC standard encoders and VVC standard decoders implement coding of motion grids according to a default order, where the default order may be one of the raster scan order, z-sequence scan order, horizontal scan order, vertical scan order, and diagonal scan order.
[0122]
[0130] It should be understood that a collated CTU may be in the best position for the time motion vector such that the signaled motion vector is equal to zero motion. Therefore, to reduce signaling overhead, VVC standard encoders and VVC standard decoders implement signaling in the syntax structure a control parameter that determines whether the position of the collated CTU is repositioned from the current CTU.
[0123]
[0131] According to other exemplary embodiments, VVC standard encoders and VVC standard decoders implement signaling sequence-level, picture-level, or slice-level syntax structure flags to indicate whether the position of a collated CTU is to be repositioned from the current CTU.
[0124]
[0132] According to other exemplary embodiments, the position of the collated CTU in the higher time layer remains invariant from the current CTU, and there is no need to signal other parameters.
[0125]
[0133] In one or more embodiments of the present disclosure, exemplary embodiments provide a time-motion vector prediction candidate selection method that utilizes an expanded selection range.
[0126]
[0134] According to the VVC and ECM described above with reference to Figure 5, the position for the time candidate can be selected only from candidates C0 and C1 shown in Figure 5. Thus, exemplary embodiments of this disclosure provide selection of time candidates from additional positions.
[0127]
[0135] According to some exemplary embodiments, VVC standard encoders and VVC standard decoders implement the selection of a time candidate from a set of combinations of C0 and C1 ("time candidates") shown in Figure 5 and A0, A1, B0, B1, and B2 ("spatial candidates") shown in Figure 2 (i.e., spatially neighboring blocks as described with reference to Figure 2, with respect to the position of the current CU shown in Figure 5).
[0128]
[0136] According to some exemplary embodiments, VVC standard encoders and VVC standard decoders implement the selection of time candidates according to a default order. The default order is C0, then C1, then B0, then A0, then A1, then B1, then B2. If the CU at position C0 is unavailable, intracoded, or outside the current row of the CTU, position C1 is checked. Otherwise, position C0 is used in the derivation of the TMVP candidate, and the search ends. Similarly, if the CU at position C1 is unavailable, intracoded, or outside the current row of the CTU, position B0 is checked, and so on. The default order according to exemplary embodiments of this disclosure is not limited and may be any combination of C0, C1, A0, A1, B0, B1, B2.
[0129]
[0137] According to other exemplary embodiments, VVC standard encoders and VVC standard decoders implement the deriving of a TMVP candidate by averaging all time motion vectors of time candidates. The time motion vectors of time candidates are normalized by scaling to a fixed reference picture, and then the normalized time motion vectors are averaged.
[0130]
[0138] According to other exemplary embodiments, VVC standard encoders and VVC standard decoders implement the derivation of TMVP candidates by comparing time candidate time motion vectors with spatial merge candidates. The time motion vector that produces the largest motion vector difference when compared is used in the derivation of the TMVP candidate.
[0131]
[0139] According to other exemplary embodiments, VVC standard encoders and VVC standard decoders implement the selection of TMVP candidates from time motion vectors of time candidates according to the respective cost values of template matching. The template matching cost of time candidates is measured by the SAD between the template samples of the current block and their corresponding reference samples. The template includes a set of reconstructed samples in the neighborhood of the current block. The reference samples of the template are positioned by the motion information of the time candidates.
[0132]
[0140] According to other exemplary embodiments, instead of indirectly determining where the TMVP candidates are selected from, the derivation of TMVP candidates may be explicitly signaled. VVC standard encoders and VVC standard decoders implement obtaining time candidates from a set containing {C0, C1, A0, A1, B0, B1, B2}. The time candidates are {C0, C1, A0, A1, B0, B 1, It can be any subset of B2. Time candidates can also be obtained from any position within the collated CTU.
[0133]
[0141] VVC standard encoders and decoders implement signaling indices in the syntax structure of each CTU to identify the derivation of TMVP candidates. For example, a signaled index of 0 identifies that the TMVP candidate is derived from position C0, a signaled index of 1 identifies that the TMVP candidate is derived from position C1, and so on. Indices can be signaled in syntax structures of different granularities, such as sequence level, picture level, slice level, 64x64 grid level, 32x32 grid level, and 16x16 grid level.
[0134]
[0142] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a time motion vector prediction candidate selection method utilizing the unconditional derivation of scaled motion vectors.
[0135]
[0143] According to ECM, the time-motion vector is scaled from either L0 time-motion or L1 time-motion, conditionally depending on whether the current picture is a low-latency picture. According to exemplary embodiments of this disclosure, this condition may be omitted, and therefore, when deriving time-merge candidates, the VVC standard encoder and VVC standard decoder implement deriving a scaled motion vector from one of the motion vectors of the collated CU. The one of the motion vectors of the collated CU is determined according to the following steps:
[0136]
[0144] If the motion vector of the collated CU is a bipredictive motion vector, the L0 motion vector of the TMVP candidate is scaled from the L0 motion vector of the collated CU, and the L1 motion vector of the TMVP candidate is scaled from the L1 motion vector of the collated CU, regardless of whether the current picture is a low-latency picture or not.
[0137]
[0145] Instead, if the motion vector of the collated CU is the L0 predicted motion vector, both the L0 motion vector and L1 motion vector of the TMVP candidate are scaled from the L0 motion of the collated CU, regardless of whether the current picture is a low-latency picture or not. Similarly, if the motion vector of the collated CU is the L1 predicted motion vector, both the L0 motion vector and L1 motion vector of the TMVP candidate are scaled from the L1 motion vector of the collated CU.
[0138]
[0146] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a time motion vector prediction candidate selection method that omits scaling single prediction motion vectors to biprediction motion vectors.
[0139]
[0147] According to ECM, the motion vector of a TMVP candidate is always derived from the bipredictive motion, regardless of whether the motion of the collated CU is single-predictive or bipredictive. Scaling the single-predictive motion vector to the bipredictive motion vector is not preferable because the scaling process is not accurate.
[0140]
[0148] According to some exemplary embodiments, it is proposed to omit the scaling process of converting a single-prediction motion vector to a double-prediction motion vector. The VVC standard encoder and VVC standard decoder implement the deriving of a scaled motion vector from one of the motion vectors of a collated CU when deriving a time merge candidate. The one of the motion vectors of a collated CU is determined according to the following steps:
[0141]
[0149] If the motion vector of the collated CU is a bipredictive motion vector and the current picture is a low-latency picture, the L0 motion vector of the TMVP candidate is scaled from the L0 motion vector of the collated CU, and the L1 motion vector of the TMVP candidate is scaled from the L1 motion vector of the collated CU.
[0142]
[0150] Instead, if the motion of the collated CU is a bipredictive motion vector and the current picture is a non-low-latency picture, then which of the two motion vectors of the collated CU is used to perform scaling is determined according to the collated CU's reference picture list. That is, if the collated CU is from the L0 reference picture list, then both the L0 and L1 motion vectors of the TMVP candidate are scaled from the L1 motion vector of the collated CU. Similarly, if the collated CU is from the L1 reference picture list, then both the L0 and L1 motion vectors of the TMVP candidate are scaled from the L0 motion vector of the collated CU.
[0143]
[0151] Instead, if the motion vector of the collated CU is the L0 predicted motion vector, the L0 motion vector of the TMVP candidate is scaled from the L0 motion vector of the collated CU, regardless of whether the current picture is a low-latency picture or not, while the L1 motion vector of the TMVP candidate is set to be unavailable. Similarly, if the motion of the collated CU is the L1 predicted motion, the L1 motion vector of the TMVP candidate is scaled from the L1 motion vector of the collated CU, while the L0 motion vector of the TMVP candidate is set to be unavailable.
[0144]
[0152] According to other exemplary embodiments, the scaling process for converting single-predictive motion to double-predictive motion may be omitted only for the lowest time layer, rather than for all time layers. For example, scaling is omitted when the time layer is lower than layer 3.
[0145]
[0153] According to other exemplary embodiments, the scaling process for converting single-prediction motion to double-prediction motion may be omitted only for the lowest time layer and only for non-low-latency pictures.
[0146]
[0154] According to other exemplary embodiments, the scaling process for converting single-prediction motion to dual-prediction motion is omitted for only a few merge modes, which may be any, some, or all of the following: normal merge mode, merge with MVD, geometric segmentation mode, combined inter and intra modes, subblock-based time motion vector prediction, affine merge mode, and template matching mode.
[0147]
[0155] According to other exemplary embodiments, whether the motion vector of a TMVP candidate is derived by single-prediction motion or by bi-prediction motion is determined according to the template matching cost value. VVC standard encoders and VVC standard decoders implement measuring the template matching cost by the sum of absolute differences between the template samples of the current block and their corresponding reference samples. The template includes a set of reconstructed samples in the neighborhood of the current block. The reference samples of the template are positioned by the L0-prediction motion information, L1-prediction motion information, and bi-prediction motion information of the time candidate.
[0148]
[0156] It should be understood that template matching to determine whether a TMVP candidate is a single-prediction motion vector or a bi-prediction motion vector is performed only when ARMC is enabled. Furthermore, to simplify the implementation, when constructing a merge candidate list, the VVC standard encoder and VVC standard decoder first implement scaling the TMVP candidates to bi-prediction motion vectors. Then, when ARMC is applied, the TMVP candidates may be converted to single-prediction motion vectors based on the cost value of the template matching.
[0149]
[0157] According to other exemplary embodiments, VVC standard encoders and VVC standard decoders implement the addition of additional single-prediction TMVP candidates when the motion of a collated CU is single-prediction.
[0150]
[0158] In one or more embodiments of the present disclosure, exemplary embodiments of the present disclosure provide VVC standard encoders and VVC standard decoders that implement a time motion vector prediction candidate selection method that utilizes multiple options when setting a reference picture index.
[0151]
[0159] According to ECM, the reference picture index for time merge candidates is set to equal to 0. According to exemplary embodiments of this disclosure, different reference picture indices may be selected.
[0152]
[0160] According to some exemplary embodiments, the selected reference picture index is the reference picture index of the collated picture whose scaling factor (i.e., tb / td shown in Figure 4) is closest to 1.
[0153]
[0161] According to other exemplary embodiments, the selected reference picture index is the reference picture index most frequently selected for spatially neighboring blocks. Spatially neighboring blocks may be spatial candidates, HMVP candidates, or non-adjacent candidates, as described above with respect to the VVC standard and ECM.
[0154]
[0162] According to other exemplary embodiments, VVC standard encoders and VVC standard decoders implement signaling of a reference picture index in sequence-level, picture-level, slice-level, or CTU-level syntax structures.
[0155]
[0163] According to other exemplary embodiments, VVC standard encoders and VVC standard decoders implement the selection of a different reference picture index for each subblock when the blocks are coded using subblock-based time motion vector prediction ("SbTMVP") mode.
[0156]
[0164] According to other exemplary embodiments, VVC standard encoders and VVC standard decoders implement determining a consensus reference picture index for each subblock based on the reference picture index selection for each subblock when a block is coded using SbTMVP mode. For each subblock, a reference picture index is first selected, and the selected reference picture index is the reference index of the collated picture whose scaling factor is closest to 1. In that case, the reference picture index for the entire block is the reference picture index that is most frequently selected among the subblocks.
[0157]
[0165] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a time-motion vector prediction candidate selection method utilizing a scaling factor offsetting.
[0158]
[0166] According to the ECM, the scaling factor is calculated using the POC distance, as described above with respect to the VVC standard and the ECM, but the scaling factor calculation is found to be inaccurate. Exemplary embodiments of this disclosure provide a VVC standard encoder and a VVC standard decoder that implement a scaling factor offset to improve accuracy.
[0159]
[0167] According to some exemplary embodiments, the scaling factor may be offset as follows:
[0160]
number
[0161]
[0168] In this specification, tb represents the POC difference between the current picture and its reference picture, td represents the POC difference between the colocated picture and its reference picture, and N represents a non-zero integer (for example, N is equal to ±8, ±16). Assuming a negative N, the scaling factor is adjusted to be smaller. Assuming a positive N, the scaling factor is adjusted to be larger.
[0162]
[0169] According to other exemplary embodiments, the offset scaling factor described above is affected to be closer to 1. Assuming a scaling factor less than 1, the N defined above is set to a positive number. Assuming a scaling factor greater than 1, N is set to a negative number.
[0163]
[0170] According to other exemplary embodiments, VVC standard encoders and VVC standard decoders implement signaling an offset (i.e., a number N) in sequence-level, picture-level, slice-level, or CTU-level syntax structures.
[0164]
[0171] For example, when signaling an offset, both the absolute value and the sign of the offset are signaled.
[0165]
[0172] As another example, when signaling an offset, only the sign of the offset is signaled. The absolute value is fixed to a default number.
[0166]
[0173] As another example, the absolute value and sign of an offset can be signaled at different levels: the absolute value is signaled in the sequence-level syntax structure, and the sign is signaled in the CTU-level syntax structure.
[0167]
[0174] According to other exemplary embodiments, the scaling factor may or may not be offset individually for each CU. For each CU, the scaling factor is:
[0168]
number
[0169]
[0175] According to other exemplary embodiments, the scaling factor may be offset for each time layer; that is, each time layer may have a different scaling factor offset.
[0170]
[0176] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a merge candidate list creation method for omitting time motion vector prediction candidates.
[0171]
[0177] According to ECM, time merge candidates are added to the normal merge mode, geometric segmentation mode ("GPM"), merge mode with MVD ("MMVD"), combination of inter and intra prediction ("CIIP"), SbTMVP, and affine mode. It has been observed that time merge candidates are not always ultimately used in coding.
[0172]
[0178] According to some exemplary embodiments, VVC standard encoders and VVC standard decoders implement conditional addition of time merge candidates to a merge candidate list based on the motion information of neighboring blocks. When the time motion of a neighboring block is similar to the time motion of the current block, and the neighboring block's motion cannot be obtained from the time motion, the TMVP candidate for the current block is treated as unavailable for addition to the merge candidate list.
[0173]
[0179] Figure 8 illustrates how, according to an exemplary embodiment of the present disclosure, time merge candidates are conditionally added to the merge candidate list according to the motion information of neighboring blocks. In Figure 8, motion vectors
[0174]
number
[0175]
number
[0176]
number
[0177]
number
[0178]
number
[0179]
number
[0180]
number
[0181]
[0180] A neighborhood block can be any subset of {A0, A1, B0, B1, B2}. A neighborhood block can also be a candidate for a non-adjacent spatial merge or an HMVP merge.
[0182]
[0181] The similarity between the time motion vector of the current block and the time motion vector of a neighboring block is compared to a default threshold. When the motion vector difference is smaller than the default threshold, the two time motion vectors are treated as similar.
[0183]
[0182] The default threshold is an integer greater than 0. The default threshold may be set to a different value depending on the coding mode of the current block or the size of the current block. For example, the default threshold is normally set to 1 in merge mode and to 16 in template matching mode.
[0184]
[0183] According to other exemplary embodiments, the VVC standard encoder and VVC standard decoder set the adaptive merge list construction order according to the time layer, picture type (e.g., low-latency picture or non-low-latency picture), or coding mode of the current CU. In one example, for higher time layers, TMVP candidates have a higher priority, which ensures that TMVP candidates are added preferentially before spatial merge candidates, thereby overriding the merge candidate order described above.
[0185]
[0184] In one or more embodiments of the present disclosure, exemplary embodiments provide VVC standard encoders and VVC standard decoders that implement a picture reconstruction method utilizing motion information refinement.
[0186]
[0185] According to some exemplary embodiments, after the current block coded in intermode is reconstructed, motion information is refined, including the interpretation direction (i.e., L0 prediction, L1 prediction, or biprediction), a reference picture index, and a motion vector. The reconstructed sample of the current block is used as a template and motion estimation is performed.
[0187]
[0186] When performing motion estimation, only strain is considered. The refined motion information is then used as time motion for future coding pictures. The reconstructed samples used in the motion estimation process may be samples before or after the loop filtering process.
[0188]
[0187] According to some exemplary embodiments, when constructing merge candidate lists for normal merge mode, CIIP, GPM, MMVD and template matching mode, the VVC standard encoder and VVC standard decoder implement treating the TMVP candidate for the current block as unavailable if the time motion of a neighboring block is similar to the time motion of the current block and the motion of the neighboring block is not obtained from the time motion. If a TMVP candidate is added to the merge candidate list, the TMVP candidate is derived as follows:
[0189]
[0188] If the motion of the collated CU is bipredictive motion, the L0 motion vector of the TMVP candidate is scaled from the L0 motion vector of the collated CU, and the L1 motion vector of the TMVP candidate is scaled from the L1 motion vector of the collated CU, regardless of whether the current picture is a low-latency picture or not.
[0190]
[0189] Instead, if the motion vector of the collated CU is the L0 predicted motion vector, both the L0 motion vector and L1 motion vector of the TMVP candidate are scaled from the L0 motion of the collated CU, regardless of whether the current picture is a low-latency picture or not. Similarly, if the motion vector of the collated CU is the L1 predicted motion vector, both the L0 motion vector and L1 motion vector of the TMVP candidate are scaled from the L1 motion vector of the collated CU.
[0191]
[0190] Furthermore, the VVC standard encoder and VVC standard decoder implement setting the reference picture index of the TMVP candidate to the reference picture index of the collated picture whose scaling factor (i.e., tb / td shown in Figure 4) is closest to 1. Furthermore, when ARMC is enabled and the current picture is a non-low-latency picture, the VVC standard encoder and VVC standard decoder implement applying a template matching method to determine whether the TMVP candidate is single-predictive or bi-predictive, and to select the optimal offset or non-offset scaling factor.
[0192]
[0191] A person skilled in the art will recognize that all of the above aspects of the Disclosure may be implemented simultaneously in any combination thereof, and that all aspects of the Disclosure may be implemented in combination as another embodiment of the Disclosure.
[0193]
[0192] Figure 9 shows an exemplary system 900 for implementing the processes and methods described above for implementing refined time-motion candidate behaviors.
[0194]
[0193] The techniques and mechanisms described herein may be implemented by multiple instances of System 900, as well as by any other computing devices, systems, and / or environments. System 900 shown in Figure 9 is merely an example of a system and does not imply any limitation on the scope of use or functionality of any computing device used to perform the processes and / or procedures described above. Other well-known computing devices, systems, environments, and / or configurations that may be suitable for use with the embodiments include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, game consoles, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and implementations using field-programmable gate arrays ("FPGAs") and application-specific integrated circuits ("ASICs").
[0195]
[0194] The system 900 may include one or more processors 902 and a system memory 904 communicably coupled to one or more processors 902. One or more processors 902 may execute one or more modules and / or processes to cause one or more processors 902 to perform various functions. In some embodiments, one or more processors 902 may include a central processing unit ("CPU"), a graphics processing unit ("GPU"), both a CPU and a GPU, or other processing units or components known in the art. Furthermore, each of the one or more processors 902 may have its own local memory, which may also store program modules, program data, and / or one or more operating systems.
[0196]
[0195] Depending on the exact configuration and type of system 900, the system memory 904 may be volatile, such as RAM; non-volatile, such as ROM; flash memory; a small hard drive; a memory card; or any combination thereof. The system memory 904 may include one or more computer executable modules 906 that can be run by one or more processors 902.
[0197]
[0196] Module 906 may include, but is not limited to, one or more of the encoder 908 and the decoder 910.
[0198]
[0197] The encoder 908 may be a VVC standard encoder that implements some, some, or all aspects of any, some exemplary embodiments of the present disclosure described above, and can be run by one or more processors 902 to configure one or more processors 902 to perform the operations described above.
[0199]
[0198] The decoder 910 may be a VVC standard encoder that implements some, some, or all aspects of any, some exemplary embodiments of the present disclosure described above, and is executable by one or more processors 902 to configure one or more processors 902 to perform the operations described above.
[0200]
[0199] The system 900 may further include an input / output (I / O) interface 940 for receiving image source data and bitstream data, and for outputting the reconstructed picture to a reference picture buffer or DBP and / or display buffer. The system 900 may also include a communications module 950 that enables the system 900 to communicate with other devices (not shown) over a network (not shown). The network may include the internet, a wired medium (such as a wired network or direct wired connection), or a wireless medium (such as acoustic, radio frequency ("RF"), infrared, and other wireless media).
[0201]
[0200] Some or all of the operations of the methods described above may be carried out by the execution of computer-readable instructions stored in a computer-readable storage medium, as defined below. As used herein and in the claims, the term “computer-readable instructions” includes routines, applications, application modules, program modules, programs, components, data structures, algorithms, etc. Computer-readable instructions may be implemented on a variety of system configurations, including single-processor or multi-processor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based systems, programmable consumer electronics, and combinations thereof.
[0202]
[0201] Computer-readable storage media may include volatile memory (such as random access memory ("RAM")) and / or non-volatile memory (such as read-only memory ("ROM"), flash memory, etc.). Computer-readable storage media may also include, but are not limited to, additional removable storage and / or non-removable storage, including flash memory, magnetic storage, optical storage, and / or tape storage that can provide non-volatile storage such as computer-readable instructions, data structures, program modules, etc.
[0203]
[0202] Non-temporary computer-readable storage media are an example of computer-readable media. Computer-readable media include at least two types of computer-readable media, namely computer-readable storage media and communication media. Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented in any process or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, phase-change memory ("PRAM"), static random-access memory ("SRAM"), dynamic random-access memory ("DRAM"), other types of random-access memory ("RAM"), read-only memory ("ROM"), electrically erasable programmable read-only memory ("EEPROM"), flash memory or other memory technologies, compact disc read-only memory ("CD-ROM"), digital versatile disc ("DVD") or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that may be used to store information for access by computing devices. In contrast, communication media may embody computer-readable instructions, data structures, program modules or other data in modulated data signals, such as carrier waves, or in other transmission mechanisms. The computer-readable storage media used herein shall not be interpreted as transient signals themselves, such as radio waves or other free-propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (such as light pulses passing through fiber optic cables), or electrical signals propagating through wires.
[0204]
[0203] Computer-readable instructions stored in one or more non-temporary computer-readable storage media, wherein when the computer-readable instructions are executed by one or more processors, they can perform the operations described above with reference to Figures 1A to 8. Generally, computer-readable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular abstract data type. The order in which the operations are described is not to be interpreted as limiting, and any number of described operations can be combined in any order and / or in parallel to implement a process.
[0205]
[0204] While the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter as defined in the attached claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms of implementing the claims.
[0206]
[0205] Exemplary embodiments of the present disclosure are further described by at least the following clauses.
[0207]
[0206] A. A method comprising selecting a plurality of motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein one or more processors (902) derive a time motion vector prediction candidate ("TMVP candidate") for the current CU of the current CTU (704) of the current picture (706) from a repositioned collated CTU (708) of a collated picture (710) according to a motion vector (702) of the motion grid of the current picture (706), the repositioned collated CTU (708) being positioned in the collated picture (710) by the motion vector (702) for the current CTU (704) in the current picture (706).
[0208]
[0207] B. The method of paragraph A, further comprising signaling the motion grid size of the current picture (706) in a sequence-level, picture-level, or slice-level syntax structure, where the motion grid size of the current CU and collated CU differs from the sizes of the current CTU(704) and collated CTU(708).
[0209]
[0208] C. The method of paragraph B, wherein signaling the motion grid size of the current CU includes signaling the motion grid size of the highest time layer of the current picture (706) in a sequence-level syntax structure.
[0210]
[0209] D. The method of paragraph B, wherein signaling the grid size of the current CU includes decreasing the grid size for every lower time layer of the current picture (706), starting from the highest time layer of the current picture (706).
[0211]
[0210] E. The method of paragraph B, wherein signaling the grid size of the current CU includes increasing the grid size for every higher time layer of the current picture (706), starting from the lowest time layer of the current picture (706).
[0212]
[0211] F. The method according to paragraph A, wherein the motion vector (702) is not signaled in the syntax structure, and a parameter indicating the identity of the motion vector (702) with another motion vector in a neighboring motion grid is signaled in the syntax structure.
[0213]
[0212] G. The method according to paragraph A, wherein the motion vector (702) is not signaled in the syntax structure, and a parameter indicating the difference between the motion vector (702) and another motion vector in a neighboring motion grid is signaled in the syntax structure.
[0214]
[0213] H. The method according to paragraph A, wherein motion vectors (702) are signaled in the syntax structure, and blocks of motion grid are coded according to a default order of one of the raster scan order, z-sequence scan order, horizontal scan order, vertical scan order, and diagonal scan order.
[0215]
[0214] I. The method of paragraph A, further comprising signaling a parameter in a sequence-level, picture-level, or slice-level syntax structure flag indicating that the position of a collated CTU(708) is to be repositioned from the current CTU(704).
[0216]
[0215] J. The method of paragraph A, wherein the sequence-level, picture-level, or slice-level syntax structure flag does not signal a parameter indicating that the position of the collated CTU(708) is to be repositioned from the current CTU(704).
[0217]
[0216] K. A method comprising selecting a plurality of motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein one or more processors (902) select a time motion vector prediction candidate ("TMVP candidate") for the current CU from a set of motion candidates that include blocks in a spatially adjacent area of a collated picture (710).
[0218]
[0217] L. The method according to paragraph K, wherein one or more processors (902) derive TMVP candidates by averaging the time motion vectors of multiple time candidates.
[0219]
[0218] M. The method according to paragraph K, wherein one or more processors (902) derive a TMVP candidate by one of several time motion vectors of time candidates which produces the largest motion vector difference compared to a spatial merge candidate.
[0220]
[0219] N. The method according to paragraph K, wherein one or more processors (902) derive TMVP candidates by the motion vector of the time candidate having the lowest cost value of template matching, the cost value of template matching includes the absolute difference sum between a sample of the template in the current block and a corresponding reference sample in the neighborhood of the current block.
[0221]
[0220] O. The method according to paragraph K, further comprising signaling an index identifying a derived TMVP candidate in the syntax structure of the CTU by one or more processors (902), wherein the syntax structure is at the sequence level, picture level, slice level, 64x64 grid level, 32x32 grid level, or 16x16 grid level.
[0222]
[0221] P. The process involves one or more processors (902) of the computing system (900) selecting a number of motion candidates for the current coding unit ("CU") of the current picture (706), wherein one or more processors (902) select time motion vector prediction candidates ("TMVP candidates") by deriving scaled motion vectors from the motion vectors of the collated CU, wherein the scaled motion vectors include an L0 motion vector and an L1 motion vector, and for the bipredicted motion vectors of the collated CU, the L0 motion vector is collated regardless of whether the current picture (706) is a low-latency picture. A method in which the L0 motion vector of a collated CU is scaled from the L0 motion vector of the collated CU, the L1 motion vector is scaled from the L1 motion vector of the collated CU, and for the L0 predicted motion vector of the collated CU, both the L0 motion vector and the L1 motion vector are scaled from the L0 motion vector of the collated CU, regardless of whether the current picture (706) is a low-latency picture, and for the L1 predicted motion vector of the collated CU, both the L0 motion vector and the L1 motion vector are scaled from the L1 motion vector of the collated CU, regardless of whether the current picture (706) is a low-latency picture.
[0223]
[0222] Q. A method comprising selecting a number of motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein one or more processors (902) select time motion vector prediction candidates ("TMVP candidates") by deriving scaled motion vectors from the motion vectors of a collated CU, wherein the scaled motion vectors include an L0 motion vector and an L1 motion vector, and for the dual prediction motion vectors of the collated CU, if the current picture (706) is a low-latency picture, the L0 motion vector is scaled from the L0 motion vector of the collated CU and the L1 motion vector is scaled from the L1 motion vector of the collated CU. Or, if the collated CU is from an L0 reference picture list and the current picture (706) is a low-latency picture, both the L0 motion vector and the L1 motion vector are scaled from the L0 motion vector of the collated CU. Alternatively, if the colocated CU is from the L1 reference picture list and the current picture (706) is a low-latency picture, both the L0 motion vector and the L1 motion vector are scaled from the L1 motion vector of the colocated CU. For the L0 predicted motion vector of the colocated CU, regardless of whether the current picture (706) is a low-latency picture, the L0 motion vector is scaled from the L0 motion vector of the colocated CU and the L1 motion vector is set to unavailable. For the L1 predicted motion vector of the colocated CU, regardless of whether the current picture (706) is a low-latency picture, the L0 motion vector is set to unavailable and the L1 motion vector is scaled from the L1 motion vector of the colocated CU.
[0224]
[0223] R. The method described in paragraph Q, where the current CU's time layer is lower than layer 3.
[0225]
[0224] S. A method comprising selecting a number of motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein one or more processors (902) select time motion vector prediction candidates ("TMVP candidates") by deriving scaled motion vectors from the motion vectors of a collated CU, wherein the scaled motion vectors include an L0 motion vector and an L1 motion vector, and for the dual prediction motion vectors of the collated CU, if the current picture (706) is a non-low-latency picture, the L0 motion vector is scaled from the L0 motion vector of the collated CU and the L1 motion vector is scaled from the L1 motion vector of the collated CU. Or, if the collated CU is from an L0 reference picture list and the current picture (706) is a low-latency picture, both the L0 motion vector and the L1 motion vector are scaled from the L0 motion vector of the collated CU. Alternatively, if the colocated CU is from the L1 reference picture list and the current picture (706) is a non-low-latency picture, both the L0 motion vector and the L1 motion vector are scaled from the L1 motion vector of the colocated CU. For the L0 predicted motion vector of the colocated CU, if the current picture (706) is a non-low-latency picture, the L0 motion vector is scaled from the L0 motion vector of the colocated CU and the L1 motion vector is set to unavailable. For the L1 predicted motion vector of the colocated CU, if the current picture (706) is a non-low-latency picture, the L0 motion vector is set to unavailable and the L1 motion vector is scaled from the L1 motion vector of the colocated CU.
[0226]
[0225] T. The method described in paragraph S, where the current time layer of CU is lower than layer 3.
[0227]
[0226] U. A method in any one of paragraphs Q, R, S, or T, wherein multiple motion candidates are selected for one of the following modes: normal merge mode, merge with MVD, geometric segmentation mode, combined inter and intra modes, subblock-based time motion vector prediction, affine merge mode, and template matching mode.
[0228]
[0227] V. A method comprising selecting a plurality of motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein one or more processors (902) select time motion vector prediction candidates ("TMVP candidates") by deriving a scaled motion vector from the motion vector of a collated CU by either single prediction motion or bi-prediction motion, depending on the lowest cost value of template matching, wherein the cost value of template matching includes the absolute difference sum between a sample of the template of the current block and a corresponding reference sample in the neighborhood of the current block.
[0229]
[0228] The method described in paragraph V, wherein the W. TMVP candidate is scaled to a bipredictive motion vector, and then the bipredictive motion vector is converted to a singlepredictive motion vector to derive the scaled motion vector.
[0230]
[0229] X. The method described in paragraph V, wherein the TMVP candidate includes two TMVP candidates each derived by a single predictive movement.
[0231]
[0230] Y. A method comprising selecting a number of motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein one or more processors (902) select a time motion vector prediction candidate ("TMVP candidate") and set the reference picture index of the TVMP candidate to a non-zero value.
[0232]
[0231] Z. The method according to paragraph Y, wherein the reference picture index includes the reference picture index of the collated picture whose scaling factor is closest to 1.
[0233]
[0232] AA. The method of paragraph Y, wherein the reference picture index includes one of several reference picture indices that is most frequently selected for a spatially nearby block.
[0234]
[0233] AB. The method according to paragraph Y, further comprising signaling a reference picture index in a sequence-level, picture-level, slice-level, or CTU-level syntax structure by one or more processors (902).
[0235]
[0234] AC. The method according to paragraph Y, wherein the current block is coded using subblock-based time motion vector prediction ("SbTMVP") mode, and the method further comprises selecting at least several different reference picture indices for different subblocks of the current block.
[0236]
[0235] AD. The method described in paragraph Y, wherein the current block is coded using subblock-based time-motion vector prediction ("SbTMVP") mode, and the method further comprises: selecting a reference picture index of a collated picture whose scaling factor is closest to 1 for each subblock of the current block; and determining the reference picture index that was most frequently selected for the subblocks of the current block as the consensus reference picture index for the current block.
[0237]
[0236] AE. A method comprising selecting a plurality of motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein one or more processors (902) select time motion vector prediction candidates ("TMVP candidates") by a scaling factor, wherein the scaling factor is offset by adding an offset to a scaling factor.
[0238]
[0237] AF. The method of paragraph AE, further comprising signaling an offset in a sequence-level, picture-level, slice-level, or CTU-level syntax structure.
[0239]
[0238] AG. A method for paragraph AF in which both the absolute value and the sign of the offset are signaled.
[0240]
[0239] AH. A method for paragraph AF in which absolute value and sign are signaled at different levels of syntactic structure.
[0241]
[0240] AI. A method for paragraph AF in which the sign of the offset is signaled, but the absolute value of the offset is not signaled.
[0242]
[0241] AJ. A method of paragraph AE wherein the scaling factor is offset individually for each CU of the current picture (706) or not.
[0243]
[0242] AK. A method of paragraph AE in which the scaling factor is different for each time layer of the current picture (706).
[0244]
[0243] AL. A method comprising selecting a number of motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein one or more processors (902) does not select a time motion vector prediction candidate ("TMVP candidate") if the motion vector of the current CU is similar to the motion vector of a neighboring block of the collated CU, and the motion vector of a neighboring block of the current CU is not a scaled motion vector derived from a neighboring block of the collated CU.
[0245]
[0244] AM. The method according to paragraph AL, wherein the similarity is determined according to a similarity threshold based on at least one of the coding mode of the current block and the size of the current block.
[0246]
[0245] AN. A method comprising selecting a plurality of motion candidates for a current coding unit ("CU") of a current picture (706) by one or more processors (902) of a computing system (900), wherein the one or more processors (902) adaptively change the order in which they select the plurality of merge candidates based on at least one of the time layer of the current picture (706) and the coding mode of the current block, regardless of whether the current picture (706) is low-latency or non-low-latency.
[0247]
[0246] AO. The method according to paragraph AN, wherein one or more processors (902) preferentially select time motion vector prediction candidates ("TMVP candidates") before spatial motion vector prediction candidates for a current picture (706) having a high time layer.
[0248]
[0247] AP. A method comprising: reconstructing an interpredictively coded current block by one or more processors (902) of a computing system (900); performing motion estimation by one or more processors (902) using a sample of the reconstructed block as a template to generate refined motion information of the reconstructed block; and selecting time motion vector prediction candidates ("TMVP candidates") derived from the refined motion information of the reconstructed block by one or more processors (902).
Claims
1. A method for encoding a video sequence, Receiving a video sequence, The aforementioned video sequence, Selecting multiple merge candidates for the current coding unit ("CU") of the current picture. By doing so, A method in which a time motion vector prediction candidate ("TMVP candidate") is selected by setting the reference picture index of the TMVP candidate to the reference picture index of the collated picture whose scaling factor is closest to 1.
2. The scaling factor is The difference in picture order count ("POC") between the reference picture of the current picture and the current picture, The POC difference between the collated picture and the reference picture of the collated picture and The method according to claim 1, including a ratio between the two.
3. The encoding is, The method according to claim 1, further comprising sorting a plurality of merge candidates according to an adaptive sorting of merge candidates ("ARMC").
4. The method according to claim 1, wherein the plurality of TMVP candidates of the plurality of merge candidates are sorted by cost values based on template matching, and the template matching cost values of the merge candidates include the absolute difference sum ("SAD") between the template samples of the current block and their respective reference samples.
5. A method for signaling a bitstream, Receiving a video sequence, The aforementioned video sequence, Selecting multiple merge candidates for the current coding unit ("CU") of the current picture. By encoding, Signaling the bitstream generated based on the aforementioned encoding. A method comprising selecting a time motion vector prediction candidate ("TMVP candidate") by setting the reference picture index of the TMVP candidate to the reference picture index of a collated picture whose scaling factor is closest to 1.
6. The scaling factor is The difference in picture order count ("POC") between the reference picture of the current picture and the current picture, The POC difference between the collated picture and the reference picture of the collated picture and The method according to claim 5, including a ratio between the two.
7. Encoding the video sequence is The method according to claim 5, further comprising sorting a plurality of merge candidates according to an adaptive sorting of merge candidates ("ARMC").
8. The method according to claim 5, wherein the plurality of TMVP candidates of the plurality of merge candidates are sorted by cost values based on template matching, and the template matching cost values of the merge candidates include the absolute difference sum ("SAD") between the template samples of the current block and their respective reference samples.
9. A method for decoding a bitstream, Receiving a bitstream and Decoding the bitstream in order to output a video sequence Including the above, the decryption is Parsing the merge index signaled in the syntax structure of the current coding unit ("CU") of the current picture, The selection of a plurality of merge candidates for the CU is such that the time motion vector prediction candidate ("TMVP candidate") is selected by setting the reference picture index of the TMVP candidate to the reference picture index of the collated picture whose scaling factor is closest to 1. Interpretation of the CU is performed by applying the merge mode based on the merge index and the plurality of merge candidates. Methods that include...
10. The scaling factor is The difference in picture order count ("POC") between the reference picture of the current picture and the current picture, The POC difference between the collated picture and the reference picture of the collated picture and The method according to claim 9, which includes a ratio between the two.
11. The method according to claim 9, wherein the decoding further comprises sorting a plurality of merge candidates according to an adaptive sorting of merge candidates ("ARMC").
12. The method according to claim 9, wherein the plurality of TMVP candidates of the plurality of merge candidates are sorted by cost values based on template matching, and the template matching cost values of the merge candidates include the absolute difference sum ("SAD") between the template samples of the current block and their respective reference samples.