Improving adjacent spatial motion vector predictor candidates

By prioritizing spatial and temporal motion vector predictors from a reference bank in the construction of MVP lists, the method addresses the challenge of generating accurate motion vector predictions, thereby improving coding efficiency and compression performance.

JP2025517263APending Publication Date: 2025-06-05TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024524008
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-09
Filing Date
2022-09-23
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in generating accurate motion vector prediction (MVP) lists, which affects coding efficiency and compression performance.

Method used

The method involves obtaining motion vector candidates from a reference motion vector bank, determining the optimal position for inserting these candidates into the MVP list, and inserting them based on their accuracy and relevance, prioritizing spatial motion vector predictors (SMVPs) and temporal MV candidates.

Benefits of technology

This approach leads to improved accuracy and coding efficiency of MVP lists, enhancing the compression capabilities of video coding technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517263000001_ABST
    Figure 2025517263000001_ABST
Patent Text Reader

Abstract

A method, a device, and a non-transitory storage medium are provided for motion vector prediction (MVP) list construction for video coding. One or more MV candidates may be obtained from a reference motion vector (MV) bank, and the one or more MV candidates are associated with a current block. A position for inserting the one or more MV candidates from the reference MV bank into the MVP list associated with the current block is determined. The one or more MV candidates from the reference MV bank are inserted into the MVP list associated with the current block based on the position.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 342,739, filed on May 17, 2022, and U.S. Patent Application No. 17 / 941,513, filed on September 9, 2022, in the United States Patent and Trademark Office, the disclosures of which are incorporated by reference in their entireties into this specification.

[0002]

[0002] Embodiments of the present disclosure relate to image and video coding techniques. More particularly, embodiments of the present disclosure relate to improved generation of motion vector prediction (MVP) lists for coding and decoding video data. [Background technology]

[0003]

[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium founded in 2015 that includes semiconductor companies, video-on-demand providers, video content producers, software developers, and web browser vendors. Many of the components of the AV1 project were derived from previous research efforts by members of the Alliance. Individual contributors started experimental technology platforms many years ago. Xiph / Mozilla's Daala had already released its code in 2010, Google's experimental VP9 evolution project, VP10, was announced on September 12, 2014, and Cisco's Thor on August 11, 2015. Built on the VP9 code base, AV1 incorporates additional techniques, some of which were developed in these experimental formats. The first version 0.1.0 of the AV1 reference codec was released on April 7, 2016. The Alliance announced the release of the AV1 Bitstream Specification, along with reference, software-based encoders and decoders, on March 28, 2018. Validated version 1.0.0 of the specification was released on June 25, 2018. Validated version 1.0.0, including specification errata 1, was released on January 8, 2019. The AV1 Bitstream Specification includes reference video codecs.

[0004]

[0004] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Since then, these organizations have been studying the potential need for standardization of future video coding technologies that could significantly exceed HEVC in compression capabilities. In October 2017, these organizations issued a joint Call for Proposal (CfP) for video compression with capabilities exceeding HEVC. By February 15, 2018, a total of 22 CfP responses for standard dynamic range (SDR), 12 CfP responses for high dynamic range (HDR), and 12 CfP responses for 360 video categories had been submitted, respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th JVET (Joint Video Exploration Team - Joint Video Experts Team) Conference. After careful evaluation, JVET officially launched the standardization of the next generation video coding beyond HEVC, the so-called Versatile Video Coding (VVC). Summary of the Invention [Means for solving the problem]

[0005]

[0005] According to an embodiment, a method for motion vector prediction (MVP) list construction for video coding may be provided. The method may be executed by at least one processor and may include the steps of obtaining one or more MV candidates from a reference motion vector (MV) bank, where the one or more MV candidates include spatial motion vector predictor (SMVP) candidates neighboring a current block, determining a position for inserting the one or more MV candidates from the reference MV bank into an MVP list associated with the current block, and inserting the one or more MV candidates including the SMVP candidates neighboring the current block from the reference MV bank into an MVP list associated with the current block based on the position.

[0006]

[0006] According to an embodiment, an apparatus for motion vector prediction (MVP) list construction for video coding may be provided. The apparatus may include at least one memory configured to store a program code, and at least one processor configured to read the program code and operate as instructed by the program code. The program code may include: an acquisition code configured to cause the at least one processor to acquire one or more MV candidates from a reference motion vector (MV) bank, the one or more MV candidates including spatial motion vector prediction (SMVP) candidates adjacent to a current block; a decision code configured to cause the at least one processor to determine a position to insert the one or more MV candidates from the reference MV bank into an MVP list associated with the current block; and an insertion code configured to cause the at least one processor to insert the one or more MV candidates including the SMVP candidates adjacent to the current block from the reference MV bank into an MVP list associated with the current block based on a position.

[0007]

[0007] According to an embodiment, a non-transitory computer-readable medium may be provided that stores instructions. The instructions, when executed by at least one processor of a device that performs motion vector prediction (MVP) list construction for video coding, may cause the at least one processor to obtain one or more MV candidates from a reference motion vector (MV) bank, where the one or more MV candidates include spatial motion vector predictor (SMVP) candidates neighboring a current block, determine a position for inserting the one or more MV candidates from the reference MV bank into an MVP list associated with the current block, and insert the one or more MV candidates from the reference MV bank, including the SMVP candidates neighboring the current block, into an MVP list associated with the current block based on the position. [Brief description of the drawings]

[0008] [Figure 1A]

[0008] FIG. 1 is a diagram illustrating an example of a partition tree under AV1 and VPN frameworks according to one embodiment of the present disclosure. [Figure 1B]

[0009] FIG. 2 illustrates an example of block partitions and tree structures using quadtree and binary tree block partitioning, according to one embodiment of the present disclosure. [Figure 1C]

[0010] FIG. 1 illustrates an example of vertical central triple tree partitioning and horizontal central triple tree partitioning according to one embodiment of the present disclosure. [Figure 1D]

[0011] FIG. 13 is a diagram illustrating an example of search points for a merge mode using motion vector difference according to one embodiment of the present disclosure. [Figure 1E]

[0012] FIG. 2 illustrates an example of a spatial motion vector neighborhood, according to one embodiment of the present disclosure. [Figure 1F]

[0013] FIG. 2 illustrates an example of motion field estimation by linear projection, according to one embodiment of the present disclosure. [Figure 1G]

[0014] FIG. 2 illustrates an example of block locations for deriving a temporal motion vector predictor according to one embodiment of the present disclosure. [Figure 1H]

[0015] FIG. 2 illustrates an example of generating additional motion vector candidates for a block with a single reference, according to one embodiment of the present disclosure. [Figure 1I]

[0016] FIG. 13 illustrates an example of generating additional motion vector candidates for a block with mixed references according to one embodiment of the present disclosure. [Diagram 2]

[0017] FIG. 2 illustrates a reference motion vector candidate update process in the related art according to an embodiment of the present disclosure. [Diagram 3]

[0018] 4 is a flow chart for building a motion vector candidate list according to one embodiment of the present disclosure. [Figure 4]

[0019] FIG. 1 is a simplified block diagram of a communication system according to one embodiment of the present disclosure. [Diagram 5]

[0020] FIG. 2 illustrates an arrangement of a video encoder and a video decoder in a streaming environment. [Figure 6]

[0021] FIG. 2 is a functional block diagram of a video decoder according to one embodiment of the present disclosure. [Figure 7]

[0022] FIG. 2 is a functional block diagram of a video encoder according to one embodiment of the present disclosure. [Figure 8]

[0023] 1 is a flowchart of an example process of motion vector prediction (MVP) list construction for video coding and decoding, according to one embodiment of this disclosure. [Figure 9]

[0024] FIG. 1 is a diagram of a computer system according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009]

[0025] Aspects of the disclosed embodiments may be used individually or in combination.Embodiments of the present disclosure relate to improvements in generating, maintaining, and updating one or more motion vector prediction (MVP) lists.

[0010]

[0026] According to an embodiment of the present disclosure, since candidates from a reference motion vector (MV) candidate bank are usually more accurate, candidates from a reference MV candidate bank may be inserted before other candidates, such as temporal MV candidates, combined MV candidates (e.g., scaled MV candidates), and padded candidates, in an MVP list. Since adjacent spatial MV candidates are usually more accurate and MV candidates from a reference MV candidate bank are mainly from spatial neighborhoods, an embodiment of the present disclosure generates and maintains a more accurate MVP list. This improved accuracy leads to improved coding efficiency.

[0011]

[0027] According to an embodiment, a motion vector (MV) candidate may include an n-dimensional (e.g., 2-dimensional) vector that may be used to predict coordinates in a current picture from a reference picture. A motion vector prediction (MVP) list for a block may include multiple motion vector candidates, from neighboring blocks, and in some embodiments from non-neighboring blocks, that may be used to predict the motion vector of the block.

[0012]

[0028] According to one embodiment, candidates from the reference MV candidate bank may be inserted before additional candidates in the MVP list. As an example, the MVP list may be constructed and maintained in the following order: adjacent spatial motion vector predictors (SMVPs), temporal motion vector predictors (TMVPs), non-adjacent SMVPs, reordering process for existing candidates, derived candidates, candidates from the reference MV candidate bank, and additional candidates.

[0013]

[0029] According to another embodiment, candidates from the reference MV candidate bank may be inserted before the derived candidates in the MVP list. As an example, the MVP list may be constructed and maintained in the following order: adjacent SMVPs, TMVPs, non-adjacent SMVPs, a sorting process for existing candidates, candidates from the reference MV candidate bank, derived candidates, and additional candidates.

[0014]

[0030] In some embodiments, a candidate from the reference MV candidate bank may be inserted before a non-adjacent SMVP in the MVP. When a candidate from the reference MV candidate bank is inserted before a non-adjacent SMVP in the MVP, the sorting process based on the predefined weighting may be removed or inhibited. In such embodiments, the MVP list may be constructed in the following order: adjacent SMVPs, TMVPs, candidates from the reference MV candidate bank, non-adjacent SMVPs, derived candidates, and additional candidates.

[0015]

[0031] In some embodiments, candidates from the reference MV candidate bank may be inserted before a TMVP in the MVP. When candidates from the reference MV candidate bank are inserted before a TMVP in the MVP, the sorting process based on the predefined weighting may be removed or inhibited. In such embodiments, the MVP list may be constructed in the following order: adjacent SMVPs, candidates from the reference MV candidate bank, TMVPs, non-adjacent SMVPs, derived candidates, and additional candidates.

[0016]

[0032] According to one embodiment, a candidate from the reference MV candidate bank may only be inserted before other candidates under one or more specific conditions or one or more predefined conditions. The conditions may include, but are not limited to, whether the current block uses a single reference mode or a mixed mode, block size, block coding mode, or one or more signaled parameters. The position of the candidate from the reference MV candidate bank may include, but is not limited to, for example, before derived candidates, additional candidates, TMVP, and non-adjacent SMVP. As an example, if the current block uses a mixed mode (i.e., the current block has two reference pictures), the candidate from the reference MV candidate bank may be inserted before the derived candidate in the MVP list. In some embodiments, if the current block uses a single reference picture, the candidate from the reference MV candidate bank may still be at the end in the MVP list. As another example, if the current block uses a single reference picture, the candidate from the reference MV candidate bank may be inserted before the derived candidate in the MVP list. In some embodiments, when the current block uses mixed mode (ie, the current block has two reference pictures), the candidate from the reference MV candidate bank may still be last in the MVP list.

[0017]

[0033] Embodiments of the present disclosure relate to improving predefined weighting based sorting of candidates in an MVP list.

[0018]

[0034] In the related art, adjacent SMVP candidates, TMVP candidates, and non-adjacent candidates may be sorted according to predefined weighting of these candidates. However, additional search and newly added modes (i.e., derived candidates and candidates from the reference MV candidate bank) are not involved in the sorting process. This reduces the accuracy of the predicted motion vector of the current block, since some related candidates may not be weighted correctly in the MVP list.

[0019]

[0035] According to an embodiment of the present disclosure, the sorting process depends on the output value of a function. The input of this weighting function can be, but is not limited to, the parameters of the candidate (e.g., the block size of the candidate, the magnitude / direction of the MVP, etc.) or already constructed samples (from the reference picture and / or the current picture). This weighting function may be called a sorting score function. Then, after all the MVP candidates are inserted into the MVP list, the list can be sorted based on the output value (i.e., score) of the sorting score function to most accurately capture the relevance of the candidates.

[0020]

[0036] According to one embodiment, the input value of the reordering score function may be the block size. In some embodiments, the larger the candidate block size, the higher the score of this candidate may be. In this embodiment, the block size may also be stored in the reference MV candidate bank. According to another embodiment, the input value of the reordering score function may be the size of the MVP, so that the larger the size of the MVP that the current block has, the higher the score this block may get.

[0021]

[0037] According to one embodiment, the reordering score function may be template matching, and the reordering score may be a template matching distortion cost. As an example, a template may be generated around the reference block pointed to by the MVP, and another template may be generated around the current block, and the distortion between these two templates may be used as the reordering store. The template matching based method may be applied to both single reference and mixed reference modes. A template refers to a region of spatially adjacent samples reconstructed before the current block.

[0022]

[0038] According to one embodiment, the permutation score function may be bilateral matching, and the permutation score may be the bilateral matching distortion cost. The distortion between two predictors generated by the MVP candidate may be calculated and used as the permutation score. The bilateral matching based method may be applied to both single reference and mixed reference modes.

[0023]

[0039] It will be appreciated that any number of suitable reordering score functions may be defined, including but not limited to the reordering score functions described above. The output scores may be weighted and averaged and used in the reordering process.

[0024]

[0040] Block Partitioning in VP9 and AV1

[0041] FIG. 1A is a diagram 1100 of an example partitioning tree under VP9 and AV1. As shown in the top half of FIG. 1A, VP9 may use a 4-way partition tree starting from the 64×64 level down to the 4×4 level, with some additional restrictions for blocks 8×8. The partitions labeled "R" refer to recursive partitions, i.e., partitions where the same partition tree may be repeated at lower scales until the lowest 4×4 level is reached.

[0025]

[0042] As shown in the lower half of FIG. 1A, AV1 can not only extend the partition tree to a 10-way structure, but AV1 may also increase the maximum size (called superblock in VP9 / AV1 terminology) to start from 128×128. It will be understood that the 10-way structure may include 4:1 / 1:4 rectangular partitions, which did not exist in VP9. Rectangular partitions cannot be further subdivided. In addition, AV1 provides more flexibility for the use of partitions below the 8×8 level in the sense that 2×2 chroma inter prediction is possible in certain cases.

[0026]

[0043] Block Partitioning in HEVC

[0044] In HEVC, to accommodate various local characteristics, a coding tree unit (CTU) is partitioned into coding units (CU) by using a quad tree (QT) structure, denoted as a coding tree. The decision on whether to code a picture region using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU may be further partitioned into one, two, or four PUs according to a prediction unit (PU) partition type. Within one PU, the same prediction process is applied, and related information is sent to the decoder for each PU. After obtaining the residual block by applying a prediction process based on the PU partition type, the CU may be partitioned into transform units (TUs) according to another quad tree structure, such as a coding tree for CUs. The feature of the HEVC structure is to have multiple partition concepts, including CUs, PUs, and TUs. In HEVC, a CU or TU can only be square, while a PU may be square or rectangular for an inter-prediction block. In HEVC, one coding block may be further divided into four square sub-blocks, and a transform is performed on each sub-block, i.e., TU. Each TU may be further divided recursively (using quad-tree partitioning) into smaller TUs called Residual Quad Trees (RQTs). At picture boundaries, HEVC employs implicit quad-tree partitioning, such that a block continues quad-tree partitioning until its size fits the picture boundary. One of the key features of the HEVC structure is that it has multiple partition concepts, including CUs, PUs, and TUs.

[0027]

[0045] Block Partitioning in Versatile Video Coding (VVC)

[0046] Block partitioning structures using quad tree (QT) and binary tree (BT)

[0047] The QTBT structure may include the concept of multiple partition types, i.e., the QTBT structure may eliminate the separation of the concepts of CU, PU, ​​and TU, and support further flexibility of CU partition shapes. In the QTBT block structure, a CU may have either a square or rectangular shape. As shown in FIG. 1B using a tree 1205, a coding tree unit (CTU) is first partitioned by a quadtree structure. The quadtree leaf node is further partitioned by a binary tree structure. There are two partition types in the binary tree partition: symmetric horizontal partition and symmetric vertical partition. The binary tree leaf node is called a coding unit (CU), and its segmentation is used for prediction and transform processing without further partitioning. Therefore, the CU, PU, ​​and TU have the same block size in the QTBT coding block structure. In JEM, a CU may be composed of coding blocks (CBs) of different color components, e.g., for P slices and B slices in 4:2:0 chrominance format, one CU includes one luma CB and two chroma CBs, and sometimes includes a single component CB, e.g., for I slices, one CU includes only one luma CB or only two chroma CBs. For the QTBT partitioning scheme, the following parameters may be defined: CTU size: quad-tree root node size, which is the same concept as in HEVC; MinQTSize: minimum allowed quad-tree leaf node size; MaxBTSize: maximum allowed binary tree root node size; MaxBTDepth: maximum allowed binary tree depth; and MinBTSize: minimum allowed binary tree leaf node size.

[0028]

[0048] In one example of a QTBT partitioning structure, the CTU size may be set as 128×128 luma samples with two corresponding 64×64 blocks of chroma samples, MinQTSize may be set as 16×16, MaxBTSize may be set as 64×64, MinBTSize (both width and height) may be set to 4×4, and MaxBTDepth is set to 4. Quad-tree partitioning may first be applied to the CTU to generate quad-tree leaf nodes. The quad-tree leaf nodes may have sizes from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the leaf quad-tree node is 128×128, it is not further split by the binary tree since the size exceeds MaxBTSize (i.e., 64×64). Otherwise, the leaf quad-tree node may be further partitioned by the binary tree. Therefore, the quadtree leaf node is also the root node of the binary tree and has the binary tree depth as 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further splits are considered. When a binary tree node has a width equal to MinBTSize (i.e., 4), no further horizontal splits are considered. Similarly, when a binary tree node has a height equal to MinBTSize, no further vertical splits are considered. The leaf node of the binary tree is further processed by the prediction and transformation process without further partitioning. In JEM, the maximum CTU size is 256x256 luma samples.

[0029]

[0049] As seen in FIG. 1B, block 1205 shows an example of block partitioning by using QTBT, and binary tree 1210 in FIG. 1B shows the corresponding tree representation. Solid lines indicate quadtree partitioning, and dotted lines indicate binary tree partitioning. At each partition (i.e., non-leaf) node of the binary tree, one flag may be signaled to indicate which partition type (i.e., horizontal or vertical) is used, where 0 may indicate horizontal partitioning and 1 may indicate vertical partitioning. In the case of quadtree partitioning, there is no need to indicate the partition type, since quadtree partitioning always partitions a block both horizontally and vertically to create four sub-blocks with equal size.

[0030]

[0050] In addition, the QTBT scheme supports the flexibility for luma and chroma to have separate QTBT structures. In the current related art, for P slices and B slices, the luma CTB and chroma CTB in one CTU share the same QTBT structure. However, for I slices, the luma CTB is partitioned into CUs by one QTBT structure and the chroma CTB is partitioned into chroma CUs by another QTBT structure. This means that a CU in an I slice may be composed of a coding block of a luma component or a coding block of two chroma components, and a CU in a P slice or B slice may be composed of coding blocks of all three chroma components.

[0031]

[0051] In HEVC, inter prediction for small blocks may be restricted to reduce memory access for motion compensation, resulting in bi-prediction not being supported for 4x8 and 8x4 blocks, and inter prediction not being supported for 4x4 blocks. These restrictions have been removed in the QTBT implementation in JEM-7.0.

[0032]

[0052] Block partitioning structure using ternary tree (TT)

[0053] In VVC, a Multi-Type-Tree (MTT) structure may be included that adds horizontal and vertical center-side triple trees on top of the QTBT in blocks 1305 and 1310, respectively, as shown in FIG. 1C. The advantage of TT partitioning is that it complements quad-tree and binary-tree partitioning, and while quad-tree and binary-tree always split along block centers, TT partitioning is also capable of capturing objects located at block centers. Moreover, the width and height of the proposed TT partitions are always powers of two, so no additional transformations are required. The design of the two-level tree is primarily motivated by reduced complexity. In theory, the complexity of traversing a tree is reduced by 1 / T. D where T denotes the number of split types and D is the depth of the tree.

[0033]

[0054] Merge mode with motion vector difference (MMVD)

[0055] In addition to the merge mode in which the implicitly derived motion information is directly used to generate the prediction sample of the current CU, VVC introduces a merge mode using motion vector difference (MMVD). Immediately after sending the skip flag and the merge flag, an MMVD flag may be signaled to specify whether the MMVD mode is used for the CU.

[0034]

[0056] In MMVD, after a merge candidate is selected, the merge candidates may be further narrowed down by signaled MVD information. The further information may include a merge candidate flag, an index to specify the magnitude of motion, and an index to indicate the direction of motion. In MMVD mode, one of the first two candidates in the merge list is selected to use as the MV basis. A merge candidate flag may be signaled to specify which one is used.

[0035]

[0057] The distance index specifies the magnitude information of the motion and indicates a predefined offset from the starting point. As shown in FIG. 1D, the offset may be added to either the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1.

[0036] [Table 1]

[0037]

[0058] The direction index represents the direction of the MVD relative to the starting point. The direction index can represent the four directions shown in Table 2.

[0038] [Table 2]

[0039]

[0059] Please note that the meaning of the MVD code may change depending on the information of the starting MV. When the starting MV is a uni-predictive MV or a bi-predictive MV, and both lists point to the same side of the current picture, the code in Table 2 specifies the sign of the MV offset added to the starting MV. As an example, when the POCs of the two references are both larger than the POC of the current picture, or both smaller than the POC of the current picture. When the starting MV is a bi-predictive MV with two MVs pointing to different sides of the current picture, and the difference of the POCs in list 0 is larger than the difference of the POCs in list 1, the code in Table 2 specifies the sign of the MV offset added to the MV components of the starting MV of list 0, and the code of the MV in list 1 has an opposite value. As an example, when the POC of one reference is larger than the POC of the current picture and the POC of the other reference is smaller than the POC of the current picture, the code of the MV offset added to the MV components of the starting MV of list 0, and the code of the MV in list 1 has an opposite value. Otherwise, if the POC difference in list 1 is greater than list 0, then the signs in Table 2 specify the signs of the MV offsets added to the MV components of the starting MV of list 1, and the signs of the MVs in list 0 have the opposite value.

[0040]

[0060] The MVD may be scaled according to the difference in POC in each direction. If the difference in POC in both lists is the same, no scaling may be done. Otherwise, if the difference in POC in list 0 is greater than the difference in POC of list 1, the MVD of list 1 may be scaled. If the difference in POC of L1 is greater than L0, the MVD of list 0 may also be scaled in the same way. If the starting MV is uni-predicted, the MVD may be added to the available MV.

[0041]

[0061] Symmetric MVD coding

[0062] In VVC, in addition to the normal unidirectional and bidirectional prediction mode MVD signaling, a symmetric MVD mode for bidirectional MVD signaling may be applied. In the symmetric MVD mode, motion information including reference picture indexes of both list 0 and list 1 and the MVD of list 1 may be derived instead of being signaled.

[0042]

[0063] The decoding process of the symmetric MVD mode at the slice level may be as follows: At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived. If mvd_l1_zero_flag is 1, BiDirPredFlag may be set equal to 0. Otherwise, if the closest reference picture in list 0 and the closest reference picture in list 1 form a forward and backward pair of reference pictures or a backward and forward pair of reference pictures, BiDirPredFlag may be set to 1, and the reference pictures in both list 0 and list 1 may be short-term reference pictures. Otherwise, BiDirPredFlag is set to 0.

[0043]

[0064] The decoding process of symmetric MVD mode at the CU level may be as follows: At the CU level, if the CU is bi-predictively coded and BiDirPredFlag is equal to 1, a symmetric mode flag is explicitly signaled to indicate whether symmetric mode may be used. If the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indexes of list 0 and list 1 are set to be equal to a pair of reference pictures, respectively. MVD1 is set to be equal to (-MVD0).

[0044]

[0065] Intermode Coding in CWG-B018

[0066] In AV1, for each coded block between frames, if the mode of the current block is an inter-coding mode rather than a skip mode, a separate flag may be signaled to indicate whether a single or a mixed reference mode is used for the current block.

[0045]

[0067] A single mode may include a prediction block generated by one motion vector in a single reference mode. In the single reference case, the following modes may be signaled: (1) NEARMV - use one of the motion vector predictors (MVPs) in the list indicated by the DRL (dynamic reference list) index, (2) NEWMV - use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and apply a delta to the MVP, and (3) GLOBALMV - use a motion vector based on a frame-level global motion parameter.

[0046]

[0068] The prediction block generated by weighted averaging of two prediction blocks may be derived from two motion vectors in mixed reference mode. In the mixed reference case, the following modes may be signaled: (1) NEAR_NEARMV - use one of the motion vector predictors (MVPs) in the list signaled by the DRL index, (2) NEAR_NEWMV - use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and send a delta MV for the second MV, (3) NEW_NEARMV - use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and send a delta MV for the first MV, (4) NEW_NEWMV - use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and send a delta MV for both MVs, and (5) GLOBAL_GLOBALMV - use a MV from each reference based on the frame-level global motion parameters.

[0047]

[0069] Motion Vector Difference Coding in AV1

[0070] AV1 allows for 1 / 8 pixel motion vector precision (or accuracy), and the following syntax may be used to signal motion vector differences within list 0 or list 1 of a reference frame: (1) mv_joint specifies which components of the motion vector difference are non-zero: 0 indicates that there is no non-zero MVD along either the horizontal or vertical direction, 1 indicates that there is non-zero MVD only along the horizontal direction, 2 indicates that there is non-zero MVD only along the vertical direction, and 3 indicates that there is non-zero MVD along both the horizontal and vertical directions, (2) mv_sign specifies whether the motion vector difference is positive or negative, and (3) mv_class specifies the class of the motion vector difference. As shown in Table 3, a higher class means a larger magnitude of the motion vector difference, (4) mv_bit specifies the integer part of the offset between the motion vector difference and the starting magnitude of each MV class, (5) mv_fr specifies the first two fractional bits of the motion vector difference, and (6) mv_hp specifies the third fractional bit of the motion vector difference.

[0048] [Table 3]

[0049]

[0071] Adaptive MVD Resolution in CWG-B092

[0072] For NEW_NEARMV and NEAR_NEWMV modes, the precision of the MVD may depend on the associated class and the magnitude of the MVD. First, fractional MVD is only allowed if the magnitude of the MVD is 1 pixel or less. Second, only one MVD value is allowed if the value of the associated MV class is MV_CLASS_1 or greater. The MVD value in each MV class is derived as 4, 8, 16, 32, 64 for MV class 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5). In addition, one context may be used to signal the mv_joint or mv_class if the current block is coded as NEW_NEARMV or NEAR_NEWMV mode. Otherwise, another context may be used to signal the mv_joint or mv_class.

[0050] [Table 4]

[0051]

[0073] Joint MVD coding (JMVD) in CWG-B092

[0074] A new inter-coding mode named JOINT_NEWMV may be applied to indicate whether the MVDs of the two reference lists are jointly signaled. When the inter-prediction mode is equal to the JOINT_NEWMV mode, the MVDs of reference list 0 and reference list 1 are jointly signaled. Thus, only one MVD named joint_mvd is signaled and sent to the decoder, and the delta MVs of reference list 0 and reference list 1 are derived from joint_mvd.

[0052]

[0075] The JOINT_NEWMV mode is signaled along with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. No further context is added.

[0053]

[0076] When JOINT_NEWMV mode is signaled and the POC distances between the two reference frames and the current frame are different, the MVD is scaled for reference list 0 or reference list 1 based on the POC distance. Specifically, the distance between reference frame list 0 and the current frame can be td0, and the distance between reference frame list 1 and the current frame is denoted as td1. If td0 is greater than or equal to td1, then joint_mvd is used as is for reference list 0, and the mvd of reference list 1 is derived from joint_mvd based on equation (1).

[0054]

number

[0055]

[0077] Otherwise, if td1 is greater than or equal to td0, then for reference list 1, joint_mvd is used as is, and mvd for reference list 0 is derived from joint_mvd based on equation (2).

[0056]

number

[0057]

[0078] Improvement of Adaptive MVD Resolution in CWG-C011

[0079] A new inter-coding mode named AMVDMV may be added to the single reference case. When the AMVDMV mode is selected, it indicates that AMVD is applied to signal MVD. One flag named amvd_flag is added under the JOINT_NEWMV mode to indicate whether AMVD is applied to the joint MVD coding mode. When the adaptive MVD resolution is applied to the joint MVD coding mode named joint AMVD coding, the MVDs of two reference frames may be jointly signaled, and the accuracy of the MVD is implicitly determined by the size of the MVD. Otherwise, the MVDs of two (or more) reference frames are jointly signaled, and conventional MVD coding is applied.

[0058]

[0080] Adaptive motion vector resolution (AMVR) in CWG-C012 and CWG-C020

[0081] AMVR was first proposed in CWG-C012, and supports a total of seven MV precisions: 8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8. For each prediction block, the AVM encoder may explore all supported precision values ​​and signal the best precision to the decoder.

[0059]

[0082] To reduce the encoder runtime, two precision sets are supported. Each precision set contains four predefined precisions. The precision set may be adaptively selected at the frame level based on the maximum precision value of the frame. Similar to AV1, the maximum precision may be signaled in the frame header. Table 5 summarizes the supported precision values ​​based on the frame-level maximum precision.

[0060] [Table 5]

[0061]

[0083] Current AVM software (similar to AV1) has a frame-level flag that indicates whether the MV of a frame contains sub-pel precision or not. AMVR is only enabled if the value of the cur_frame_force_integer_mv flag is 0. In AMVR, if the precision of a block is less than the maximum precision, the motion model and the interpolation filter may not be signaled. If the precision of a block is less than the maximum precision, the motion mode may be inferred to be translational motion and the interpolation filter may be inferred to be a REGULAR interpolation filter. Similarly, if the precision of a block is 4-pel or 8-pel, the inter-intra mode is not signaled and is inferred to be 0.

[0062]

[0084] Motion Vector Predictor List in AV1 and AVM

[0085] Spatial motion vector predictors (SMVPs, both adjacent and non-adjacent SMVPs), temporal motion vector predictors, additional MV candidates and additional derived MVPs in AV1, and reference bank MVPs are added to the AVM design. To store the MVPs, a fixed-size stack known as a motion vector predictor list is generated on both the encoder and decoder sides.

[0063]

[0086] Spatial Motion Vector Predictor (SMVP)

[0087] A spatial motion vector (MV) predictor is derived from spatial neighboring blocks, including adjacent spatial neighboring blocks that are direct neighbors above and to the left of the current block, and non-adjacent spatial neighboring blocks that are close to the current block but not directly adjacent to the current block. An example of a set of spatial neighboring blocks for a luminance block is shown in FIG. 1E, where each spatial neighboring block is an 8×8 block. The spatial neighboring blocks are examined to find one or more MVs associated with the same reference frame index as the current block. For the current block, the search order of the 8×8 luminance blocks in the spatial neighborhood is as shown by numbers 1 to 8 in FIG. 5. (1) the top adjacent row is checked from left to right, (2) the left adjacent column is checked from top to bottom, (3) the top right neighbor block is checked, (4) the top left neighbor block is checked, (5) the top non-adjacent 1st row is checked from left to right, (6) the left non-adjacent 1st column is checked from top to bottom, (7) the top non-adjacent 2nd row is checked from left to right, and (8) the left non-adjacent 2nd column is checked from top to bottom.

[0064]

[0088] Adjacent candidates (1-3 in Fig. 1E) are put into the MV predictor list before TMVP, and non-adjacent candidates (also known as outer candidates, i.e. candidates 4-8 in Fig. 1E) are put into the MV predictor list after TMVP. All SMVP candidates should have the same reference picture as the current block. That is, if the current block has a single reference picture, and an MVP candidate with a single reference picture and this reference picture is the same as the reference picture of the current block, or an MVP candidate with a mixed reference picture (two reference pictures) and one of the reference pictures is the same as the reference picture of the current block, this MVP candidate will be put into the MV predictor list. If the current block has two reference pictures, then only an MVP candidate with two reference pictures and these two reference pictures are the same as the reference picture of the current block will be put into the predictor list.

[0065]

[0089] Temporal motion vector predictor (TMVP)

[0090] In addition to spatially neighboring blocks, a MV predictor, known as a temporal MV predictor, may also be derived using co-located blocks in a reference frame. To generate a temporal MV predictor, first, the MVs of the reference frames are stored together with the reference indexes associated with the respective reference frames. Then, for each 8x8 block of the current frame, the MVs of the reference frames whose trajectories pass through the 8x8 block are identified and stored together with the reference frame index in the temporal MV buffer. In the case of inter prediction using a single reference frame, the MVs are stored in 8x8 units regardless of whether the reference frame is a forward reference frame or a backward reference frame to perform temporal motion vector prediction of future frames. In the case of hybrid inter prediction, only the forward MVs are stored in 8x8 units to perform temporal motion vector prediction of future frames.

[0066]

[0091] Referring to FIG. 1G, the MV of reference frame 1 (R1) 1620, i.e., MVref 1650, points out from frame 1 (R1) 1620. In doing so, MVref 1650 goes through an 8×8 block (blocks in current frame 1615 and blocks in reference frame 0 1610 of the current frame). MVref is stored in the temporal MV buffer in association with this 8×8 block. During the motion projection process to derive the temporal MV predictor, the reference frames may be scanned in a predefined order, i.e., LAST_FRAME, BWDREF_FRAME, ALTREF_FRAME, ALTREF2_FRAME, and LAST2_FRAME. MVs from higher indexed reference frames (in scan order) do not replace previously identified MVs assigned by lower indexed reference frames (in scan order).

[0067]

[0092] Given predefined block coordinates, the associated MVs stored in the temporal MV buffer are identified and projected onto the current block to derive a temporal MV predictor that points from the current block to its reference frame, e.g., MV0 in Figure 1F.

[0068]

[0093] Referring to Figure 1G, predefined block positions for deriving the temporal MV predictor of a 16x16 block are shown. Up to seven blocks are checked for valid temporal MV predictors. The temporal MV predictor is checked after adjacent spatial MV predictors and before non-adjacent spatial MV predictors.

[0069]

[0094] In deriving the MV predictor, all spatial and temporal MV candidates may be pooled and each predictor may be assigned a weight determined during scanning of spatial and temporal neighboring blocks. The candidates may be sorted and ranked based on their associated weights, and up to four candidates are identified and added to the MV predictor list. This list of MV predictors, also called the Dynamic Reference List (DRL), is further used in the dynamic MV prediction mode, as described in the next subsection.

[0070]

[0095] Additional exploration MVPs for additional MVP candidates

[0096] If the MVP list is still not full, additional searches are performed and additional MVP candidates are used to fill the MVP list, including, for example, global MVs, zero MVs, combined composite MVs without scaling, etc.

[0071]

[0097] MVP candidate sorting process

[0098] The adjacent SMVP candidates, TMVP candidates, and non-adjacent SMVP candidates added in the MVP list are reordered. Based on the current design in AV1 and AVM, the reordering process is based on the weight of each candidate. The candidate weight is predefined according to the overlapping area between the current block and the candidate block.

[0072]

[0099] Derived MVP candidates

[0100] The derived MVP candidates have been adopted in the AVM reference software by proposal CWG-B049, which includes both MVPs derived for single reference pictures and MVPs derived for mixed modes.

[0073]

[0101] Single Inter Prediction

[0102] If the reference frame of a neighboring block is different from that of the current block, but these reference frames are in the same direction, a temporal scaling algorithm can be used to scale its MV to its reference frame to form the MVP of the motion vector of the current block. As shown in Figure 1H, mv1 1855 from the neighboring blocks (shaded blocks) can be used to derive the MVP of the motion vector mv0 1850 of the current block with temporal scaling.

[0074]

[0103] Combined Inter Prediction

[0104] To derive the MVP of the current block, the synthesized MVs from different neighboring blocks are utilized, but the reference frames of the synthesized MVs need to be the same as the current block. As shown in Figure 1I, the synthesized MVs (mv2 1960, mv3 1965) have the same reference frames as the current block, but these reference frames are from different neighboring blocks.

[0075]

[0105] Reference motion vector candidate bank

[0106] Each buffer corresponds to a unique reference frame type, which corresponds to a single reference frame or a pair of reference frames covering single and mixed inter modes, respectively. All buffers are the same size. When a new MV is added to a full buffer, existing MVs may be evicted to make space for the new MV.

[0076]

[0107] A coding block may refer to the MV candidate bank to collect reference MV candidates, in addition to the reference MV candidates obtained in the conventional AV1 reference MV list generation. After coding a superblock, the MV bank may be updated with the MVs used by the coding block of the superblock.

[0077]

[0108] Each tile may have an independent MV reference bank that may be utilized by all superblocks in the tile. At the beginning of encoding each tile, the corresponding bank may be emptied. MVs from the bank may then be used as MV reference candidates when coding each superblock in that tile. At the end of encoding a superblock, the bank may be updated.

[0078]

[0109] Bank Update

[0110] As shown in diagram 200 of Figure 2, the bank update process may be based on a superblock. That is, after a superblock is coded, the first (up to 64) candidate MVs used by each coding block in the superblock are added to the bank. During the update, a pruning process may also be involved during the update.

[0079]

[0111] Bank Reference

[0112] After scanning the traditional AV1 or new AV2 reference MV candidates, if there are free slots in the candidate list, the codec may refer to the MV candidate bank (in the buffer with a matching reference frame type) for additional MV candidates. Working backwards from the end of the buffer to the beginning, an MV in the bank buffer may be added to the candidate list if it is not already present in the candidate list.

[0080]

[0113] The MVP list building process in cutting-edge design

[0114] In the related art, an MVP list may be constructed using full pruning, as shown in flowchart 300 of Fig. 3. An example may include operation 305 including adding adjacent SMVPs, operation 310 including adding TMVPs, operation 315 including adding non-adjacent SMVPs, operation 320 including adding a sorting process for existing candidates, operation 325 including adding derived candidates, operation 330 including adding additional MVPs, and finally operation 355 including adding candidates from a reference MV candidate bank.

[0081]

[0115] FIG. 4 shows a simplified block diagram of a communication system 400 according to an embodiment of the present disclosure. The communication system 400 may include at least two terminals 410-420 interconnected via a network 450. In the case of a unidirectional transmission of data, the first terminal 410 may code video data at a local location for transmission to the other terminal 420 via the network 450. The second terminal 420 may receive the coded video data of the other terminal from the network 450, decode the coded data, and display the restored video data. The unidirectional data transmission may be common in media serving applications, etc.

[0082]

[0116] 4 shows a second pair of terminals 430, 440 provided to support two-way transmission of coded video, such as may occur during a video conference. For two-way transmission of data, each terminal 430, 440 may code captured video data at a local location for transmission to the other terminal over network 450. Each terminal 430, 440 may also receive coded video data transmitted by the other terminal, decode the coded data, and display the recovered video data on a local display device.

[0083]

[0117] In FIG. 4, terminals 410-440 are shown as servers, personal computers, and smartphones, although the principles of the present disclosure are not so limited. Embodiments of the present disclosure find use with laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 450 represents any number of networks that convey coded video data between terminals 410-440, including, for example, wired and / or wireless communication networks. Communications network 450 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 450 may not be important to the operation of the present disclosure, unless described below.

[0084]

[0118] 5 illustrates, as an example of an application of the disclosed subject matter, an arrangement of a video encoder and a video decoder in a streaming environment, e.g., streaming system 500. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0085]

[0119] The streaming system can include a capture subsystem 513, which can include a video source 501, e.g., a digital camera, that creates, e.g., an uncompressed video sample stream 502. The sample stream 502 is illustrated as a thick line to emphasize its large amount of data when compared to an encoded video bitstream and can be processed by an encoder 503 coupled to the camera 501. The encoder 503 can include hardware, software, or combinations thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream 504 is illustrated as a thin line to emphasize its small amount of data when compared to the sample stream and can be stored on a streaming server 505 for future use. One or more streaming clients 506, 508 can access the streaming server 505 to obtain copies 507, 509 of the encoded video bitstream 504. The client 506 may include a video decoder 510 that decodes a copy of the incoming encoded video bitstream 507 and creates an outgoing video sample stream 511 that may be rendered on a display 512 or other rendering device not shown. In some streaming systems, the video bitstreams 504, 507, 509 may be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. A video coding standard informally known as Versatile Video Coding (VVC) is under development. The disclosed subject matter may be used in the context of VVC.

[0086]

[0120] FIG. 6 may be a functional block diagram of a video decoder 510 according to one embodiment of the present invention.

[0087]

[0121] The receiver 610 may receive one or more codec video sequences to be decoded by the decoder 510, where in the same or another embodiment, one coded video sequence is decoded at a time, where the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequences may be received from a channel 612, which may be a hardware / software link to a storage device that stores the encoded video data. The receiver 610 may receive the encoded video data along with other data, e.g., coded audio data and / or auxiliary data streams, that may be forwarded to respective using entities, not shown. The receiver 610 may separate the coded video sequences from the other data. To address network jitter, a buffer memory 615 may be coupled between the receiver 610 and the entropy decoder / analyzer 620, hereafter "analyzer." If the receiver 610 is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer 615 may not be needed or may be small. For use in a best effort packet network such as the Internet, buffer 615 may be required and may be relatively large and advantageously of an adaptable size.

[0088]

[0122] The video decoder 510 may include a parser 620 for reconstructing symbols 621 from the entropy coded video sequence. These categories of symbols include information used to manage the operation of the decoder 510 and, in some cases, information for controlling a rendering device, such as a display 512, which is not an integral part of the decoder but may be coupled to the decoder, as shown in FIG. 6. The control information for the rendering device may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment, not shown. The parser 620 may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may conform to a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser 620 may extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups may include groups of pictures GOP, pictures, tiles, slices, macroblocks, coding units CU, blocks, transform units TU, prediction units PU, etc. The entropy decoder / analyzer may also extract information from the coded video sequence, such as transform coefficients, quantization parameter QP values, motion vectors, etc.

[0089]

[0123] The analyzer 620 may perform an entropy decoding / parsing operation on the video sequence received from the buffer 615 to produce symbols 621. The analyzer 620 may receive the encoded data and selectively decode a particular symbol 621. Additionally, the analyzer 620 may determine whether a particular symbol 621 should be provided to a motion compensated prediction unit 653, a scaler / inverse transform unit 651, an intra prediction unit 652, or a loop filter 656.

[0090]

[0124] The reconstruction of symbols 621 may involve different units depending on the type of coded video picture or part thereof, such as inter-picture and intra-picture, inter-block and intra-block, and other factors. Which units are involved and how may be controlled by subgroup control information parsed by parser 620 from the coded video sequence. For simplicity, the flow of such subgroup control information between parser 620 and the following units is not shown.

[0091]

[0125] In addition to the functional blocks already mentioned, the decoder 510 may be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units may closely interact with each other and may be at least partially integrated with each other. However, the following conceptual subdivision into functional units is adequate for describing the disclosed subject matter.

[0092]

[0126] The first unit is a scalar / inverse transform unit 651. The scalar / inverse transform unit 651 receives quantized transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. as symbols 621 from the analyzer 620. The scalar / inverse transform unit 651 may output a block containing sample values ​​that may be input to an aggregator 655.

[0093]

[0127] In some cases, the output samples of the scalar / inverse transform unit 651 may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. The intra-picture prediction unit 652 may provide such prediction information. In some cases, the intra-picture prediction unit 652 uses surrounding already reconstructed information fetched from the current partially reconstructed picture 658 to generate blocks of the same size and shape as the block being reconstructed. The aggregator 655 adds, in some cases, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit 652 to the output sample information provided by the scalar / inverse transform unit 651.

[0094]

[0128] In other cases, the output samples of the scalar / inverse transform unit 651 may relate to an inter-coded, possibly motion-compensated block. In such cases, the motion compensated prediction unit 653 may access the reference picture memory 657 to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols 621 related to the block, the aggregator 655 may add these samples, referred to in this case as residual samples or residual signals, to the output of the scalar / inverse transform unit to generate output sample information. The addresses in the reference picture memory from which the motion compensation unit fetches the prediction samples may be controlled by a motion vector, available to the motion compensation unit in the form of the symbols 621, which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0095]

[0129] The output samples of aggregator 655 may be subject to various loop filtering techniques in loop filter unit 656. Video compression techniques may include in-loop filter techniques controlled by parameters available to loop filter unit 656 as symbols 621 from analyzer 620 contained in the coded video bitstream, but may also correspond to meta-information obtained during decoding of previous portions in decoding order of the coded picture or coded video sequence, as well as to previously reconstructed loop filtered sample values.

[0096]

[0130] The output of the loop filter unit 656 may be a sample stream that is output to the render device 512 as well as stored in a reference picture memory 658 for use in future inter-picture prediction.

[0097]

[0131] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture, for example by the analyzer 620, the current reference picture 658 can become part of the reference picture buffer 657, and a new current picture memory can be reallocated before starting reconstruction of the next coded picture.

[0098]

[0132] The video decoder 510 may perform decoding operations according to a given video compression technique, which may be documented in a standard such as ITU-T Rec. H.265. The coded video sequence may conform to a syntax specified by the video compression technique or standard being used in the sense that the coded video sequence conforms to the syntax of the video compression technique or standard specified in the video compression technique document or standard, in particular in a profile document therein. Compliance may also require that the complexity of the coded video sequence be within a range defined by the level of the video compression technique or standard. In some instances, the level limits the maximum picture size, the maximum frame rate, the maximum reconstructed sample rate, e.g., measured in megasamples per second, the maximum reference picture size, etc. The limits set by the level may be further limited in some instances through the specification of a Hypothetical Reference Decoder (HRD) and metadata for HRD buffer management signaled in the coded video sequence.

[0099]

[0133] In one embodiment, the receiver 610 may receive additional redundant data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder 510 to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0100]

[0134] FIG. 7 may be a functional block diagram of a video encoder 503 according to one embodiment of the present invention.

[0101]

[0135] The encoder 503 may receive video samples from a video source 501 that is not part of the encoder and that may capture the video images that are coded by the encoder 503 .

[0102]

[0136] The video source 501 may provide a source video sequence to be coded by the encoder 503 in the form of a digital video sample stream, which may be of any suitable bit depth, e.g., 8-bit, 10-bit, 12-bit, ..., any color space, e.g., BT.601 Y CrCB, RGB, ..., and any suitable sampling structure, e.g., Y CrCb 4:2:0, Y CrCb 4:4:4. In a media serving system, the video source 501 may be a storage device that stores previously prepared video. In a video conferencing system, the video source 503 may be a camera that captures local image information as a video sequence. The video data may be provided as a number of individual pictures that give motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each pixel may contain one or more samples, depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0103]

[0137] According to an embodiment, the encoder 503 may code and compress pictures of a source video sequence into a coded video sequence 743 in real-time or under any other time constraint as required by an application. Enforcing an appropriate coding rate is one of the functions of the controller 750. The controller controls and is operatively coupled to other functional units as described below. For simplicity, couplings are not shown. Parameters set by the controller may include picture skip, quantizer, rate control related parameters such as lambda value for rate distortion optimization techniques, picture size, group of pictures GOP (group of pictures) layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller 750 that may be relevant to a video encoder 503 optimized for a particular system design.

[0104]

[0138] Some video encoders operate in what those skilled in the art will easily recognize as a "coding loop." As a very simplified explanation, the coding loop consists of a part of the encoder 730, hereafter the "source coder," responsible for creating symbols based on the input picture to be coded and the reference pictures, and a local decoder 733 built into the encoder 503, which reconstructs the symbols to create sample data, which the remote decoder also creates, since the video compression techniques considered in the disclosed subject matter allow for lossless compression between the symbols and the coded video bitstream. That reconstructed sample stream is input to a reference picture memory 734. The contents of the reference picture buffer are also bit-exact between the local and remote encoders, since decoding of the symbol stream produces bit-exact results regardless of whether the decoder is local or remote. In other words, the predictive part of the encoder "sees" exactly the same sample values ​​as the decoder would "see" when using prediction during decoding, as reference picture samples. This basic principle of reference picture synchrony and the resulting drift when synchrony cannot be maintained, for example due to channel errors, is well known to those skilled in the art.

[0105]

[0139] The operation of the "local" decoder 733 may be the same as that of the "remote" decoder 510 already described in detail above in relation to Figure 6. However, with brief reference also to Figure 7, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder 745 and the analyzer 620 may be lossless, the entropy decoding portion of the decoder 510, including the channel 612, the receiver 610, the buffer 615, and the analyzer 620, may not be fully implemented in the local decoder 733.

[0106]

[0140] At this point, it can be said that any decoder technique, other than analysis / entropy decoding, present in the decoder must also necessarily be present in the corresponding encoder in substantially the same functional form. The description of the encoder technique may be omitted, since it is the inverse of the decoder technique described generically. Only in certain areas is a more detailed description required, and is provided below.

[0107]

[0141] As part of its operation, the source coder 730 may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence, designated as “reference frames.” In this manner, the coding engine 732 codes differences between pixel blocks of the input frame and pixel blocks of reference frames that may be selected as predictive references for the input frame.

[0108]

[0142] The local video decoder 733 may decode the coded video data of frames that may be designated as reference frames based on the symbols created by the source coder 730. The operation of the coding engine 732 is advantageously a lossy process. When the coded video data may be decoded in a video decoder not shown in FIG. 7, the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder 733 may replicate the decoding process that may be performed by the video decoder on the reference frames and store the reconstructed reference frames in the reference picture cache 734. In this way, the encoder 503 may locally store copies of the reconstructed reference frames that have a common content as the reconstructed reference frames obtained without transmission errors by the far-end video decoder.

[0109]

[0143] The predictor 735 may perform a prediction search for the coding engine 732. That is, for a new frame to be coded, the predictor 735 may search the reference picture memory 734 for sample data as candidate reference pixel blocks or specific metadata such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor 735 may operate on a sample block by pixel block basis to find suitable prediction references. In some instances, as determined by the search results obtained by the predictor 735, the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory 734.

[0110]

[0144] Controller 750 may manage the coding operations of video coder 730, including setting the parameters and subgroup parameters used to encode the video data.

[0111]

[0145] The output of all the aforementioned functional units may be subject to entropy coding in the entropy coder 745. The entropy coder converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0112]

[0146] The transmitter 740 may buffer the coded video sequence created by the entropy coder 745 to prepare the coded video sequence for transmission over a communication channel 760, which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter 740 may merge the coded video data from the video coder 730 with other data to be transmitted, such as coded audio data and / or auxiliary data stream sources not shown.

[0113]

[0147] The controller 750 may manage the operation of the encoder 503. During coding, the controller 750 may assign each coded picture a particular coded picture type, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as one of the following frame types:

[0114]

[0148] An intra picture, an I-picture, may be a picture that can be coded and decoded without using any other frame in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of these variations of I-pictures, as well as their respective uses and characteristics.

[0115]

[0149] A predicted picture, a P-picture, may be a picture that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0116]

[0150] A bidirectionally predicted picture, a B-picture, may be a picture that can be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, a multi-predicted picture can use more than two reference pictures and associated metadata in the reconstruction of a single block.

[0117]

[0151] A source picture may generally be spatially subdivided into a number of sample blocks, e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each, and coded block-wise. A block may be predictively coded with reference to other already coded blocks, as determined by a coding assignment applied to the respective picture of the block. For example, a block of an I picture may be non-predictively coded, or a block of an I picture may be predictively coded with reference to already coded blocks of spatial or intra prediction of the same picture. A pixel block of a P picture may be non-predictively coded via spatial or temporal prediction with reference to one previously coded reference picture. A block of a B picture may be non-predictively coded via spatial or temporal prediction with reference to one or two previously coded reference pictures.

[0118]

[0152] Video coder 503 may perform coding operations in accordance with a given video coding technique or standard, such as ITU-T Rec. H.265. In its operations, video coder 503 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0119]

[0153] In one embodiment, the transmitter 740 may transmit additional data along with the encoded video. The video coder 730 may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0120]

[0154] FIG. 8 shows an example process 800 for generating and maintaining an MVP list for a current block 800 .

[0121]

[0155] In operation 805, one or more motion vector (MV) candidates may be obtained from a reference MV bank, and the one or more obtained MV candidates are associated with the current block.

[0122]

[0156] In operation 810, a location for inserting one or more MV candidates from the reference MV bank into an MVP list associated with the current block may be determined. The location for inserting the one or more MV candidates may include one or more of: inserting before at least one additional candidate of the one or more additional candidates, inserting before at least one derived candidate of the one or more derived candidates, inserting before at least one non-contiguous SMVP of the one or more non-contiguous SMVPs, and / or inserting before at least one TMVP of the one or more TMVPs.

[0123]

[0157] At operation 815, it may be determined whether reordering of motion vectors in the MVP list is permitted. At operation 825, it may be determined whether reordering of motion vectors in the MVP list is prohibited. In some embodiments, reordering may be prohibited based on a position being before at least one non-adjacent SMVP of the one or more non-adjacent SMVPs. In other embodiments, reordering may be prohibited based on a position being before at least one TMVP of the one or more TMVPs.

[0124]

[0158] In operation 820, it may be determined whether a condition for inserting one or more motion vectors is met. In some embodiments, if the insertion of one or more MV candidates is based on the condition being met, the condition may be that the reference mode of the current block is a single reference mode. In some embodiments, the condition may be that the reference mode of the current block is a mixed reference mode. In embodiments where the condition is not met, one or more MV candidates from the reference MV bank may be inserted at the end of the MVP list associated with the current block.

[0125]

[0159] At operation 830, one or more MV candidates from the reference MV bank may be inserted into the MVP list associated with the current block based on the position.

[0126]

[0160] In operation 840, based on the determination that reordering is allowed, one or more existing MVs in the MVP list associated with the current block are reordered based on the predefined weighting. In some embodiments, based on a condition being satisfied, one or more existing MVs in the MVP list associated with the current block are reordered based on the predefined weighting. The condition may include that the reference mode of the current block is a single reference mode or a composite mode.

[0127]

[0161] According to one aspect of the disclosure, the predefined weighting may be based on a reordering score, which may be based on a block size of the current block. If the reordering score is based on a block size of the current block, the block size of the current block may be stored in the reference MV bank. In some embodiments, the reordering score may be based on a magnitude of the MVP of the current block. In some embodiments, the reordering score may be based on a distortion cost associated with template matching. In some embodiments, the reordering score may be based on a distortion cost associated with bilateral matching.

[0128]

[0162] According to one aspect, the existing MVs in the MVP list may include any MV in the MVP list. In some embodiments, the one or more existing MVs in the MVP list may include at least one of one or more neighboring spatial MV predictors (SMVPs) associated with the current block, one or more temporal MV predictors (TMNPs) associated with the current block, and one or more non-neighboring SMVPs associated with the current block. In some embodiments, the one or more existing MVs in the MVP list may further include at least one of one or more derived candidates, one or more MV candidates from a reference MV bank, or one or more additional candidates.

[0129]

[0163] Although Figure 8 illustrates example blocks of process 800, in some implementations process 800 may include additional blocks, fewer blocks than those illustrated in Figure 8, different blocks than those illustrated in Figure 8, or blocks arranged in a different manner than those illustrated in Figure 8. Additionally or alternatively, two or more of the blocks of process 800 may be performed in parallel.

[0130]

[0164] Additionally, the proposed methods may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium to perform one or more of the proposed methods.

[0131]

[0165] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 9 illustrates a computer system 900 suitable for implementing certain embodiments of the disclosed subject matter.

[0132]

[0166] The computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to produce code containing instructions that may be executed directly by a computer central processing unit (CPU), graphics processing unit (GPU), etc., or that may be executed through interpretation, microcode execution.

[0133]

[0167] The instructions may be executed on various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0134]

[0168] 9 for computer system 900 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 900.

[0135]

[0169] The computer system 900 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users through, for example, tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (speech, music, environmental sounds, etc.), images (scanned images, photographic images obtained from still image cameras, etc.), video (2D video, 3D video including stereoscopic video, etc.), etc.

[0136]

[0170] The input human interface devices may include one or more of a keyboard 901, a mouse 902, a trackpad 903, a touch screen 910, a data glove 1204, a joystick 905, a microphone 906, a scanner 907, a camera 908 (only one of each is shown).

[0137]

[0171] The computer system 900 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the senses of a human user, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via touch screen 910, data glove 1204, or joystick 905, although there may also be haptic feedback devices that do not function as input devices), audio output devices (such as speakers 909, headphones (not shown)), visual output devices (such as screens 910, including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light emitting diode (OLED) screens, each with or without touch screen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or three or more dimensional output through means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0138]

[0172] The computer system 900 may also include human accessible storage devices and their associated media, such as optical media including a CD / DVD ROM / RW 920 with CD / DVD or similar media 921, thumb drives 922, removable hard drives or solid state drives 923, legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD based devices (not shown) such as security dongles.

[0139]

[0173] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0140]

[0174] The computer system (900) may also include an interface to one or more communication networks (955). The network (955) may be, for example, a wireless network, a wired network, an optical network. The network (955) may further be a local network, a wide area network, a metropolitan network, a vehicular and industrial network, a real-time network, a delay-tolerant network, and the like. Examples of networks (955) include local area networks such as Ethernet, wireless LAN, and the like, cellular networks including GSM, 3G, 4G, 5G, LTE, and the like, TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANBus, and the like. Certain networks (955) typically require an external network interface adapter (954) that is attached to a particular general-purpose data port or peripheral bus (949) (e.g., a USB port of the computer system (900) or the like), while other networks are typically integrated into the core of the computer system (900) by attachment to the system bus as described below (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). The computer system (900) can use any of these networks (955) to communicate with other entities. Such communications can be unidirectional receive only (e.g., broadcast TV), unidirectional transmit only (e.g., CANbus to certain CANbus devices), or bidirectional with other computer systems using local or wide area digital networks. For each of these networks (955) and network interfaces (954) described above, specific protocols and protocol stacks may be used.

[0141]

[0175] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to the core 940 of the computer system 900.

[0142]

[0176] The cores 940 may include one or more central processing units (CPUs) 941, graphics processing units (GPUs) 942, specialized programmable processing units 943 in the form of field programmable gate areas (FPGAs), hardware accelerators 944 for specific tasks, etc. These devices may be connected through a system bus 1248, along with read only memory (ROM) 945, random access memory (RAM) 946, internal mass storage 947 such as an internal hard drive, solid state drive (SSD) that is not accessible to the user, etc. In some computer systems, the system bus 1248 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus 1248 or through a peripheral bus 949. Architectures for peripheral buses include Peripheral Component Interconnect (PCI), USB, etc.

[0143]

[0177] The CPU 941, GPU 942, FPGA 943, and accelerator 944 can execute certain instructions that in combination can constitute the aforementioned computer code. That computer code can be stored in ROM 945 or RAM 946. Persistent data can be stored, for example, in internal mass storage 947, while transitory data can also be stored in RAM 946. Rapid storage and retrieval from any memory device can be made possible through the use of cache memory that can be closely associated with one or more of the CPU 941, GPU 942, mass storage 947, ROM 945, RAM 946, etc.

[0144]

[0178] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.

[0145]

[0179] By way of example and not limitation, computer system 900 having the architecture, and specifically core 940, may provide functionality as a result of processor (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be the user-accessible mass storage introduced above, and media associated with specific storage of core 940 of a non-transitory nature, such as core internal mass storage 947 or ROM 945. Software implementing various embodiments of the present disclosure may be stored in such devices and executed by core 940. Computer-readable media may include one or more memory devices or chips, depending on specific needs. The software may cause core 940, and specifically the processor therein (including CPU, GPU, FPGA, etc.) to perform certain processes or certain portions of certain processes described herein, including defining data structures stored in RAM 946 and modifying such data structures according to the software-defined processes. Additionally or alternatively, a computer system may provide functionality as a result of logic being hardwired or otherwise embodied in circuitry (e.g., accelerator 944) that can operate in place of or together with software to perform particular processes or particular portions of particular processes described herein. References to software may encompass logic, and vice versa, where appropriate. References to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that store software for execution, circuitry that embodies logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0146]

[0180] While this disclosure has described several exemplary embodiments, there are alterations, substitutions, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

1. 1. A method for motion vector prediction (MVP) list construction for video coding, the method being executed by at least one processor, comprising: obtaining one or more motion vector (MV) candidates from a reference MV bank, the one or more MV candidates including spatial motion vector predictor (SMVP) candidates neighboring a current block; determining a position for inserting the one or more MV candidates from the reference MV bank into an MVP list associated with the current block; inserting the one or more MV candidates from the reference MV bank, including the SMVP candidate adjacent to the current block, into the MVP list associated with the current block based on the location; A method comprising:

2. The method further comprising: The method of claim 1 , further comprising the step of reordering one or more existing MVs in the MVP list associated with the current block based on a predefined weighting.

3. 3. The method of claim 2, wherein the one or more existing MVs in the MVP list include at least one of: one or more adjacent SMVPs associated with the current block, one or more temporal MV predictors (TMVPs) associated with the current block, and one or more non-adjacent SMVPs associated with the current block.

4. 4. The method of claim 3, wherein the one or more existing MVs in the MVP list further include at least one of one or more derived candidates, the one or more MV candidates from the reference MV bank, or one or more additional candidates.

5. The position is before at least one additional candidate of the one or more additional candidates; prior to at least one derived candidate of the one or more derived candidates; before at least one non-adjacent SMVP of said one or more non-adjacent SMVPs; or Before at least one of the one or more TMVPs The method of claim 4, wherein the step of

6. 6. The method of claim 5, further comprising prohibiting the reordering based on the location being before the at least one non-adjacent SMVP of the one or more non-adjacent SMVPs.

7. The method of claim 5 , further comprising prohibiting the reordering based on the location being before the at least one TMVP of the one or more TMVPs.

8. The method of claim 1 , wherein the insertion of the one or more MV candidates is based on the condition being satisfied, the condition being that the reference mode of the current block is a single reference mode.

9. The method of claim 1 , wherein the insertion of the one or more MV candidates is based on a condition being satisfied, the condition being that the reference mode of the current block is a mixed reference mode.

10. 10. The method of claim 9, further comprising inserting the one or more MV candidates from the reference MV bank to the end of the MVP list associated with the current block based on the condition not being satisfied.

11. The method of claim 2 , wherein the predefined weighting is based on a reordering score, the reordering score being based on a block size of the current block.

12. The method of claim 11 , further comprising storing the block size of the current block in the reference MV bank based on the reordering score based on the block size of the current block.

13. The method of claim 2 , wherein the predefined weighting is based on a reordering score, the reordering score being based on a magnitude of an MVP of the current block.

14. The method of claim 2 , wherein the predefined weightings are based on a permutation score, the permutation score being based on a distortion cost associated with template matching.

15. The method of claim 2 , wherein the predefined weightings are based on a permutation score, the permutation score being based on a distortion cost associated with bilateral matching.

16. 1. An apparatus for motion vector prediction (MVP) list construction for video coding, comprising: at least one memory configured to store program code; at least one processor configured to read the program code and to operate as instructed by the program code; wherein the program code comprises: retrieval code configured to cause the at least one processor to retrieve one or more motion vector (MV) candidates from a reference MV bank, the one or more MV candidates including spatial motion vector prediction (SMVP) candidates neighboring a current block; and a decision code configured to cause the at least one processor to determine a location for inserting the one or more MV candidates from the reference MV bank into an MVP list associated with the current block; an insertion code configured to cause the at least one processor to insert the one or more MV candidates, including the SMVP candidate adjacent to the current block, from the reference MV bank into the MVP list associated with the current block based on the position; 13. An apparatus comprising:

17. The position is before at least one additional candidate of the one or more additional candidates in the MVP list; before at least one derived candidate of one or more derived candidates in the MVP list; before at least one non-adjacent SMVP of one or more non-adjacent SMVPs in said MVP list; or before at least one of the one or more TMVPs in the MVP list 17. The apparatus of claim 16, wherein the

18. The apparatus of claim 16 , wherein the insertion of the one or more MV candidates is based on the condition being satisfied, the condition being that a reference mode of the current block is a single reference mode.

19. The apparatus of claim 18 , further comprising: based on the condition not being satisfied, inserting the one or more MV candidates from the reference MV bank to the end of the MVP list associated with the current block.

20. 1. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a device that performs motion vector prediction (MVP) list building for video coding, cause the one or more processors to: Obtaining one or more motion vector (MV) candidates from a reference MV bank, the one or more MV candidates including spatial motion vector predictor (SMVP) candidates neighboring a current block; determining a position for inserting the one or more MV candidates from the reference MV bank into an MVP list associated with the current block; inserting the one or more MV candidates, including the SMVP candidate adjacent to the current block, from the reference MV bank into the MVP list associated with the current block based on the position; A non-transitory computer-readable medium comprising one or more instructions for causing a

Citation Information

Patent Citations

  • Merge candidates for motion vector prediction for video coding

    JP2019515587A

  • Motion Vector Prediction

    JP2020523853A

  • Video decoding method, apparatus and computer program thereof

    JP2022522398A

  • Method and apparatus of reordering motion vector prediction candidate set for video coding

    WO2018205914A1