Improvement of Candidate of Adjacent Spatial Motion Vector Predictor
The interleaved motion vector prediction method addresses inefficiencies in generating MVP lists by inserting specific spatial predictors from neighboring blocks and pruning duplicates, resulting in enhanced video coding efficiency.
Patent Information
- Application Number
- JP2024521777
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-09
- Filing Date
- 2022-09-13
- Publication Date
- 2025-06-26
AI Technical Summary
Existing image and video coding techniques face challenges in efficiently generating motion vector prediction (MVP) lists, particularly in effectively utilizing spatial motion vectors from neighboring blocks.
The proposed method involves an interleaved motion vector prediction technique for video coding, where a list of MVP candidates is generated based on spatial motion vectors from neighboring and non-neighboring blocks. This method includes inserting specific spatial motion vector predictors from left and upper neighboring blocks into the list and pruning candidates with duplicate predictors.
This approach enhances the efficiency of MVP list generation by more effectively utilizing similarities between current motion vectors and neighboring predictors, leading to improved video coding performance.
Smart Images

Figure 2025519300000008 
Figure 2025519300000009 
Figure 2025519300000010
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications
[0001] This application claims priority based on U.S. Provisional Patent Application No. 63 / 349,761, filed with the United States Patent and Trademark Office on June 7, 2022, the disclosure of which is hereby incorporated by reference in its entirety.
[0002]
[0002] Embodiments of the present disclosure relate to image and video coding techniques. More particularly, embodiments of the present disclosure relate to improvements in the generation of motion vector prediction (MVP) lists.
Background Art
[0003]
[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium established in 2015 that includes semiconductor companies, video - on - demand providers, video content production companies, software development companies, and web browser vendors. Many of the components of the AV1 project were derived from previous research activities by the Alliance's members. Individual contributors had started experimental technology platforms years earlier. Xiph / Mozilla's Daala had already made its code public in 2010, Google's experimental VP9 evolution project, VP10, was announced on September 12, 2014, and Cisco's Thor was made public on August 11, 2015. AV1, built on the VP9 codebase, incorporates additional techniques, some of which were developed in these experimental formats. The first version, 0.1.0, of the AV1 reference codec was made public on April 7, 2016. The Alliance announced the release of the AV1 bitstream specification on March 28, 2018, along with a reference, software - based encoder and decoder. On June 25, 2018, a verified version 1.0.0 of the specification was released. On January 8, 2019, a verified version 1.0.0 including errata 1 of the specification was released. The AV1 bitstream specification includes a reference video codec.
[0004]
[0004] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Since then, these groups have studied the potential need for standardization of future video coding technologies that may significantly exceed HEVC in compression capabilities. In October 2017, these groups issued a Call for Proposal (CfP) for video compression with capabilities beyond HEVC. By February 15, 2018, a total of 22 CfP responses regarding standard dynamic range (SDR), 12 CfP responses regarding high dynamic range (HDR), and 12 CfP responses regarding 360 video categories were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th JVET (Joint Video Exploration Team - Joint Video Experts Team) meeting. After careful evaluation, JVET officially started the standardization of the next-generation video coding beyond HEVC, so-called Versatile Video Coding (VVC).
Summary of the Invention
Means for Solving the Problems
[0005] According to an embodiment, a method for interleaved motion vector prediction (MVP) for video coding may be provided. The method may be executed by at least one processor, and includes generating a list of MVP candidates for a current block based on a plurality of spatial motion vectors associated with neighboring blocks and non-neighboring blocks, inserting a first spatial motion vector predictor (SMVP) associated with a left neighboring block of the current block into the list of MVP candidates, inserting a second SMVP associated with an upper neighboring block of the current block into the list of MVP candidates, and pruning one or more candidates in the list of MVP candidates based on a determination that one or more candidates have the same SMVP.
[0006] According to an embodiment, a device for interleaved motion vector prediction (MVP) for video coding may be provided. The device may include at least one memory configured to store program code, and at least one processor configured to read the program code and operate as instructed by the program code. The program code includes candidate bank code for inserting a motion vector associated with a current block into a reference motion vector candidate bank, generation code configured to cause at least one processor to generate a list of MVP candidates for the current block based on spatial motion vectors associated with neighboring blocks and non-neighboring blocks, first insertion code configured to cause at least one processor to insert a first spatial motion vector predictor (SMVP) associated with a left neighboring block of the current block into the list of MVP candidates, second insertion code configured to cause at least one processor to insert a second SMVP associated with an upper neighboring block of the current block into the list of MVP candidates, and pruning code configured to cause at least one processor to prune one or more candidates in the list of MVP candidates based on a determination that one or more candidates have the same SMVP.
[0007]
[0007] According to an embodiment, a non-transitory computer-readable medium storing instructions may be provided. When the instructions are executed by at least one processor of a device that interleaves motion vector prediction (MVP) for video coding, the at least one processor is caused to generate a list of MVP candidates for a current block based on spatial motion vectors associated with neighboring blocks and non-neighboring blocks, insert a first spatial motion vector predictor (SMVP) associated with a left neighboring block of the current block into the list of MVP candidates, insert a second SMVP associated with an upper neighboring block of the current block into the list of MVP candidates, and prune one or more candidates in the list of MVP candidates based on a determination that one or more candidates have the same SMVP.
Brief Description of the Drawings
[0008]
Figure 1A
[0008] FIG. showing an example of a partition tree under AV1 and a VPN framework according to an embodiment of the present disclosure.
Figure 1B
[0009] FIG. showing examples of block partitioning and tree structures using quadtree and binary tree block partitioning according to an embodiment of the present disclosure.
Figure 1C
[0010] FIG. showing examples of vertical center side triple tree partitioning and horizontal center side triple tree partitioning according to an embodiment of the present disclosure.
Figure 1D
[0011] FIG. showing an example of search points in a merge mode using a motion vector difference according to an embodiment of the present disclosure.
Figure 1E
[0012] FIG. showing an example of a neighborhood of a spatial motion vector according to an embodiment of the present disclosure.
Figure 1F
[0013] A diagram showing an example of motion field estimation by linear projection according to an embodiment of the present disclosure.
Figure 1G
[0014] A diagram showing an example of a block position for deriving a temporal motion vector predictor according to an embodiment of the present disclosure.
Figure 1H
[0015] A diagram showing an example of generation of additional motion vector candidates for a block with a single reference according to an embodiment of the present disclosure.
Figure 1I
[0016] A diagram showing an example of generation of additional motion vector candidates for a block with a composite reference according to an embodiment of the present disclosure.
Figure 2
[0017] A diagram showing a reference motion vector candidate update process in the related art according to an embodiment of the present disclosure.
Figure 3
[0018] A flowchart for constructing a motion vector candidate list according to an embodiment of the present disclosure.
Figure 4
[0019] A simplified block diagram of a communication system according to an embodiment of the present disclosure.
Figure 5
[0020] A diagram showing the arrangement of a video encoder and a video decoder in a streaming environment.
Figure 6
[0021] A functional block diagram of a video decoder according to an embodiment of the present disclosure.
Figure 7
[0022] A functional block diagram of a video encoder according to an embodiment of the present disclosure.
Figure 8A
[0023] An exemplary diagram showing the insertion of an interleaved MVP candidate according to an embodiment of the present disclosure.
Figure 8B
[0024] A flowchart of an exemplary process for video coding and decoding according to an embodiment of the present disclosure.
Figure 9
[0025] A diagram of a computer system according to an embodiment of the present disclosure.
DETAILED DESCRIPTION OF THE INVENTION
[0009]
[0026] The proposed methods and processes may be used individually, used separately, or used in combination. Embodiments of the present disclosure relate to improvements in the generation, maintenance, and update of one or more motion vector prediction (MVP) lists.
[0010]
[0027] According to an embodiment of the present disclosure, an interleaved scanning order for one or more adjacent and non-adjacent spatial motion vector predictors (SMVPs) associated with a current block for more efficiently utilizing the similarity between the current MV and neighboring SMVPs.
[0011]
[0028] According to one aspect, a first adjacent SMVP from near the lower left corner may be inserted into the MVP list. Further, a second adjacent SMVP from near the upper right corner may be inserted into the MVP list. The above-described operation of interleaving the insertion of adjacent SMVPs from the left with the insertion of adjacent SMVPs from the top (also referred to as "scanning") is repeated until all adjacent SMVPs have been used, scanned, inserted, or pruned. Pruning may include not inserting a motion vector into the MVP list or deleting a motion vector from the MVP list if the incoming SMVP candidate has the same MVP as one already present in the MVP list. In some embodiments, partial pruning may be performed instead of complete pruning. As an example, for a particular motion vector candidate or for motion vector candidates that meet certain conditions, the pruning operation may not be performed.
[0012]
[0029] The SMVP associated with the top-left corner candidate may be inserted at the end of the adjacent SMVP, or counted as a left or top candidate for insertion at any position, or counted as a non-adjacent candidate, and thus may not be inserted during the processing of the adjacent SMVP.
[0013]
[0030] According to one embodiment, motion vectors (also interchangeably referred to as "motion vector predictors" or "candidates") within the MVP list may be weighted, and the weights may be accumulated. In some embodiments, the weights of one or more motion vector candidates may not be accumulated. As an example, the weighting associated with one or more motion vector candidates for which the weights may not be accumulated may be set to zero. In some embodiments, one or some of the candidate motion vectors may not be counted while counting the context for syntax element context modeling (e.g., reference picture index signaling, or new MV signaling). According to one embodiment, the candidate block size may be used during the weighting process or accumulation.
[0014]
[0031] In some embodiments, in addition to the features described above, candidate interleaved scanning and / or insertion may start by inserting the first adjacent SMVP from near the top-right corner into the MVP list, and subsequently, the second adjacent SMVP from near the bottom-left corner may be inserted into the MVP list. Such embodiments may allow for more efficient utilization of the MVP list to include relevant neighborhoods when the upper neighborhood is used more frequently. As an example, this scanning order from top to left may be beneficial when the video content is vertical or includes vertical motion.
[0015]
[0032] In some embodiments, in addition to the features described above, candidate interleaved scanning and / or insertion may start by inserting the first adjacent SMVP from near the lower left corner into the MVP list, and subsequently, the second adjacent SMVP from near the upper right corner may be inserted into the MVP list. The next set of interleaved SMVPs may include SMVPs associated with non-adjacent blocks. Such embodiments may enable more efficient utilization of the MVP list to include more relevant non-adjacent neighborhoods when the adjacent neighborhoods along the upper and left edges of the current block may have similar SMVPs compared to the current block.
[0016]
[0033] In some embodiments, in addition to the features described above, candidate interleaved scanning and / or insertion may start by inserting the first adjacent SMVP from near the upper right corner into the MVP list, and subsequently, the second adjacent SMVP from near the lower left corner may be inserted into the MVP list. The next set of interleaved SMVPs may include SMVPs associated with non-adjacent blocks. Such embodiments may enable more efficient utilization of the MVP list to include more relevant non-adjacent neighborhoods when the adjacent neighborhoods along the upper and left edges of the current block may have similar SMVPs compared to the current block.
[0017]
[0034] According to one embodiment, whether the SMVP associated with the adjacent neighborhood from above is added to the MVP list before the SMVP associated with the adjacent neighborhood from the left may be based on coded information or any information that may be available to both the encoder and the decoder when coding the current block, and this information includes, but is not limited to, block shape, block size, and block aspect ratio. As an example, when the block width is greater than the block height, the left neighborhood is added to the MVP list first. As another example, when the block height is greater than the block width, the left neighborhood is added to the MVP list first.
[0018]
[0035] Block partitioning in VP9 and AV1
[0036] Figure 1A is Figure 1100 of an example of a partitioning tree under VP9 and AV1. As shown in the upper half of Figure 1A, VP9 may use a 4-way partitioning tree starting from the 64×64 level down to the 4×4 level, with some additional restrictions for blocks of 8×8. Partitions labeled "R" refer to recursive partitions, i.e., partitions for which the same partitioning tree can be repeated at a lower scale until reaching the lowest 4×4 level.
[0019]
[0037] As shown in the lower half of Figure 1A, AV1 can not only extend the partitioning tree to a 10-way structure, but AV1 may also increase the maximum size (referred to as a superblock in VP9 / AV1 terminology) to start from 128×128. It will be understood that the 10-way structure may include 4:1 / 1:4 rectangular partitions that did not exist in VP9. The rectangular partitions cannot be further subdivided. Additionally, AV1 further enhances the flexibility for the use of partitions below the 8×8 level in the sense that 2×2 chroma inter prediction becomes possible in certain cases.
[0020]
[0038] Block partitioning in HEVC
[0039] In HEVC, in order to adapt to various local characteristics, by using a quadtree (QT) structure shown as a coding tree, a coding tree unit (CTU) is divided into coding units (CUs). The decision on whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU can be further divided into one, two, or four prediction units (PUs) according to the prediction unit (PU) partitioning type. Within one PU, the same prediction process is applied, and related information is sent to the decoder for each PU. After obtaining a residual block by applying a prediction process based on the PU partitioning type, the CU can be partitioned into transform units (TUs) according to another quadtree structure like the coding tree for the CU. A feature of the HEVC structure is that it has multiple partitioning concepts including CUs, PUs, and TUs. In HEVC, a CU or TU can only be square, but in the case of an inter-prediction block, a PU can be square or rectangular. In HEVC, one coding block may be further divided into four square sub-blocks, and a transform is performed on each sub-block, i.e., TU. Each TU can be further recursively divided (using quadtree partitioning) into smaller TUs called a residual quadtree (RQT). At the picture boundary, HEVC adopts an implicit quadtree partitioning so that the block continues quadtree partitioning until its size conforms to the picture boundary. One of the important features of the HEVC structure is that the HEVC structure has multiple partitioning concepts including CUs, PUs, and TUs.
[0021]
[0040] Block Partitioning in Versatile Video Coding (VVC)
[0041] Block Partitioning Structure Using Quadtree (QT) and Binary Tree (BT)
[0042] The QTBT structure may include concepts of multiple partition types, that is, the QTBT structure may eliminate the separation of the concepts of CU, PU, and TU and support greater flexibility in the CU partition shape. In the QTBT block structure, the CU may have a shape of either a square or a rectangle. As shown in FIG. 1B using the tree 1205, the coding tree unit (CTU) is first partitioned by a quadtree structure. The quadtree leaf nodes are further partitioned by a binary tree structure. There are two types of binary tree splits, namely, symmetric horizontal split and symmetric vertical split. The binary tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transformation processing without further partitioning. Therefore, the CU, PU, and TU have the same block size within the QTBT coding block structure. In JEM, the CU may be composed of coding blocks (CBs) of different color components. For example, in the case of P slices and B slices in the 4:2:0 color difference format, one CU includes one luma CB and two chroma CBs, and in some cases, a single component CB. For example, in the case of I slices, one CU includes only one luma CB or only two chroma CBs. Due to the QTBT partitioning method, the following parameters may be defined, that is, the CTU size: the size of the quadtree root node, which is the same concept as HEVC, MinQTSize: the minimum allowable quadtree leaf node size, MaxBTSize: the maximum size of the maximum allowable binary tree root node, MaxBTDepth: the maximum allowable binary tree depth, and MinBTSize: the minimum allowable binary tree leaf node size.
[0022]
[0043] In an example of the QTBT partitioning structure, the CTU size may be set as 128×128 luma samples having two corresponding 64×64 blocks of chroma samples, MinQTSize may be set as 16×16, MaxBTSize may be set as 64×64, MinBTSize (both width and height) may be set to 4×4, and MaxBTDepth is set to 4. The quadtree partitioning may first be applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes may have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). When the leaf quadtree node is 128×128, since the size exceeds MaxBTSize (i.e., 64×64), it is not further divided by the binary tree. Otherwise, the leaf quadtree node may be further partitioned by the binary tree. Thus, the quadtree leaf node is also the root node of the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further division is considered. When the binary tree node has a width equal to MinBTSize (i.e., 4), no further horizontal division is considered. Similarly, when the binary tree node has a height equal to MinBTSize, no further vertical division is considered. The leaf nodes of the binary tree are further processed by prediction and transformation processing without further partitioning. In JEM, the maximum CTU size is 256×256 luma samples.
[0023]
[0044] As shown in FIG. 1B, block 1205 shows an example of block partitioning by using QTBT, and the binary tree 1210 in FIG. 1B shows the corresponding tree representation. Solid lines indicate quadtree partitions, and dotted lines indicate binary tree partitions. At each partition (i.e., non-leaf) node of the binary tree, one flag may be signaled to indicate which partition type (i.e., horizontal or vertical) is used. 0 may indicate a horizontal partition, and 1 may indicate a vertical partition. In the case of quadtree partitioning, since the quadtree partition always divides the block both horizontally and vertically to create four sub-blocks of equal size, there is no need to indicate the partition type.
[0024]
[0045] In addition, the QTBT scheme supports the flexibility for luminance and chrominance to have individual QTBT structures. In the current related art, for P slices and B slices, the luminance CTB and chrominance CTB within one CTU share the same QTBT structure. However, in the case of I slices, the luminance CTB is partitioned into CUs by a certain QTBT structure, and the chrominance CTB is partitioned into chrominance CUs by another QTBT structure. This means that a CU within an I slice may be composed of a coding block of a luminance component or coding blocks of two chrominance components, and a CU within a P slice or a B slice may be composed of coding blocks of all three color components.
[0025]
[0046] In HEVC, in order to reduce memory access for motion compensation, inter prediction for small blocks may be restricted. As a result, bi-prediction may not be supported for 4×8 and 8×4 blocks, and inter prediction may not be supported for 4×4 blocks. In QTBT implemented in JEM-7.0, these restrictions have been removed.
[0026]
[0047] Block partitioning structure using a ternary tree (TT)
[0048] In VVC, as shown in FIG. 1C, a multi-type-tree (MTT) structure that adds horizontal and vertical center-side triple trees on top of the QTBT may be included in each of block 1305 and block 1310. The advantage of TT partitioning is to complement quadtree and binary tree partitioning. While quadtree and binary tree are always divided along the block center, TT partitioning can also capture objects located at the block center. Furthermore, since the width and height of the proposed TT partition are always powers of 2, no additional conversion is required. The design of the two-level tree is mainly motivated by complexity reduction. Theoretically, the complexity of traversing the tree is T D where T represents the number of split types and D represents the depth of the tree.
[0027]
[0049] Merge mode with motion vector difference (MMVD)
[0050] In VVC, in addition to the merge mode where implicitly derived motion information is directly used for generating predicted samples of the current CU, a merge mode with motion vector difference (MMVD) is introduced. Immediately after transmitting the skip flag and the merge flag, an MMVD flag may be signaled to specify whether the MMVD mode is used for the CU.
[0028]
[0051] In MMVD, after a merge candidate is selected, the merge candidate may be further refined by the signaled MVD information. The additional information may include a merge candidate flag, an index for specifying the magnitude of the motion, and an index for indicating the direction of the motion. In the MMVD mode, one of the first two candidates in the merge list is selected for use as the MV base. A merge candidate flag may be signaled to specify which one is used.
[0029]
[0052] The distance index specifies the magnitude information of the motion and indicates a predefined offset from the starting point. As shown in FIG. 1D, the offset may be added to either the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1.
[0030]
Table 1
[0031]
[0053] The direction index represents the direction of the MVD with respect to the starting point. The direction index can represent the four directions shown in Table 2.
[0032]
Table 2
[0033]
[0054] Note that the meaning of the MVD sign can vary depending on the information of the starting MV. When the starting MV is a single prediction MV or a bi-prediction MV and both lists point to the same side of the current picture, the signs in Table 2 specify the signs of the MV offsets added to the starting MV. As an example, when both POCs of the two references are greater than the POC of the current picture or both are less than the POC of the current picture. When the starting MV is a bi-prediction MV in a state where the two MVs point to different sides of the current picture and the difference in POC within list 0 is greater than the difference in POC within list 1, the signs in Table 2 specify the signs of the MV offsets added to the MV component of the starting MV in list 0, and the signs of the MVs in list 1 have opposite values. As an example, one reference POC is greater than the POC of the current picture and the other reference POC is less than the POC of the current picture, which is the sign of the MV offset added to the MV component of the starting MV in list 0, and the signs of the MVs in list 1 have opposite values. Otherwise, when the difference in POC within list 1 is greater than list 0, the signs in Table 2 specify the signs of the MV offsets added to the MV component of the starting MV in list 1, and the signs of the MVs in list 0 have opposite values.
[0034]
[0055] The MVD may be scaled according to the difference in POC in each direction. If the difference in POC in both lists is the same, scaling may not be performed. Otherwise, if the difference in POC in list 0 is greater than the difference in POC in list 1, the MVD of list 1 may be scaled. If the difference in POC of L1 is greater than that of L0, the MVD of list 0 may also be scaled in the same way. When the starting MV is singly predicted, the MVD may be added to the available MV.
[0035]
[0056] Symmetric MVD coding
[0057] In VVC, in addition to the MVD signaling in the normal uni - directional prediction and bi - directional prediction modes, a symmetric MVD mode for bi - directional MVD signaling may be applied. In the symmetric MVD mode, the motion information including both the reference picture indices of list 0 and list 1 and the MVD of list 1 may be derived rather than signaled.
[0036]
[0058] The decoding process of the symmetric MVD mode at the slice level may be as follows. At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived. If mvd_l1_zero_flag is 1, BiDirPredFlag may be set to be equal to 0. Otherwise, if the closest reference picture in list 0 and the closest reference picture in list 1 form a pair of the forward and reverse directions of the reference picture or a pair of the reverse and forward directions of the reference picture, BiDirPredFlag may be set to 1, and both the reference pictures of list 0 and list 1 may be short - term reference pictures. Otherwise, BiDirPredFlag is set to 0.
[0037]
[0059] The decoding process of the symmetric MVD mode at the CU level can be as follows. At the CU level, the CU is bi-predicted coded, and when BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether the symmetric mode can be used is explicitly signaled. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indices of list 0 and list 1 are set to be equal to the pair of reference pictures respectively. MVD1 is set to be equal to (-MVD0).
[0038]
[0060] Inter-mode coding in CWG-B018
[0061] In AV1, for each coded block between frames, when the mode of the current block is an inter-coding mode rather than a skip mode, another flag may be signaled to indicate whether a single reference mode or a composite reference mode is used for the current block.
[0039]
[0062] The single mode may include a predicted block generated by one motion vector in the single reference mode. In the case of single reference, the following modes may be signaled. (1) Using one of the motion vector predictors (MVP) in the list indicated by the NEARMV - DRL (dynamic reference list) index, (2) Using one of the motion vector predictors (MVP) in the list signaled by the NEWMV - DRL index as a reference and applying a delta to the MVP, and (3) Using a motion vector based on the global motion parameters at the frame level of GLOBALMV.
[0040]
[0063] The predicted block generated by weighted-averaging two predicted blocks may be derived from two motion vectors in the composite reference mode. In the case of composite reference, the following modes may be signaled: (1) Use one of the motion vector predictors (MVPs) in the list signaled by the NEAR_NEARMV - DRL index; (2) Use one of the motion vector predictors (MVPs) in the list signaled by the NEAR_NEWMV - DRL index as a reference and send a delta MV for the second MV; (3) Use one of the motion vector predictors (MVPs) in the list signaled by the NEW_NEARMV - DRL index as a reference and send a delta MV for the first MV; (4) Use one of the motion vector predictors (MVPs) in the list signaled by the NEW_NEWMV - DRL index as a reference and send delta MVs for both MVs; and (5) Use the MVs from each reference based on the global motion parameters at the frame level of GLOBAL_GLOBALMV.
[0041]
[0064] Motion Vector Difference Coding in AV1
[0065] In AV1, 1 / 8 pixel motion vector accuracy (or precision) is possible, and the following syntax may be used to signal the motion vector differences within reference frame list 0 or list 1. (1) mv_joint specifies which components of the motion vector difference are non-zero. 0 indicates that there is no non-zero MVD along either the horizontal or vertical direction, 1 indicates that there is a non-zero MVD only along the horizontal direction, 2 indicates that there is a non-zero MVD only along the vertical direction, and 3 indicates that there is a non-zero MVD along both the horizontal and vertical directions. (2) mv_sign specifies whether the motion vector difference is positive or negative. (3) mv_class specifies the class of the motion vector difference. As shown in Table 3, the higher the class, the larger the magnitude of the motion vector difference. (4) mv_bit specifies the integer part of the offset between the motion vector difference and the magnitude of the start of each MV class. (5) mv_fr specifies the first two fractional bits of the motion vector difference, and (6) mv_hp specifies the third fractional bit of the motion vector difference.
[0042]
Table 3
[0043]
[0066] Adaptive MVD Resolution in CWG-B092
[0067] In the case of NEW_NEARMV and NEAR_NEWMV modes, the accuracy of the MVD may depend on the associated class and the size of the MVD. First, a fractional MVD is only permitted when the size of the MVD is 1 pixel or less. Second, when the value of the associated MV class is MV_CLASS_1 or higher, only one MVD value is permitted. The MVD values in each MV class are derived as 4, 8, 16, 32, 64 for MV class 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5). In addition, when the current block is coded as NEW_NEARMV or NEAR_NEWMV mode, a certain context may be used to signal mv_joint or mv_class. Otherwise, another context may be used to signal mv_joint or mv_class.
[0044]
Table 4
[0045]
[0068] Joint MVD coding (JMVD) in CWG-B092
[0069] To indicate whether the MVDs of the two reference lists are jointly signaled, a new inter-coding mode named JOINT_NEWMV may be applied. When the inter-prediction mode is equal to the JOINT_NEWMV mode, the MVDs of reference list 0 and reference list 1 are jointly signaled. Therefore, only one MVD named joint_mvd is signaled and sent to the decoder, and from joint_mvd, the delta MVs of reference list 0 and reference list 1 are derived.
[0046]
[0070] The JOINT_NEWMV mode is signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. No additional context is added.
[0047]
[0071] When the JOINT_NEWMV mode is signaled and the POC distances between the two reference frames and the current frame are different, the MVD is scaled for reference list 0 or reference list 1 based on the POC distance. Specifically, the distance between reference frame list 0 and the current frame can be denoted as td0, and the distance between reference frame list 1 and the current frame is shown as td1. If td0 is greater than or equal to td1, joint_mvd is used as it is for reference list 0, and the mvd for reference list 1 is derived from joint_mvd based on Equation (1).
[0048]
Equation
[0049]
[0072] Otherwise, if td1 is greater than or equal to td0, joint_mvd is used as it is for reference list 1, and the mvd for reference list 0 is derived from joint_mvd based on Equation (2).
[0050]
Equation
[0051]
[0073] Improvement of Adaptive MVD Resolution in CWG-C011
[0074] In the case of a single reference, a new inter-coding mode named AMVDMV may be added. When the AMVDMV mode is selected, the AMVDMV mode indicates that AMVD is applied to signal the MVD. To indicate whether AMVD is applied to the joint MVD coding mode, one flag named amvd_flag is added under the JOINT_NEWMV mode. When the adaptive MVD resolution is applied to the joint MVD coding mode named joint AMVD coding, the MVDs of two reference frames may be signaled jointly, and the accuracy of the MVD is implicitly determined by the magnitude of the MVD. Otherwise, the MVDs of two (or three or more) reference frames are signaled jointly, and the conventional MVD coding is applied.
[0052]
[0075] Adaptive motion vector resolution (AMVR) in CWG-C012 and CWG-C020
[0076] AMVR was first proposed in CWG-C012, and a total of seven MV accuracies (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8) are supported. For each prediction block, the AVM encoder may explore all supported accuracy values and signal the best accuracy to the decoder.
[0053]
[0077] To shorten the encoder's execution time, two accuracy sets are supported. Each accuracy set contains four pre-defined accuracies. The accuracy set may be adaptively selected at the frame level based on the maximum accuracy value of the frame. Similar to AV1, the maximum accuracy may be signaled in the frame header. Table 5 summarizes the supported accuracy values based on the maximum accuracy at the frame level.
[0054]
Table 5
[0055]
[0078] Current AVM software (similar to AV1) has a frame-level flag indicating whether the MV of a frame includes sub-pel accuracy. AMVR is enabled only when the value of the cur_frame_force_integer_mv flag is 0. In AMVR, when the accuracy of a block is lower than the maximum accuracy, the motion model and interpolation filter may not need to be signaled. When the accuracy of a block is lower than the maximum accuracy, the motion mode may be assumed to be a translational motion, and the interpolation filter may be assumed to be a REGULAR interpolation filter. Similarly, when the accuracy of a block is 4 pels or 8 pels, the inter-intra mode is not signaled and is inferred to be 0.
[0056]
[0079] Motion vector predictor lists in AV1 and AVM
[0080] Spatial motion vector predictors (SMVP: spatial motion vector predictor, both adjacent and non-adjacent SMVPs), temporal motion vector predictors, additional MV candidates and additionally derived MVPs in AV1, and reference bank MVPs are added to the AVM design. To store the MVPs, a fixed-size stack known as the motion vector predictor list is generated on both the encoder side and the decoder side.
[0057]
[0081] Spatial motion vector predictor (SMVP)
[0082] The spatial motion vector (MV) predictor is derived from spatial neighboring blocks including adjacent spatial neighboring blocks that are the direct neighborhood above and to the left of the current block, and non-adjacent spatial neighboring blocks that are close to the current block but not directly adjacent to the current block. An example of a set of spatial neighboring blocks for a luminance block is shown in FIG. 1E, and each spatial neighboring block is an 8×8 block. The spatial neighboring blocks are examined to find one or more MVs associated with the same reference frame index as the current block. For the current block, the search order of the 8×8 luminance blocks in the spatial neighborhood is as shown by numbers 1 to 8 in FIG. 5. (1) The upper adjacent row is checked from left to right, (2) the left adjacent column is checked from top to bottom, (3) the upper right neighboring block is checked, (4) the upper left block neighboring block is checked, (5) the first non-adjacent upper row is checked from left to right, (6) the first non-adjacent left column is checked from top to bottom, (7) the second non-adjacent upper row is checked from left to right, and (8) the second non-adjacent left column is checked from top to bottom.
[0058]
[0083] Adjacent candidates (1 to 3 in FIG. 1E) are put into the MV predictor list before the TMVP, and non-adjacent candidates (also known as outer candidates, i.e., candidates 4 to 8 in FIG. 1E) are put into the MV predictor list after the TMVP. All SMVP candidates should have the same reference picture as the current block. That is, if the current block has a single reference picture, an MVP candidate with a single reference picture and this reference picture is the same as the reference picture of the current block, or an MVP candidate with a composite reference picture (two reference pictures) and one of the reference pictures is the same as the reference picture of the current block, this MVP candidate will be put into the MV predictor list. If the current block has two reference pictures, this MVP candidate is put into the predictor list only when the MVP candidate with two reference pictures and these two reference pictures are the same as the reference pictures of the current block.
[0059]
[0084] Temporal Motion Vector Predictor (TMVP)
[0085] In addition to the spatial neighboring blocks, an MV predictor known as the temporal MV predictor may also be derived using blocks at the same position within the reference frame. To generate the temporal MV predictor, first, the MVs of the reference frames are stored together with the reference indices associated with their respective reference frames. Then, for each 8×8 block of the current frame, the MVs of the reference frames through which the trajectory passes through the 8×8 block are identified and stored in the temporal MV buffer together with the reference frame indices. In the case of inter prediction using a single reference frame, for performing temporal motion vector prediction of the future frame, regardless of whether the reference frame is a forward reference frame or a backward reference frame, the MVs are stored in 8×8 units. In the case of composite inter prediction, for performing temporal motion vector prediction of the future frame, only the forward MVs are stored in 8×8 units.
[0060]
[0086] Referring to FIG. 1G, the MV of the reference frame 1 (R1) 1620, i.e., MVref 1650, is indicated from the frame 1 (R1) 1620. In doing so, MVref 1650 passes through an 8×8 block (the block within the current frame 1615 and the block within the reference frame 0 1610 of the current frame). MVref is stored in the temporal MV buffer in association with this 8×8 block. During the motion projection process for deriving the temporal MV predictor, the reference frames may be scanned in a predefined order, i.e., in the order of LAST_FRAME, BWDREF_FRAME, ALTREF_FRAME, ALTREF2_FRAME, and LAST2_FRAME. The MVs from the reference frames with higher indices (in the scanning order) do not replace the previously identified MVs assigned by the reference frames with lower indices (in the scanning order).
[0061]
[0087] When a pre - defined block coordinate is given, to derive a temporal MV predictor, e.g., MV0 in FIG. 1F, which indicates the reference frame from the current block, the relevant MVs stored in the temporal MV buffer are identified and projected onto the current block.
[0062]
[0088] Referring to FIG. 1G, the pre - defined block positions for deriving the temporal MV predictor of a 16×16 block are shown. For a valid temporal MV predictor, up to seven blocks are checked. The temporal MV predictor is checked after the adjacent spatial MV predictor and before the non - adjacent spatial MV predictor.
[0063]
[0089] In the derivation of the MV predictor, all spatial and temporal MV candidates may be pooled, and each predictor may be assigned a weight determined during the scan of spatial and temporal neighboring blocks. The candidates may be sorted and ranked based on the associated weights, and up to four candidates are identified and added to the MV predictor list. This list of MV predictors, also called the dynamic reference list (DRL), is further used in the dynamic MV prediction mode as described in the next sub - section.
[0064]
[0090] Additional search MVPs for additional MVP candidates
[0091] If the MVP list is still not full, additional searches are performed and additional MVP candidates are used to fill the MVP list. The additional MVP candidates include, for example, global MVs, zero MVs, combined composite MVs without scaling, etc.
[0065]
[0092] The process of sorting MVP candidates
[0093] The adjacent SMVP candidates, TMVP candidates, and non - adjacent SMVP candidates added to the MVP list are sorted. Based on the current designs in AV1 and AVM, the sorting process is based on the weight of each candidate. The weight of a candidate is predefined according to the overlapping area between the current block and the candidate block.
[0066]
[0094] Derived MVP candidates
[0095] The derived MVP candidates are adopted in the AVM reference software according to Proposal CWG-B049, which includes both the MVP derived for a single reference picture and the MVP derived for the combined mode.
[0067]
[0096] Single inter prediction
[0097] If the reference frames of neighboring blocks are different from the reference frame of the current block but in the same direction, a temporal scaling algorithm can be used to scale the MV to match that reference frame in order to form the MVP of the motion vector of the current block. As shown in Figure 1H, the MVP of the motion vector mv0 1850 of the current block with temporal scaling can be derived using mv1 1855 from a neighboring block (the shaded block).
[0068]
[0098] Combined inter prediction
[0099] To derive the MVP of the current block, combined MVs from different neighboring blocks are utilized, but the reference frames of the combined MVs need to be the same as that of the current block. As shown in Figure 1I, the combined MVs (mv2 1960, mv3 1965) have the same reference frame as the current block, but these reference frames are from different neighboring blocks.
[0069]
[0100] Reference motion vector candidate bank
[0101] Each buffer corresponds to a unique reference frame type that covers a single reference frame or a pair of reference frames for single inter mode and combined inter mode respectively. All buffers are of the same size. When a new MV is added to a full buffer, existing MVs may be removed to make space for the new MV.
[0070]
[0102] For collecting reference MV candidates, the coding block may refer to the MV candidate bank in addition to the reference MV candidates obtained in the conventional AV1 reference MV list generation. After coding the superblock, the MV bank may be updated with the MVs used by the coding blocks of the superblock.
[0071]
[0103] Each tile may have an independent MV reference bank that can be utilized by all superblocks within the tile. At the start of encoding each tile, the corresponding bank may be emptied. Thereafter, when coding each superblock within that tile, the MVs from the bank may be used as MV reference candidates. At the end of encoding the superblock, the bank may be updated.
[0072]
[0104] Bank update
[0105] As shown in graphic 200 of FIG. 2, the bank update process may be based on the superblock. That is, after the superblock is coded, the first (up to 64) candidate MVs used by each coding block within the superblock are added to the bank. During the update, the pruning process may also be involved during the update.
[0073]
[0106] Bank reference
[0107] After scanning the reference MV candidates of the conventional AV1 or the new AV2, if there are empty slots in the candidate list, the codec may refer to the MV candidate bank for additional MV candidates (within the buffer with matching reference frame types). While proceeding in the reverse direction from the end to the beginning of the buffer, if the MV in the bank buffer does not yet exist in the candidate list, that MV may be added to the candidate list.
[0074]
[0108] MVP list construction process in the state-of-the-art design
[0109] In related art, as shown in flowchart 300 of FIG. 3, the MVP list may be constructed using complete pruning. One example may include operation 305 of adding adjacent SMVPs, operation 310 of adding TMVP, operation 315 of adding non - adjacent SMVPs, operation 320 of adding a sorting process for existing candidates, operation 325 of adding derived candidates, operation 330 of adding additional MVPs, and finally, operation 355 of adding candidates from the reference MV candidate bank.
[0075]
[0110] FIG. 4 shows a simplified block diagram of a communication system 400 according to an embodiment of the present disclosure. The communication system 400 may include at least two terminals 410 - 420 interconnected via a network 450. In the case of unidirectional data transmission, the first terminal 410 may code video data at a local location for transmission to other terminals 420 via the network 450. The second terminal 420 may receive the coded video data of other terminals from the network 450, decode the coded data, and display the restored video data. Unidirectional data transmission may be common in media - serving applications and the like.
[0076]
[0111] FIG. 4 shows a second pair of terminals 430, 440 provided to support two - way transmission of coded video that may occur, for example, during a video conference. In the case of two - way data transmission, each terminal 430, 440 may code the captured video data at a local location for transmission to other terminals via the network 450. Each terminal 430, 440 may also receive the coded video data transmitted by other terminals, decode the coded data, and display the restored video data on a local display device.
[0077]
[0112] In FIG. 4, terminals 410 to 440 are shown as a server, a personal computer, and a smartphone, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure find applications in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 450 represents any number of networks that transmit coded video data among terminals 410 to 440, including, for example, wired and / or wireless communication networks. Communication network 450 may exchange data in a circuit-switched channel and / or a packet-switched channel. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of network 450 may not be important for the operation of the present disclosure, unless otherwise described below.
[0078]
[0113] FIG. 5 shows the arrangement of a video encoder and a video decoder in a streaming environment, such as a streaming system 500, as an example of the use of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0079]
[0114] The streaming system can include a capture subsystem 513, and the capture subsystem 513 can include, for example, a video source 501 that creates an uncompressed video sample stream 502, such as a digital camera. The sample stream 502 is illustrated as a thick line to emphasize that it has a large amount of data when compared to the encoded video bit stream, and can be processed by an encoder 503 coupled to the camera 501. The encoder 503 can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video bit stream 504 is illustrated as a thin line to emphasize that it has a small amount of data when compared to the sample stream, and can be stored on a streaming server 505 for future use. One or more streaming clients 506, 508 can access the streaming server 505 to obtain copies 507, 509 of the encoded video bit stream 504. The client 506 can include a video decoder 510, and the video decoder 510 decodes a copy of the incoming encoded video bit stream 507 and creates an outgoing video sample stream 511, and the video sample stream 511 can be rendered on a display 512 or other rendering device not shown. In some streaming systems, the video bit streams 504, 507, 509 can be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. A video coding standard, informally known as Versatile Video Coding (VVC), is under development. The disclosed subject matter may be used in the context of VVC.
[0080]
[0115] FIG. 6 can be a functional block diagram of a video decoder 510 according to an embodiment of the present invention.
[0081]
[0116] Receiver 610 may receive one or more codec video sequences decoded by decoder 510. In the same or another embodiment, one coded video sequence is decoded at a time, in which case the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequence may be received from channel 612, which may be a hardware / software link to a storage device storing the encoded video data. Receiver 610 may receive the encoded video data together with other data, such as coded audio data and / or auxiliary data streams, that may be transferred to respective consuming entities (not shown). Receiver 610 may separate the coded video sequence from the other data. To handle network jitter, buffer memory 615 may be coupled between receiver 610 and entropy decoder / parser 620, hereinafter “parser”. If receiver 610 is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous network, buffer 615 may not be needed or may be small. When used in a best-effort packet network such as the Internet, buffer 615 may be needed, may be relatively large, and advantageously may be of an adaptable size.
[0082]
[0117] Video decoder 510 may include an analyzer 620 for reconstructing symbol 621 from an entropy-coded video sequence. The categories of these symbols include information used to manage the operation of decoder 510 and, in some cases, information for controlling a rendering device such as display 512 which, as shown in FIG. 6, is not an essential part of the decoder but can be coupled to the decoder. The control information for the rendering device may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment not shown. Analyzer 620 may analyze / entropy-decode the received coded video sequence. The coding of the coded video sequence can conform to a video coding technology or standard and can follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding, etc., regardless of the presence or absence of context dependence. Analyzer 620 may extract a set of at least one subgroup parameter of at least one subgroup of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to a group. The subgroups can include, for example, a group of pictures GOP, a picture, a tile, a slice, a macroblock, a coding unit CU, a block, a transform unit TU, a prediction unit PU, etc. The entropy decoder / analyzer may also extract information such as transform coefficients, quantization parameter QP values, motion vectors, etc. from the coded video sequence.
[0083]
[0118] The parser 620 may perform entropy decoding / parsing operations on the video sequence received from the buffer 615 to create the symbol 621. The parser 620 may receive the encoded data and selectively decode a particular symbol 621. Further, the parser 620 may determine whether a particular symbol 621 should be provided to the motion compensation prediction unit 653, the scaler / inverse transform unit 651, the intra prediction unit 652, or the loop filter 656.
[0084]
[0119] For the reconstruction of the symbol 621, multiple different units may be involved depending on the type of the coded video picture or a part thereof, such as inter pictures and intra pictures, inter blocks and intra blocks, and other factors. How each unit is involved may be controlled by subgroup control information parsed by the parser 620 from the coded video sequence. For the sake of brevity, such a flow of subgroup control information between the parser 620 and the multiple following units is not illustrated.
[0085]
[0120] In addition to the function blocks already described, the decoder 510 may be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for explaining the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate.
[0086]
[0121] The first unit is the scaler / inverse transform unit 651. The scaler / inverse transform unit 651 receives, as the symbol 621 from the parser 620, the quantized transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. The scaler / inverse transform unit 651 can output a block including sample values that can be input to the aggregator 655.
[0087]
[0122] In some cases, the output samples of the scaler / inverse transform 651 may not use prediction information from intra-coded blocks, i.e., previously reconstructed pictures, but may relate to blocks that can use prediction information from previously reconstructed parts of the current picture. The intra-picture prediction unit 652 can provide such prediction information. In some cases, the intra-picture prediction unit 652 uses the surrounding already reconstructed information fetched from the current partially reconstructed picture 658 to generate a block of the same size and shape as the block being reconstructed. The aggregator 655, in some cases, adds, for each sample, the prediction information generated by the intra-prediction unit 652 to the output sample information provided by the scaler / inverse transform unit 651.
[0088]
[0123] In other cases, the output samples of the scaler / inverse transform unit 651 may relate to inter-coded, possibly motion-compensated blocks. In such cases, the motion-compensation prediction unit 653 can access the reference picture memory 657 to fetch the samples used for prediction. After motion-compensating the fetched samples according to the symbols 621 related to the block, the aggregator 655 can, in this case, add these samples, called residual samples or residual signals, to the output of the scaler / inverse transform unit to generate the output sample information. The address in the reference picture memory from which the motion-compensation unit fetches the prediction samples can be controlled by a motion vector, for example, in the form of symbols 621 that can have X, Y, and reference picture components and are available to the motion-compensation unit. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory when an exact sub-sample motion vector is used, a motion vector prediction mechanism, etc.
[0089]
[0124] The output samples of the aggregator 655 can be subject to various loop filtering techniques in the loop filter unit 656. Video compression techniques can include in-loop filter techniques, and the in-loop filter techniques are controlled by parameters that the loop filter unit 656 can utilize as symbols 621 from the parser 620 included in the coded video bitstream, but also correspond to meta information obtained during the decoding of a previous portion in the decoding order of the coded picture or coded video sequence, and can also correspond to previously reconstructed and loop filter processed sample values.
[0090]
[0125] The output of the loop filter unit 656 can be a sample stream that is output to the render device 512 and stored in the reference picture memory 658 for use in future inter-picture prediction.
[0091]
[0126] When a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. When the coded picture is fully reconstructed and the coded picture is identified as a reference picture, for example, by the parser 620, the current reference picture 658 can become part of the reference picture buffer 657, and a new current picture memory can be reallocated before starting the reconstruction of the next coded picture.
[0092]
[0127] Video decoder 510 may perform a decoding operation according to a predetermined video compression technique that can be documented in a standard such as ITU-T Rec.H.265. The encoded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the encoded video sequence conforms to the syntax of the video compression technique or standard specified in the document or standard of the video compression technique, particularly the profile document therein. Also, in order to comply, it may be necessary that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstructed sample rate measured, for example, in megasamples per second, the maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted through the specifications of a Hypothetical Reference Decoder (HRD) and the metadata for HRD buffer management signaled in the encoded video sequence.
[0093]
[0128] In one embodiment, receiver 610 may receive additional redundant data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by video decoder 510 to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0094]
[0129] FIG. 7 may be a functional block diagram of video encoder 503 according to an embodiment of the present invention.
[0095]
[0130] The encoder 503 may receive video samples from a video source 501 that can capture a video image coded by the encoder 503 and that is not part of the encoder.
[0096]
[0131] The video source 501 may provide a source video sequence coded by the encoder 503 in the form of a digital video sample stream, and the form of the digital video sample stream can be any suitable bit depth, for example, 8 bits, 10 bits, 12 bits,..., any color space, for example, BT.601 Y CrCB, RGB,..., and any suitable sampling structure, for example, Y CrCb 4:2:0, Y CrCb 4:4:4. In a media serving system, the video source 501 may be a storage device that stores previously prepared video. In a video conferencing system, the video source 503 may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that give motion when viewed in sequence. The picture itself may be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. A person skilled in the art can easily understand the relationship between a pixel and a sample. In the following description, the focus is on the sample.
[0097]
[0132] According to one embodiment, the encoder 503 may code and compress pictures of a source video sequence in real time or under any other time constraints in response to requests by an application to obtain a coded video sequence 743. Enforcing an appropriate coding speed is one of the functions of the controller 750. The controller controls other functional units and is functionally coupled to these units as described below. For simplicity, the couplings are not shown. Parameters set by the controller may include rate control related parameters such as picture skip, quantizer, lambda value of rate distortion optimization techniques, picture size, group of pictures (GOP) layout of pictures, maximum motion vector search range, and the like. Those skilled in the art can readily identify other functions of the controller 750 that may be relevant to a video encoder 503 optimized for a particular system design.
[0098]
[0133] Some video encoders operate in what those skilled in the art would readily recognize as a "coding loop." As a very simplified explanation, the coding loop is part of encoder 730, hereinafter "source coder," which is responsible for creating symbols based on the input picture and reference pictures to be coded, and local decoder 733 incorporated in encoder 503. Local decoder 733 reconstructs the symbols to create sample data. Since in the video compression techniques contemplated by the disclosed subject matter the compression between symbols and the coded video bit stream is reversible, the remote decoder also creates this sample data. The reconstructed sample stream is input into reference picture memory 734. Since decoding the symbol stream results in bit-exact results regardless of whether the decoder is local or remote, the contents of the reference picture buffer are also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as reference picture samples as the sample values that the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronization and the drift that results when synchronization cannot be maintained, for example due to channel errors, is well known to those skilled in the art.
[0099]
[0134] The operation of "local" decoder 733 can be considered the same as the operation of "remote" decoder 510, which has already been described in detail in connection with FIG. 6. However, referring briefly to FIG. 7, since the symbol is available and the encoding / decoding of the symbol into the coded video sequence by entropy encoder 745 and parser 620 can be reversible, the entropy decoding part of decoder 510, which includes channel 612, receiver 610, buffer 615, and parser 620, may not be fully implemented in local decoder 733.
[0100]
[0135] What can be said at this point is that any decoder technology other than the parsing / entropy decoding existing in the decoder must likewise necessarily exist in substantially the same functional form in the corresponding encoder. Since the description of the encoder technology is the reverse of the decoder technology described comprehensively, it may be omitted. More detailed description is required only in specific areas, and is provided below.
[0101]
[0136] As part of its operation, source coder 730 may perform motion compensation prediction coding, which predictively codes an input frame by referring to one or more previously coded frames from the video sequence designated as the "reference frame". In this way, coding engine 732 codes the difference between a pixel block of the input frame and a pixel block of the reference frame that can be selected as a prediction reference for the input frame.
[0102]
[0137] Local video decoder 733 may decode the coded video data of a frame that can be designated as a reference frame based on the symbols created by source coder 730. The operation of coding engine 732 is, advantageously, an irreversible process. When the coded video data can be decoded in a video decoder not shown in FIG. 7, the reconstructed video sequence may typically be a reproduction of the source video sequence with some errors. Local video decoder 733 may replicate the decoding process that can be performed by the video decoder for the reference frame and store the reconstructed reference frame in reference picture cache 734. In this way, encoder 503 may locally store a copy of the reconstructed reference frame having common content as the reconstructed reference frame obtained without transmission error by the remote video decoder.
[0103]
[0138] Predictor 735 may perform predictive search for the coding engine 732. That is, for a new frame to be coded, predictor 735 may seek sample data as candidate reference pixel blocks, or specific metadata such as reference picture motion vectors, block shapes, etc. that can function as appropriate predictive references for the new picture, and search the reference picture memory 734. Predictor 735 may operate on a per sample block x pixel block basis to find an appropriate predictive reference. In some cases, the input picture may have predictive references drawn from a plurality of reference pictures stored in the reference picture memory 734, as determined by the search results obtained by predictor 735.
[0104]
[0139] Controller 750 may manage the coding operations of video coder 730, including setting parameters and subgroup parameters used to encode video data.
[0105]
[0140] The outputs of all of the aforementioned functional units may be subject to entropy coding in entropy coder 745. The entropy coder converts the symbols generated by the various functional units into a coded video sequence by reversibly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0106]
[0141] The transmitter 740 may buffer the coded video sequence created by the entropy coder 745 to prepare the coded video sequence for transmission via the communication channel 760, which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter 740 may merge the coded video data from the video coder 730 with other data to be transmitted, such as coded audio data and / or an auxiliary data stream source (not shown).
[0107]
[0142] The controller 750 may manage the operation of the encoder 503. During coding, the controller 750 may assign a specific coded picture type to each coded picture, which may affect the coding technique applicable to each picture. For example, often a picture may be assigned as one of the following frame types.
[0108]
[0143] An intra picture, an I picture, may be a picture that can be coded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of these variations of I pictures, as well as their respective uses and characteristics.
[0109]
[0144] A predicted picture, a P picture, may be a picture that can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0110]
[0145] Bidirectional prediction pictures, B pictures, may be pictures that can be coded and decoded using intra prediction or inter prediction that use up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple prediction pictures can use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0111]
[0146] The source picture may generally be spatially subdivided into a plurality of sample blocks, for example blocks of 4×4, 8×8, 4×8, or 16×16 samples each, and may be coded in block units. The blocks may be coded predictively by referring to other already-coded blocks as determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively, or blocks of an I picture may be coded predictively by referring to already-coded blocks of spatial prediction or intra prediction of the same picture. Pixel blocks of a P picture may be coded non-predictively via spatial prediction or temporal prediction by referring to one previously-coded reference picture. Blocks of a B picture may be coded non-predictively via spatial prediction or temporal prediction by referring to one or two previously-coded reference pictures.
[0112]
[0147] Video coder 503 may perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec.H.265. In its operation, video coder 503 may perform various compression operations including predictive coding operations that utilize temporal and spatial redundancies in the input video sequence. Thus, the coded video data may conform to the syntax specified by the video coding technology or standard being used.
[0113]
[0148] In one embodiment, transmitter 740 may transmit additional data along with the encoded video. Video coder 730 may include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and redundant slices, supplementary enhancement information (SEI) messages, visual user usability information (VUI) parameter set fragments, and the like.
[0114]
[0149] FIG. 8A shows an exemplary interleaved MVP candidate insertion for block 800.
[0115]
[0150] The present disclosure relates to an interleaved scan order of spatial neighboring motion vector predictors for more efficiently utilizing the similarity between the current MV and spatial neighboring MVs.
[0116]
[0151] In one embodiment, a first adjacent SMVP from the left (or top) neighborhood may be inserted into the MVP list, then an SMVP from the top (or left) may be inserted, and subsequently, the insertion of interleaved adjacent SMVPs may continue until all adjacent SMVPs are used or pruned. As an example, referring to block 800 in FIG. 8A, the adjacent SMVPs from the neighboring blocks may be scanned in the order of {1, 2, 3, 4, 5, 6, 7}. It may be assumed that all candidates may be of a predefined size.
[0117]
[0152] It can be understood that the SMVP associated with the top left corner candidate may be inserted last among the adjacent SMVPs, or counted as a left or top candidate for insertion at any position, or counted as a non - adjacent candidate and thus may not be inserted during the processing of adjacent SMVPs.
[0118]
[0153] According to one embodiment, motion vectors within the MVP list (which may also be interchangeably referred to as "motion vector predictors" or "candidates") may be weighted, and the weights may be accumulated. In some embodiments, the weights of one or more motion vector candidates may not be accumulated. As an example, the weighting associated with one or more motion vector candidates (e.g., candidate 7) for which the weights may not be accumulated may be set to zero. In some embodiments, one or more candidate motion vectors (e.g., candidate 7) may not be counted while counting the context for context modeling of syntax elements (e.g., reference picture index signaling, or new MV signaling). According to one embodiment, the block size of the candidate may be used during the weighting process or accumulation.
[0119]
[0154] Pruning may include not inserting a motion vector into the MVP list or removing a motion vector from the MVP list if an incoming SMVP candidate has the same MVP as one already present in the MVP list. In some embodiments, partial pruning may be performed instead of full pruning. As an example, for a particular motion vector candidate (e.g., candidate 4), or for motion vector candidates that meet certain conditions, the pruning operation may not be performed.
[0120]
[0155] In some embodiments, in addition to the above features, candidate interleaved scanning and / or insertion may start by inserting the first adjacent SMVP from near the upper right corner into the MVP list, and subsequently, the second adjacent SMVP from near the lower left corner may be inserted into the MVP list. Such embodiments can enable more efficient utilization of the MVP list to include relevant neighborhoods when the upper neighborhood is used more frequently. As an example, this scanning order from top to left may be beneficial when the video content is vertical or includes vertical motion. As an example, referring to block 800 in FIG. 8A, the adjacent SMVPs from neighboring blocks may be scanned in the order of {2, 1, 4, 3, 6, 5, 7}.
[0121]
[0156] In one embodiment, referring to block 800 in FIG. 8A, the adjacent SMVPs from neighboring blocks may be scanned in the order of {1, 2, 5, 6, 7, 4, 3}. In this embodiment, more different MVPs (assuming that 3 is likely to include an MVP similar to 1 and 5 is likely to be different from 1) can be inserted into the MVP list. As another embodiment, referring to block 800 in FIG. 8A, the adjacent SMVPs from neighboring blocks may be scanned in the order of {2, 1, 6, 5, 7, 3, 4}. The advantage of this embodiment is that the MVPs are more efficient when the upper neighborhood is used more frequently (e.g., when the video content is vertical or the video includes vertical motion).
[0122]
[0157] According to one embodiment, whether the SMVP associated with the upper neighboring neighborhood is added to the MVP list before the SMVP associated with the left neighboring neighborhood may be based on coded information or any information that may be available to both the encoder and the decoder when coding the current block. This information may include, but is not limited to, block shape, block size, and block aspect ratio. As an example, when the block width is greater than the block height, the left neighborhood is added to the MVP list first. Referring to block 800 in FIG. 8A, the scanning order of adjacent SMVPs from neighboring blocks is {2, 1, 4, 3, 6, 5, 7}. As another example, when the block height is greater than the block width, the left neighborhood is added to the MVP list first. Referring to block 800 in FIG. 8A, the scanning order of adjacent SMVPs from neighboring blocks is {1, 2, 3, 4, 5, 6, 7}.
[0123]
[0158] In one embodiment, referring to block 800 in FIG. 8A, the motion vector of the upper neighboring block located between candidate 4 and candidate 2 and / or the motion vector of the left neighboring block located between candidate 3 and candidate 1 may be inserted into the MVP candidate list after the motion vectors of some or all of the seven positions are listed in the MVP candidate list. In some embodiments, whether to insert the motion vector of the upper neighboring block located between candidate 4 and candidate 2 and / or the motion vector of the left neighboring block located between candidate 3 and candidate 1 into the MVP list may depend on coded information or any information that is available to both the encoder and the decoder when coding the current block. This information may include, but is not limited to, block shape, block size, block aspect ratio, and whether the left or upper neighborhood is selected as an MVP candidate for the neighboring block.
[0124]
[0159] In one embodiment, when the block width is greater than or equal to the block height, the motion vector of the upper neighboring block located between candidate 4 and candidate 2 may be inserted into the MVP candidate list after the motion vectors of the neighboring blocks located at all seven positions.
[0125]
[0160] In another embodiment, when the block width is less than or equal to the block height, the motion vector of the left neighboring block located between candidate 3 and candidate 1 is inserted into the MVP candidate list after the motion vectors of the neighboring blocks located at all seven positions.
[0126]
[0161] FIG. 8B shows an exemplary process 850 for generating the MVP list of the current block 800 based on the interleaving of adjacent and non - adjacent candidates.
[0127]
[0162] In operation 855, an MVP candidate list for the current block may be generated based on the spatial motion vectors associated with neighboring and non - neighboring blocks. In some embodiments, each candidate motion vector predictor in the MVP candidate list may be associated with a cumulative weight, and the cumulative weight may indicate the importance associated with each candidate motion vector predictor. The weight accumulation associated with at least one candidate motion vector predictor may be prohibited based on a determination that the cumulative weight associated with at least one candidate motion vector predictor is zero. In some embodiments, neighboring and non - neighboring blocks have a predetermined size.
[0128]
[0163] In operation 860, a first spatial motion vector predictor (SMVP) associated with the left neighboring block of the current block may be inserted into the MVP candidate list.
[0129]
[0164] In operation 865, a second SMVP associated with the upper neighboring block of the current block may be inserted into the MVP candidate list.
[0130]
[0165] In some embodiments, based on the first condition being satisfied, the first SMVP may be inserted before the second SMVP. As an example, the first condition is one of the current block having a width greater than its height and the current block having a height greater than its width.
[0131]
[0166] In some embodiments, based on the second condition being satisfied, the second SMVP may be inserted before the first SMVP. As an example, the second condition is one of the current block having a width greater than its height and the current block having a height greater than its width.
[0132]
[0167] In some embodiments, a fifth SMVP associated with a neighboring block in the immediate vicinity of the upper left corner of the current block may be inserted into the MVP candidate list. The fifth SMVP may be inserted into the MVP candidate list before the third SMVP and the fourth SMVP.
[0133]
[0168] In operation 870, a third SMVP associated with a non-neighboring block to the left of the current block may be inserted into the MVP candidate list. In operation 875, a fourth SMVP associated with a non-neighboring block above the current block may be inserted into the MVP candidate list.
[0134]
[0169] In operation 880, based on a determination that one or more candidates have the same SMVP, one or more candidates in the MVP candidate list may be pruned. Pruning may include not inserting a motion vector into the MVP list or deleting a motion vector from the MVP list when an incoming SMVP candidate has the same MVP as one already present in the MVP list. In some embodiments, partial pruning may be performed instead of complete pruning. As an example, for a particular motion vector candidate or for motion vector candidates that meet certain conditions, the pruning operation may not be performed.
[0135]
[0170] Figure 8B shows exemplary blocks of process 850, but in some implementations, process 800 may include additional blocks, fewer blocks than those illustrated in FIG. 8, blocks different from those illustrated in FIG. 8, or blocks arranged in a different way than those illustrated in FIG. 8. Additionally or alternatively, two or more of the blocks of process 800 may be executed in parallel.
[0136]
[0171] Further, the proposed method may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium to perform one or more of the proposed methods.
[0137]
[0172] The techniques described above may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media. For example, FIG. 9 shows a computer system 900 suitable for implementing a particular embodiment of the disclosed subject matter.
[0138]
[0173] The computer software may include code that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or code that can be the subject of assembly, compilation, linking, or similar mechanisms to create instructions that can be executed through interpretation, microcode execution.
[0139]
[0174] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet-of-things devices, etc.
[0140]
[0175] The components shown in FIG. 9 for computer system 900 are exemplary in nature and do not imply any limitation regarding the use or functionality of the computer software implementing embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement related to any one or combination of the components shown in the exemplary embodiments of computer system 900.
[0141]
[0176] Computer system 900 may include specific human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, through tactile input (keystrokes, swipes, movement of a data glove, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). The human interface device may also be used to capture certain media that is not necessarily directly related to conscious input by humans, such as audio (speech, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.).
[0142]
[0177] The input human interface device may include one or more of keyboard 901, mouse 902, trackpad 903, touch screen 910, data glove 1204, joystick 905, microphone 906, scanner 907, and camera 908 (only one of each is shown).
[0143]
[0178] Computer system 900 may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by touch screen 910, data glove 1204, or joystick 905, although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speaker 909, headphones (not shown), etc.), visual output devices (screen 910, etc., including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light emitting diode (OLED) screens, regardless of whether each has touch screen input capabilities and regardless of whether each has tactile feedback capabilities, and some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher output through means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown) may also be included.
[0144]
[0179] Computer system 900 can also include human-accessible storage devices and associated media such as optical media including CD / DVD ROM / RW 920 having a CD / DVD or similar medium 921, thumb drive 922, removable hard drive or solid state drive 923, legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.
[0145]
[0180] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.
[0146]
[0181] The computer system (900) can also include an interface to one or more communication networks (955). The network (955) can be, for example, a wireless network, a wired network, or an optical network. The network (955) can further be a local network, a wide area network, a metropolitan network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of the network (955) include local area networks such as Ethernet and wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., and wired or wireless wide area digital networks for TV including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial networks including CANBus, etc. A particular network (955) typically requires an external network interface adapter (954) attached to a particular general-purpose data port or peripheral bus (949) (such as a USB port of the computer system (900)), and other networks are typically integrated into the core of the computer system (900) by attaching to the system bus as described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). The computer system (900) can communicate with other entities using any of these networks (955). Such communication can be only unidirectional reception (such as broadcast TV), only unidirectional transmission (such as CANbus to a particular CANbus device), or bidirectional with other computer systems using local or wide area digital networks. For each of these networks (955) and network interfaces (954) described above, a particular protocol and protocol stack can be used.
[0147]
[0182] The foregoing human interface device, the human-accessible memory device, and the network interface can be attached to the core 940 of the computer system 900.
[0148]
[0183] The core 940 can include one or more central processing units (CPUs) 941, a graphics processing unit (GPU) 942, a specialized programmable processing unit 943 in the form of a field programmable gate array (FPGA), a hardware accelerator 944 for specific tasks, and the like. These devices may be connected through a system bus 1248 together with a read-only memory (ROM) 945, a random access memory (RAM) 946, an internal hard drive that is not accessible to the user, an internal mass storage 947 such as a solid state drive (SSD), etc. In some computer systems, the system bus 1248 can be made accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core's system bus 1248 or through a peripheral bus 949. Architectures for peripheral buses include Peripheral Component Interconnect (PCI), USB, and the like.
[0149]
[0184] The CPU 941, GPU 942, FPGA 943, and accelerator 944 can execute specific instructions that can together constitute the aforementioned computer code. The computer code can be stored in the ROM 945 or the RAM 946. Migration data can also be stored in the RAM 946, while persistent data can be stored, for example, in the internal mass storage 947. Fast storage and retrieval for any memory device can be enabled through the use of cache memory that can be closely associated with one or more CPUs 941, GPUs 942, mass storage 947, ROM 945, RAM 946, etc.
[0150]
[0185] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or can be of the kind well known and available to those having skill in the computer software arts.
[0151]
[0186] By way of example and not limitation, a computer system 900 having an architecture, specifically a core 940, can provide functionality as a result of software embodied in one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be the user-accessible mass storage introduced above, and media associated with specific storage of the core 940 that is non-transitory in nature, such as the core internal mass storage 947 or ROM 945. The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core 940. The computer-readable media can include one or more memory devices or chips depending on specific needs. The software can cause the core 940, specifically a processor (including a CPU, GPU, FPGA, etc.) therein, to define data structures stored in the RAM 946 and modify such data structures according to processes defined by the software, thereby executing a specific process described herein or a specific part of a specific process. Additionally, or alternatively, the computer system can provide functionality as a result of logic being hardwired or otherwise embodied within a circuit (e.g., accelerator 944) that operates instead of or in conjunction with software to execute a specific process described herein or a specific part of a specific process. References to software can, as necessary, include logic, and vice versa. References to computer-readable media can, as necessary, include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0152]
[0187] Although several exemplary embodiments have been described, there are changes, substitutions, and various alternative equivalents that are within the scope of the present disclosure. Thus, it will be understood that, although not explicitly illustrated or described herein, many systems and methods embodying the principles of the present disclosure and thus within the spirit and scope of the present disclosure can be devised by those skilled in the art.
Claims
1. A method for interleaved motion vector prediction (MVP) for video coding, executed by at least one processor, generating an MVP candidate list for a current block based on a plurality of spatial motion vector predictors (SMVPs) associated with neighboring blocks and non-neighboring blocks; inserting a first SMVP associated with a first neighboring block to the left of the current block into the MVP candidate list; inserting a second SMVP associated with a second neighboring block above the current block into the MVP candidate list, wherein the SMVP associated with the first neighboring block to the left of the current block is interleaved with the SMVP associated with the second neighboring block above the current block; pruning the one or more candidates in the MVP candidate list based on a determination that one or more candidates have the same SMVP; A method comprising:
2. inserting a third SMVP associated with a first non-neighboring block to the left of the current block into the MVP candidate list; inserting a fourth SMVP associated with a second non-neighboring block above the current block into the MVP candidate list; The method according to claim 1, further comprising:
3. The method according to claim 1, wherein the first SMVP is inserted before the second SMVP based on a first condition being satisfied.
4. The first condition is the width of the current block is greater than the height of the current block, and the height of the current block is greater than the width of the current block One of the above, the method according to claim 3.
5. The method according to claim 1, wherein the second SMVP is inserted before the first SMVP based on a second condition being satisfied.
6. The second condition is the width of the current block is greater than the height of the current block, and the height of the current block is greater than the width of the current block One of the above, the method according to claim 5.
7. The method according to claim 2, further comprising inserting a fifth SMVP associated with a third neighboring block at the upper left corner of the current block into the MVP candidate list.
8. The method according to claim 7, wherein the fifth SMVP is inserted into the MVP candidate list before the third SMVP and the fourth SMVP.
9. The method according to claim 1, wherein each candidate motion vector predictor in the MVP candidate list is associated with an accumulated weight, and the accumulated weight indicates the importance associated with each candidate motion vector predictor.
10. The method according to claim 9, wherein weight accumulation associated with the at least one candidate motion vector predictor is prohibited based on a determination that the accumulated weight associated with the at least one candidate motion vector predictor is zero.
11. The method according to claim 1, wherein the neighboring block and the non-neighboring block have a predetermined size.
12. A device for interleaved motion vector prediction (MVP) for video coding, the device comprising: at least one memory configured to store program code; at least one processor configured to read the program code and operate as commanded by the program code; wherein the program code comprises: generation code configured to cause the at least one processor to generate an MVP candidate list for a current block based on a plurality of spatial motion vector predictors (SMVPs) associated with neighboring blocks and non-neighboring blocks; first insertion code configured to cause the at least one processor to insert a first SMVP associated with a first neighboring block to the left of the current block into the MVP candidate list; second insertion code configured to cause the at least one processor to insert a second SMVP associated with a second neighboring block above the current block into the MVP candidate list, wherein the SMVP associated with the first neighboring block to the left of the current block is interleaved with the SMVP associated with the second neighboring block above the current block. pruning code configured to cause the at least one processor to prune the one or more candidates in the MVP candidate list based on a determination that the one or more candidates have the same SMVP A device comprising the same. **Claim 13** wherein the program code third insertion code configured to cause the at least one processor to insert a third SMVP associated with a first non-adjacent neighboring block to the left of the current block into the MVP candidate list; and fourth insertion code configured to cause the at least one processor to insert a fourth SMVP associated with a second non-adjacent neighboring block above the current block into the MVP candidate list The device according to claim 12, further comprising the same. **Claim 14** The device according to claim 12, wherein the first SMVP is inserted before the second SMVP based on the first condition being satisfied. **Claim 15** The device according to claim 12, wherein the second SMVP is inserted before the first SMVP based on the second condition being satisfied. **Claim 16** The device according to claim 13, wherein the program code further comprises fifth insertion code configured to cause the at least one processor to insert a fifth SMVP associated with a third adjacent neighboring block at the upper left corner of the current block into the MVP candidate list. **Claim 17** The device according to claim 16, wherein the fifth SMVP is inserted into the MVP candidate list before the third SMVP and the fourth SMVP. **Claim 18** The device according to claim 12, wherein each candidate motion vector predictor in the MVP candidate list is associated with an accumulated weight, and the accumulated weight indicates the importance associated with each candidate motion vector predictor. **Claim 19** The device according to claim 18, wherein weight accumulation associated with the at least one candidate motion vector predictor is prohibited based on a determination that the accumulated weight associated with the at least one candidate motion vector predictor is zero. **Claim 20** A non-transitory computer-readable medium storing instructions, which, when executed by one or more processors of a device that interleaves motion vector prediction (MVP) for video coding, cause the one or more processors to generate a list of MVP candidates for a current block based on a plurality of spatial motion vector predictors (SMVPs) associated with neighboring and non-neighboring blocks; insert a first SMVP associated with a first neighboring block to the left of the current block into the MVP candidate list; insert a second SMVP associated with a second neighboring block above the current block into the MVP candidate list, wherein the SMVP associated with the first neighboring block to the left of the current block is interleaved with the SMVP associated with the second neighboring block above the current block; prune the one or more candidates in the MVP candidate list based on a determination that one or more candidates have the same SMVP; A non-transitory computer-readable medium comprising one or more instructions that cause the above to be performed.
Citation Information
Patent Citations
Vector predictor list generation
US20200084468A1
Encoding device, decoding device, encoding method, and decoding method
WO2019221103A1