Improvement of the reference motion vector candidate bank

By updating the reference motion vector candidate bank with motion vectors from the same superblock and improving the consideration of other directions, the proposed method addresses the inefficiencies in existing video coding techniques, leading to enhanced prediction accuracy and compression efficiency.

JP2025518431APending Publication Date: 2025-06-17TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024516620
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-09
Filing Date
2022-09-13
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing video coding techniques, such as those used in AV1 and HEVC, do not efficiently update the reference motion vector candidate bank, leading to suboptimal results due to the exclusion of motion vector candidates from the same superblock and limited consideration of other possible directions and blocks.

Method used

The proposed method involves updating the reference motion vector candidate bank by inserting motion vectors associated with current blocks into the bank, while also considering motion vector predictors from the same superblock, thereby improving the generation and analysis of reference motion vector candidate banks.

Benefits of technology

This approach enhances the efficiency of video coding by allowing for more accurate prediction of motion vectors and improving the overall compression capabilities, reducing latency and increasing the relevance of distant motion vector candidates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025518431000001_ABST
    Figure 2025518431000001_ABST
Patent Text Reader

Abstract

A method, a device, and a non-transitory storage medium for decoding video data are provided. One or more vector predictors associated with a current block may be obtained from a reference motion vector candidate bank, and the one or more obtained motion vector predictors include at least one or more motion vectors associated with one or more already decoded blocks, and the one or more already decoded blocks belong to the same super block as the current block. Based on the one or more obtained motion vector predictors, a motion vector associated with the current block is determined, and based on the determined motion vector, the current block is decoded. The reference motion vector candidate bank may be updated by inserting a motion vector associated with the current block into the reference motion vector candidate bank.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications

[0001] This application claims priority based on U.S. Provisional Patent Application No. 63 / 342,744, filed with the United States Patent and Trademark Office on May 17, 2022, the disclosure of which is hereby incorporated by reference in its entirety.

[0002]

[0002] Embodiments of the present disclosure relate to image and video coding techniques. More particularly, embodiments of the present disclosure relate to improving the generation and analysis of reference motion vector candidate banks.

Background Art

[0003]

[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium established in 2015 that includes semiconductor companies, video - on - demand providers, video content production companies, software development companies, and web browser vendors. Many of the components of the AV1 project were derived from previous research activities by the Alliance's members. Individual contributors had started experimental technology platforms years earlier. Xiph / Mozilla's Daala had already made its code public in 2010, Google's experimental VP9 evolution project, VP10, was announced on September 12, 2014, and Cisco's Thor was made public on August 11, 2015. AV1, built on the VP9 codebase, incorporates additional techniques, some of which were developed in these experimental formats. The first version, 0.1.0, of the AV1 reference codec was made public on April 7, 2016. The Alliance announced the release of the AV1 bitstream specification on March 28, 2018, along with a reference, software - based encoder and decoder. On June 25, 2018, a verified version 1.0.0 of the specification was released. On January 8, 2019, a verified version 1.0.0 including errata 1 of the specification was released. The AV1 bitstream specification includes a reference video codec.

[0004]

[0004] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Since then, these groups have studied the potential need for standardizing future video coding technologies that could significantly outperform HEVC in terms of compression capabilities. In October 2017, these groups issued a Call for Proposal (CfP) for video compression with capabilities beyond HEVC. By February 15, 2018, a total of 22 CfP responses regarding standard dynamic range (SDR), 12 CfP responses regarding high dynamic range (HDR), and 12 CfP responses regarding 360 video categories were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th JVET (Joint Video Exploration Team - Joint Video Experts Team) meeting. After careful evaluation, JVET officially started the standardization of next-generation video coding beyond HEVC, namely Versatile Video Coding (VVC).

[0005]

[0005] In the AVM reference software by proposal CWG-B023, a reference motion vector MV candidate bank was adopted. The reference MV candidate bank may be used as a buffer for collecting reference MV candidates. According to the above-mentioned standard, the reference MV candidate bank is updated using the MVs used by coding blocks. However, the reference MV candidate bank update process does not consider MV candidates from the same superblock, and the reference process cannot consider other possible directions and blocks, resulting in suboptimal results.

Summary of the Invention

Problems to be Solved by the Invention

[0006]

[0006] Accordingly, there is a need for a method, system, device, and / or non-transitory storage medium for improving video coding by improving the generation and update of a reference MV candidate bank.

Means for Solving the Problem

[0007]

[0007] According to an embodiment, a method for decoding video data may be provided. The method may be executed by at least one processor, and includes obtaining, from a reference motion vector candidate bank, a plurality of motion vector predictors associated with a current block, wherein the plurality of obtained motion vector predictors are associated with a plurality of already decoded blocks, and the plurality of already decoded blocks belong to the same super block as the current block; determining a motion vector associated with the current block based on the plurality of obtained motion vector predictors; updating the reference motion vector candidate bank by inserting the motion vector associated with the current block into the reference motion vector candidate bank; and decoding the current block based on the determined motion vector associated with the current block.

[0008] According to an embodiment, a device for decoding video data may be provided. The device may include at least one memory configured to store program code, and at least one processor configured to read the program code and operate as commanded by the program code. The program code is acquisition code configured to cause at least one processor to acquire a plurality of motion vector predictors associated with a current block from a reference motion vector candidate bank, wherein the plurality of acquired motion vector predictors are associated with a plurality of already decoded blocks, and the plurality of already decoded blocks belong to the same super block as the current block; determination code configured to cause at least one processor to determine a motion vector associated with the current block based on the plurality of acquired motion vector predictors; first update code configured to cause at least one processor to update the reference motion vector candidate bank by inserting the motion vector associated with the current block into the reference motion vector candidate bank; and first decoding code configured to cause at least one processor to decode the current block based on the determined motion vector associated with the current block.

[0009] According to an embodiment, a non-transitory computer-readable medium storing instructions may be provided. When the instructions are executed by at least one processor for decoding video data, the at least one processor is caused to: obtain, from a reference motion vector candidate bank, a plurality of motion vector predictors associated with a current block, wherein the plurality of obtained motion vector predictors are associated with a plurality of already decoded blocks, and the plurality of already decoded blocks belong to the same super block as the current block; determine a motion vector associated with the current block based on the plurality of obtained motion vector predictors; update the reference motion vector candidate bank by inserting the motion vector associated with the current block into the reference motion vector candidate bank; and decode the current block based on the determined motion vector associated with the current block.

Brief Description of the Drawings

[0010]

Figure 1A

[0010] FIG. showing an example of a partition tree under AV1 and VPN frameworks according to an embodiment of the present disclosure.

Figure 1B

[0011] FIG. showing an example of a block partition and a tree structure using quadtree and binary tree block partitioning according to an embodiment of the present disclosure.

Figure 1C

[0012] FIG. showing an example of vertical center side triple tree partitioning and horizontal center side triple tree partitioning according to an embodiment of the present disclosure.

Figure 1D

[0013] FIG. showing an example of search points in a merge mode using a motion vector difference according to an embodiment of the present disclosure.

Figure 1E

[0014] FIG. showing an example of a neighborhood of a spatial motion vector according to an embodiment of the present disclosure.

Figure 1F

[0015] FIG. is a diagram showing an example of motion field estimation by linear projection according to an embodiment of the present disclosure.

Figure 1G

[0016] FIG. is a diagram showing an example of a block position for deriving a temporal motion vector predictor according to an embodiment of the present disclosure.

Figure 1H

[0017] FIG. is a diagram showing an example of generation of additional motion vector candidates for a block having a single reference according to an embodiment of the present disclosure.

Figure 1I

[0018] FIG. is a diagram showing an example of generation of additional motion vector candidates for a block having a composite reference according to an embodiment of the present disclosure.

Figure 2

[0019] FIG. is a diagram showing a reference motion vector candidate update process in related art according to an embodiment of the present disclosure.

Figure 3

[0020] FIG. is a flowchart for constructing a motion vector candidate list according to an embodiment of the present disclosure.

Figure 4

[0021] FIG. is a simplified block diagram of a communication system according to an embodiment of the present disclosure.

Figure 5

[0022] FIG. shows the arrangement of a video encoder and a video decoder in a streaming environment.

Figure 6

[0023] FIG. is a functional block diagram of a video decoder according to an embodiment of the present disclosure.

Figure 7

[0024] FIG. is a functional block diagram of a video encoder according to an embodiment of the present disclosure.

Figure 8

[0025] FIG. is a flowchart of an exemplary process for video coding and decoding according to an embodiment of the present disclosure.

Figure 9

[0026] FIG. is a diagram of a computer system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011]

[0027] The functions of the proposed methods, devices, and processes may be used individually or in combination. Embodiments of the present disclosure relate to improvements in the generation, maintenance, and update of a reference motion vector (MV) candidate bank.

[0012]

[0028] According to embodiments of the present disclosure, instead of, or in addition to, a superblock candidate MVP that can be updated at the superblock level, a candidate motion vector predictor (MVP) may be updated at the coding block level. The advantage of a coding block level candidate MVP is that more candidate MVPs from the same superblock can be used to more efficiently predict the MV of the current block compared to using candidate MVPs that are far from the superblock to the left, above, or even further ahead. Additionally, in addition to the candidate MVP being updated at the coding block level, the candidate MVP may be updated at a specific level.

[0013]

[0029] According to one embodiment, the MV of a coding block may be updated with respect to the reference MV candidate bank immediately after the MV of that coding block is analyzed. The MV of the coding block may then be used as an MVP candidate for any subsequent coding block. In the related art, the MV of a block is only updated after the entire superblock has been analyzed or inserted into the reference MV candidate bank, which is an inefficient use of the reference MV candidate bank.

[0014]

[0030] According to one embodiment of the present disclosure, the reference MV candidate bank may be updated after the blocks in the pre-defined area are analyzed or at a specific level. As an example, the reference MV candidate bank may be updated after all the blocks in a 64×64 non-overlapping area are analyzed or at a level one level after the super-block level of the quadtree (QT). During coding or decoding, after all the coding blocks in this area are analyzed, all the MVs in the pre-defined area or QT may be updated for the reference MV candidate bank and made available for reference.

[0015]

[0031] The advantage of the above embodiment is that there is no need to frequently update the reference MV bank. In a specific hardware design where the reference MV bank is not stored in the cache, this design will reduce the latency. In addition, the above embodiment avoids updating MVs that are too close to the current coding block. These MVs may overlap with the spatial motion vector predictor (SMVP) of the current block.

[0016]

[0032] The above embodiments may be combined. As an example, two or more reference MV banks are maintained. The first reference MV candidate bank may be updated at the coding block level, and the second reference MV candidate bank may be updated at a predefined level. During reference, which of the first bank and the second bank to use as a reference may be implicitly selected or explicitly signaled. As a non-limiting example, two reference MV banks are maintained. One bank may be updated at the coding block level. That is, immediately after the block MV is analyzed, the MV of the analyzed block is inserted into that one bank (e.g., the first bank). The other bank may be updated at the superblock level. That is, after all coding blocks of the superblock of the current block are analyzed, the MVs of these coding blocks may be inserted into the other bank. During reference, in one example, if the block is larger than a specific block size (e.g., 16×16), the bank updated at the superblock level may be used. Otherwise, if the block size is less than or equal to a specific block size (e.g., 16×16), the bank updated at the block level may be used. It should be understood that combinations of multiple banks, or banks between several different specific levels (e.g., 32×32 and 64×64) are also possible.

[0017]

[0033] Embodiments of the present invention also relate to reversing the reference order of the reference MV candidate bank. In the related art, the reference of the reference MV candidate bank always started from the end of the bank towards the beginning of the bank. According to one embodiment of the present disclosure, the reference order of the reference MV candidate bank can be, by design, in the order from the beginning of the bank towards the end of the bank. The advantage of the reversed order is that the relevance of the distant MV candidates can be higher, which may be beneficial for certain content types, such as screen content or content without translational movement. It should be understood that a combination of the reversed reference and the reference starting from the end of the bank towards the beginning is also possible. As an example, a flag or a condition (e.g., block size, QP value, hash hit rate of screen content) may be used to determine which reference order can be used.

[0018]

[0034] According to one embodiment, the level of the reference MV candidate bank, flag, condition, or index may be signaled in a high-level syntax, and the high-level syntax includes, but is not limited to, a sequence parameter set, a video parameter set, a picture parameter set, a slice header, a frame header, an APS, and a tile header.

[0019]

[0035] In one embodiment, the MV of a coding block may be conditionally updated with respect to a reference MV candidate bank immediately after the MV of the coding block is analyzed, whereby the MV can be used as an MVP candidate for any subsequent coding block. Other conditions may also be used individually or in combination. Other conditions include, but are not limited to, whether the analyzed MV is associated with a block coded in a specific coding mode, whether the analyzed MV is associated with a block larger than, smaller than, or equal to a specific pre-defined block size, etc. Compared with related techniques where the MV of a block must be updated with respect to the reference MV candidate bank after the entire superblock is analyzed, the embodiments described herein can use the reference MV candidate bank more efficiently because not all analyzed MVs can be used for updating the reference MV candidate bank, and some MVs may be analyzed but not used for updating the reference MV candidate bank, thereby reducing latency and improving the efficiency of the reference MV candidate bank.

[0020]

[0036] Block partitioning in VP9 and AV1

[0037] FIG. 1A is a diagram 1100 of an example of a partitioning tree under VP9 and AV1. As shown in the upper half of FIG. 1A, VP9 may use a 4-way partitioning tree starting from the 64×64 level to the 4×4 level, and for blocks of 8×8, there are some additional restrictions. A partition labeled "R" refers to a recursive partition, i.e., a partition where the same partitioning tree can be repeated at a lower scale until the lowest 4×4 level is reached.

[0021]

[0038] As shown in the lower half of FIG. 1A, AV1 can not only extend the partition tree to a 10-way structure, but AV1 may also increase the maximum size (referred to as a superblock in VP9 / AV1 terminology) to start from 128×128. It will be understood that the 10-way structure may include 4:1 / 1:4 rectangular partitions that did not exist in VP9. The rectangular partition cannot be further subdivided. In addition, AV1 further enhances the flexibility for the use of partitions below the 8×8 level in the sense that 2×2 chrominance inter prediction becomes possible in certain cases.

[0022]

[0039] Block Partitioning in HEVC

[0040] In HEVC, in order to adapt to various local characteristics, the coding tree unit (CTU) is divided into coding units (CUs) by using a quadtree (QT) structure shown as a coding tree. The decision on whether to code the picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU can be further divided into one, two, or four prediction units (PUs) according to the prediction unit (PU) division type. Within one PU, the same prediction process is applied, and related information is sent to the decoder for each PU. After obtaining the residual block by applying the prediction process based on the PU division type, the CU can be partitioned into transform units (TUs) according to another quadtree structure like the coding tree for the CU. A feature of the HEVC structure is that it has a plurality of partitioning concepts including CUs, PUs, and TUs. In HEVC, a CU or TU can only be square, but in the case of an inter-prediction block, the PU can be square or rectangular. In HEVC, one coding block may be further divided into four square sub-blocks, and a transform is performed on each sub-block, that is, on the TU. Each TU can be further recursively divided (using quadtree partitioning) into smaller TUs called the Residual Quad Tree (RQT). At the picture boundary, HEVC adopts an implicit quadtree partitioning so that the block continues to be quadtree partitioned until its size conforms to the picture boundary. One of the important features of the HEVC structure is that the HEVC structure has a plurality of partitioning concepts including CUs, PUs, and TUs.

[0023]

[0041] Block Partitioning in Versatile Video Coding (VVC)

[0042] Block Partitioning Structure Using Quadtree (QT) and Binary Tree (BT)

[0043] The QTBT structure may include concepts of multiple partition types, that is, the QTBT structure may eliminate the separation of the concepts of CU, PU, and TU and support greater flexibility in the CU partition shape. In the QTBT block structure, the CU may have a shape of either a square or a rectangle. As shown in FIG. 1B using the tree 1205, the coding tree unit (CTU) is first partitioned by a quadtree structure. The quadtree leaf nodes are further partitioned by a binary tree structure. There are two types of binary tree splits, namely, symmetric horizontal split and symmetric vertical split. The binary tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transformation processing without further partitioning. Therefore, the CUs, PUs, and TUs have the same block size within the QTBT coding block structure. In JEM, the CU may be composed of coding blocks (CBs) of different color components. For example, in the case of P slices and B slices in the 4:2:0 color difference format, one CU includes one luma CB and two chroma CBs, and in some cases, it includes a single component CB. For example, in the case of I slices, one CU includes only one luma CB or only two chroma CBs. Due to the QTBT partitioning method, the following parameters may be defined, namely, the CTU size: the size of the quadtree root node, which is the same concept as HEVC, MinQTSize: the minimum allowable quadtree leaf node size, MaxBTSize: the maximum size of the maximum allowable binary tree root node, MaxBTDepth: the maximum allowable binary tree depth, and MinBTSize: the minimum allowable binary tree leaf node size.

[0024]

[0044] In an example of the QTBT partitioning structure, the CTU size may be set as 128×128 luminance samples having two corresponding 64×64 blocks of chroma samples, MinQTSize may be set as 16×16, MaxBTSize may be set as 64×64, MinBTSize (both width and height) may be set to 4×4, and MaxBTDepth is set to 4. The quadtree partitioning may first be applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes may have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). When the leaf quadtree node is 128×128, since the size exceeds MaxBTSize (i.e., 64×64), it is not further divided by the binary tree. Otherwise, the leaf quadtree node may be further partitioned by the binary tree. Thus, the quadtree leaf node is also the root node of the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further division is considered. When the binary tree node has a width equal to MinBTSize (i.e., 4), no further horizontal division is considered. Similarly, when the binary tree node has a height equal to MinBTSize, no further vertical division is considered. The leaf nodes of the binary tree are further processed by prediction and transformation processing without further partitioning. In JEM, the maximum CTU size is 256×256 luminance samples.

[0025]

[0045] As shown in FIG. 1B, block 1205 shows an example of block partitioning by using QTBT, and the binary tree 1210 in FIG. 1B shows the corresponding tree representation. Solid lines indicate quadtree partitions, and dotted lines indicate binary tree partitions. At each partition (i.e., non-leaf) node of the binary tree, one flag may be signaled to indicate which partition type (i.e., horizontal or vertical) is used. 0 may indicate a horizontal partition, and 1 may indicate a vertical partition. In the case of a quadtree partition, since the quadtree partition always divides the block both horizontally and vertically to create four sub-blocks of equal size, there is no need to indicate the partition type.

[0026]

[0046] In addition, the QTBT scheme supports the flexibility for luminance and chrominance to have individual QTBT structures. In the current related art, for P slices and B slices, the luminance CTB and chrominance CTB within one CTU share the same QTBT structure. However, for I slices, the luminance CTB is partitioned into CUs by a certain QTBT structure and the chrominance CTB is partitioned into chrominance CUs by another QTBT structure. This means that a CU within an I slice may be composed of a coding block of the luminance component or coding blocks of two chrominance components, and a CU within a P slice or a B slice may be composed of coding blocks of all three color components.

[0027]

[0047] In HEVC, in order to reduce memory access for motion compensation, inter prediction for small blocks may be restricted. As a result, bi-prediction may not be supported for 4×8 and 8×4 blocks, and inter prediction may not be supported for 4×4 blocks. In QTBT implemented in JEM-7.0, these restrictions are removed.

[0028]

[0048] Block partitioning structure using a ternary tree (TT)

[0049] In VVC, as shown in FIG. 1C, a Multi-Type-Tree (MTT) structure that adds horizontal and vertical center-side triple trees on top of the QTBT may be included in each of block 1305 and block 1310. The advantage of TT partitioning is to complement quadtree and binary tree partitioning. While quadtree and binary tree are always split along the block center, TT partitioning can also capture objects located at the block center. Furthermore, since the width and height of the proposed TT partition are always powers of 2, no additional conversion is required. The design of the two-level tree is mainly motivated by complexity reduction. Theoretically, the complexity of traversing the tree is T D where T represents the number of split types and D represents the depth of the tree.

[0029]

[0050] Merge mode with motion vector difference (MMVD)

[0051] In VVC, in addition to the merge mode where implicitly derived motion information is directly used for generating prediction samples of the current CU, a merge mode with motion vector difference (MMVD) is introduced. Immediately after transmitting the skip flag and the merge flag, an MMVD flag may be signaled to specify whether the MMVD mode is used for the CU.

[0030]

[0052] In MMVD, after a merge candidate is selected, the merge candidate may be further narrowed down by the signaled MVD information. The additional information may include a merge candidate flag, an index for specifying the magnitude of the motion, and an index for indicating the direction of the motion. In the MMVD mode, one of the first two candidates in the merge list is selected for use as the MV base. A merge candidate flag may be signaled to specify which one is used.

[0031]

[0053] The distance index specifies the magnitude information of the motion and indicates a predefined offset from the starting point. As shown in FIG. 1D, the offset may be added to either the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1.

[0032]

Table 1

[0033]

[0054] The direction index represents the direction of the MVD with respect to the starting point. The direction index can represent the four directions shown in Table 2.

[0034]

Table 2

[0035]

[0055] Note that the meaning of the MVD sign may vary depending on the information of the starting MV. When the starting MV is a single-prediction MV or a bi-prediction MV and both lists point to the same side of the current picture, the signs in Table 2 specify the signs of the MV offsets added to the starting MV. As an example, when the POCs of both references are either greater than the POC of the current picture or both less than the POC of the current picture. When the starting MV is a bi-prediction MV in a state where the two MVs point to different sides of the current picture and the difference in POC in list 0 is greater than the difference in POC in list 1, the signs in Table 2 specify the signs of the MV offsets added to the MV component of the starting MV in list 0, and the signs of the MVs in list 1 have opposite values. As an example, when the POC of one reference is greater than the POC of the current picture and the POC of the other reference is less than the POC of the current picture, it is the sign of the MV offset added to the MV component of the starting MV in list 0, and the signs of the MVs in list 1 have opposite values. Otherwise, if the difference in POC in list 1 is greater than that in list 0, the signs in Table 2 specify the signs of the MV offsets added to the MV component of the starting MV in list 1, and the signs of the MVs in list 0 have opposite values.

[0036]

[0056] The MVD may be scaled according to the difference in POC in each direction. If the difference in POC in both lists is the same, scaling may not be performed. Otherwise, if the difference in POC in list 0 is greater than the difference in POC in list 1, the MVD of list 1 may be scaled. If the difference in POC of L1 is greater than that of L0, the MVD of list 0 may also be scaled in the same way. When the starting MV is singly predicted, the MVD may be added to the available MV.

[0037]

[0057] Symmetric MVD Coding

[0058] In VVC, in addition to the MVD signaling for normal unidirectional prediction and bidirectional prediction modes, a symmetric MVD mode for bidirectional MVD signaling may be applied. In the symmetric MVD mode, the motion information including both the reference picture indices of list 0 and list 1 and the MVD of list 1 may be derived instead of being signaled.

[0038]

[0059] The decoding process of the symmetric MVD mode at the slice level may be as follows. At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived. If mvd_l1_zero_flag is 1, BiDirPredFlag may be set to be equal to 0. Otherwise, if the closest reference picture in list 0 and the closest reference picture in list 1 form a pair of the forward and reverse directions of the reference picture, or a pair of the reverse and forward directions of the reference picture, BiDirPredFlag may be set to 1, and both the reference pictures of list 0 and list 1 may be short-term reference pictures. Otherwise, BiDirPredFlag is set to 0.

[0039]

[0060] The decoding process of the symmetric MVD mode at the CU level can be as follows. At the CU level, the CU is bi-predicted coded, and when BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether the symmetric mode can be used is explicitly signaled. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indices for list 0 and list 1 are set to be equal to the pair of reference pictures, respectively. MVD1 is set to be equal to (-MVD0).

[0040]

[0061] Inter-mode coding in CWG-B018

[0062] In AV1, for each coded block between frames, when the mode of the current block is an inter-coding mode rather than a skip mode, another flag may be signaled to indicate whether a single-reference mode or a composite-reference mode is used for the current block.

[0041]

[0063] The single mode may include a predicted block generated by one motion vector in the single-reference mode. In the case of single-reference, the following modes may be signaled: (1) using one of the motion vector predictors (MVPs) in the list indicated by the NEARMV - DRL (dynamic reference list) index; (2) using one of the motion vector predictors (MVPs) in the list signaled by the NEWMV - DRL index as a reference and applying a delta to the MVP; and (3) using a motion vector based on the global motion parameters at the frame level of GLOBALMV.

[0042]

[0064] The predicted block generated by weighted-averaging two predicted blocks may be derived from two motion vectors in the composite reference mode. In the case of composite reference, the following modes may be signaled: (1) Use one of the motion vector predictors (MVPs) in the list signaled by the NEAR_NEARMV - DRL index; (2) Use one of the motion vector predictors (MVPs) in the list signaled by the NEAR_NEWMV - DRL index as a reference and send a delta MV for the second MV; (3) Use one of the motion vector predictors (MVPs) in the list signaled by the NEW_NEARMV - DRL index as a reference and send a delta MV for the first MV; (4) Use one of the motion vector predictors (MVPs) in the list signaled by the NEW_NEWMV - DRL index as a reference and send delta MVs for both MVs; and (5) Use the MVs from each reference based on the global motion parameters at the frame level of GLOBAL_GLOBALMV.

[0043]

[0065] Motion Vector Difference Coding in AV1

[0066] In AV1, 1 / 8 pixel motion vector accuracy (or precision) is possible, and the following syntax may be used to signal the motion vector difference within reference frame list 0 or list 1. (1) mv_joint specifies which components of the motion vector difference are non-zero. 0 indicates that there is no non-zero MVD along either the horizontal or vertical direction, 1 indicates that there is a non-zero MVD only along the horizontal direction, 2 indicates that there is a non-zero MVD only along the vertical direction, and 3 indicates that there is a non-zero MVD along both the horizontal and vertical directions. (2) mv_sign specifies whether the motion vector difference is positive or negative. (3) mv_class specifies the class of the motion vector difference. As shown in Table 3, the higher the class, the larger the magnitude of the motion vector difference. (4) mv_bit specifies the integer part of the offset between the motion vector difference and the magnitude of the start of each MV class. (5) mv_fr specifies the first two fractional bits of the motion vector difference, and (6) mv_hp specifies the third fractional bit of the motion vector difference.

[0044]

Table 3

[0045]

[0067] Adaptive MVD Resolution in CWG-B092

[0068] In the case of NEW_NEARMV and NEAR_NEWMV modes, the accuracy of the MVD may depend on the related class and the size of the MVD. First, a fractional MVD is only permitted when the size of the MVD is 1 pixel or less. Second, when the value of the related MV class is MV_CLASS_1 or higher, only one MVD value is permitted. The MVD values in each MV class are derived as 4, 8, 16, 32, 64 for MV class 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5). In addition, when the current block is coded as the NEW_NEARMV or NEAR_NEWMV mode, a certain context may be used to signal mv_joint or mv_class. Otherwise, another context may be used to signal mv_joint or mv_class.

[0046]

Table 4

[0047]

[0069] Joint MVD coding (JMVD) in CWG-B092

[0070] To indicate whether the MVDs of the two reference lists are jointly signaled, a new inter-coding mode named JOINT_NEWMV may be applied. When the inter-prediction mode is equal to the JOINT_NEWMV mode, the MVDs of reference list 0 and reference list 1 are jointly signaled. Therefore, only one MVD named joint_mvd is signaled and sent to the decoder, and from joint_mvd, the delta MVs of reference list 0 and reference list 1 are derived.

[0048]

[0071] The JOINT_NEWMV mode is signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. No additional context is added.

[0049]

[0072] When the JOINT_NEWMV mode is signaled and the POC distances between the two reference frames and the current frame are different, the MVD is scaled for reference list 0 or reference list 1 based on the POC distance. Specifically, the distance between reference frame list 0 and the current frame can be denoted as td0, and the distance between reference frame list 1 and the current frame is shown as td1. If td0 is greater than or equal to td1, joint_mvd is used as it is for reference list 0, and the mvd of reference list 1 is derived from joint_mvd based on Equation (1).

[0050]

Equation

[0051]

[0073] Otherwise, if td1 is greater than or equal to td0, joint_mvd is used as it is for reference list 1, and the mvd of reference list 0 is derived from joint_mvd based on Equation (2).

[0052]

Equation

[0053]

[0074] Improvement of Adaptive MVD Resolution in CWG-C011

[0075] In the case of a single reference, a new inter-coding mode named AMVDMV may be added. When the AMVDMV mode is selected, the AMVDMV mode indicates that AMVD is applied to signal the MVD. To indicate whether AMVD is applied to the joint MVD coding mode, one flag named amvd_flag is added under the JOINT_NEWMV mode. When the adaptive MVD resolution is applied to the joint MVD coding mode named joint AMVD coding, the MVDs of two reference frames may be signaled jointly, and the accuracy of the MVD is implicitly determined by the size of the MVD. Otherwise, the MVDs of two (or three or more) reference frames are signaled jointly and the conventional MVD coding is applied.

[0054]

[0076] Adaptive motion vector resolution (AMVR) in CWG-C012 and CWG-C020

[0077] AMVR was first proposed in CWG-C012, and a total of seven MV accuracies (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8) are supported. For each prediction block, the AVM encoder may explore all the supported accuracy values and signal the best accuracy to the decoder.

[0055]

[0078] To reduce the encoder's execution time, two accuracy sets are supported. Each accuracy set contains four pre-defined accuracies. The accuracy set may be adaptively selected at the frame level based on the value of the maximum accuracy of the frame. Similar to AV1, the maximum accuracy may be signaled in the frame header. Table 5 summarizes the supported accuracy values based on the maximum accuracy at the frame level.

[0056]

Table 5

[0057]

[0079] Current AVM software (similar to AV1) has a frame-level flag indicating whether the MV of a frame includes sub-pel accuracy. AMVR is enabled only when the value of the cur_frame_force_integer_mv flag is 0. In AMVR, when the accuracy of a block is lower than the maximum accuracy, the motion model and interpolation filter may not need to be signaled. When the accuracy of a block is lower than the maximum accuracy, the motion mode may be presumed to be translational motion, and the interpolation filter may be presumed to be the REGULAR interpolation filter. Similarly, when the accuracy of a block is 4 pels or 8 pels, the inter-intra mode is not signaled and is inferred to be 0.

[0058]

[0080] Motion vector predictor lists in AV1 and AVM

[0081] Spatial motion vector predictors (SMVP: spatial motion vector predictor), both adjacent and non-adjacent SMVPs), temporal motion vector predictors, additional MV candidates in AV1 and additional derived MVPs, and reference bank MVPs are added to the AVM design. To store the MVPs, a fixed-size stack known as the motion vector predictor list is generated on both the encoder side and the decoder side.

[0059]

[0082] Spatial motion vector predictor (SMVP)

[0083] The spatial motion vector (MV) predictor is derived from spatial neighboring blocks including adjacent spatial neighboring blocks which are the direct neighborhood above and to the left of the current block, and non-adjacent spatial neighboring blocks which are close to the current block but not directly adjacent to the current block. An example of a set of spatial neighboring blocks for a luminance block is shown in FIG. 1E, and each spatial neighboring block is an 8×8 block. The spatial neighboring blocks are examined to find one or more MVs associated with the same reference frame index as the current block. For the current block, the search order of the 8×8 luminance blocks in the spatial neighborhood is as shown by numbers 1 to 8 in FIG. 5. (1) The upper adjacent row is checked from left to right, (2) the left adjacent column is checked from top to bottom, (3) the upper right neighboring block is checked, (4) the upper left block neighboring block is checked, (5) the first non-adjacent upper row is checked from left to right, (6) the first non-adjacent left column is checked from top to bottom, (7) the second non-adjacent upper row is checked from left to right, (8) the second non-adjacent left column is checked from top to bottom.

[0060]

[0084] Adjacent candidates (1 to 3 in FIG. 1E) are put into the MV predictor list before the TMVP, and non-adjacent candidates (also known as outer candidates, i.e., candidates 4 to 8 in FIG. 1E) are put into the MV predictor list after the TMVP. All SMVP candidates should have the same reference picture as the current block. That is, if the current block has a single reference picture, an MVP candidate with a single reference picture and this reference picture is the same as the reference picture of the current block, or an MVP candidate with a composite reference picture (two reference pictures) and one of the reference pictures is the same as the reference picture of the current block, this MVP candidate will be put into the MV predictor list. If the current block has two reference pictures, this MVP candidate is put into the predictor list only when the MVP candidate with two reference pictures and these two reference pictures are the same as the reference pictures of the current block.

[0061]

[0085] Temporal Motion Vector Predictor (TMVP)

[0086] In addition to the spatial neighboring blocks, an MV predictor known as a temporal MV predictor may also be derived using blocks at the same position within the reference frame. To generate the temporal MV predictor, first, the MVs of the reference frames are stored together with the reference indices associated with their respective reference frames. Then, for each 8×8 block in the current frame, the MVs of the reference frames through which the trajectory passes through the 8×8 block are identified and stored in the temporal MV buffer together with the reference frame indices. In the case of inter prediction using a single reference frame, for performing the temporal motion vector prediction of the future frame, regardless of whether the reference frame is a forward reference frame or a backward reference frame, the MVs are stored in 8×8 units. In the case of composite inter prediction, for performing the temporal motion vector prediction of the future frame, only the forward MVs are stored in 8×8 units.

[0062]

[0087] Referring to FIG. 1G, the MV of reference frame 1 (R1) 1620, i.e., MVref 1650, is indicated from frame 1 (R1) 1620. In so doing, MVref 1650 passes through an 8×8 block (the block within the current frame 1615 and the block within reference frame 0 1610 of the current frame). MVref is stored in the temporal MV buffer in association with this 8×8 block. During the motion projection process for deriving the temporal MV predictor, the reference frames may be scanned in a predefined order, i.e., in the order of LAST_FRAME, BWDREF_FRAME, ALTREF_FRAME, ALTREF2_FRAME, and LAST2_FRAME. The MVs from the reference frames with higher indices (in the scanning order) do not replace the previously identified MVs assigned by the reference frames with lower indices (in the scanning order).

[0063]

[0088] When given a pre - defined block coordinate, to derive a temporal MV predictor, e.g., MV0 in FIG. 1F, which indicates the reference frame from the current block, the relevant MVs stored in the temporal MV buffer are identified and projected onto the current block.

[0064]

[0089] Referring to FIG. 1G, the pre - defined block positions for deriving the temporal MV predictor of a 16×16 block are shown. For a valid temporal MV predictor, up to seven blocks are checked. The temporal MV predictor is checked after the adjacent spatial MV predictor and before the non - adjacent spatial MV predictor.

[0065]

[0090] In the derivation of the MV predictor, all spatial and temporal MV candidates may be pooled, and each predictor may be assigned a weight determined during the scan of spatial and temporal neighboring blocks. The candidates may be sorted and ranked based on the associated weights, and up to four candidates are identified and added to the MV predictor list. This list of MV predictors, also called the dynamic reference list (DRL), is further used in the dynamic MV prediction mode as described in the next sub - section.

[0066]

[0091] Additional search MVPs for additional MVP candidates

[0092] If the MVP list is still not full, additional searches are performed and additional MVP candidates are used to fill the MVP list. The additional MVP candidates include, for example, global MVs, zero MVs, combined composite MVs without scaling, etc.

[0067]

[0093] The process of sorting MVP candidates

[0094] The adjacent SMVP candidates, TMVP candidates, and non - adjacent SMVP candidates added to the MVP list are sorted. Based on the current designs in AV1 and AVM, the sorting process is based on the weight of each candidate. The weight of a candidate is pre - defined according to the overlapping area between the current block and the candidate block.

[0068]

[0095] Derived MVP candidates

[0096] The derived MVP candidates are adopted in the AVM reference software according to Proposal CWG-B049, which includes both the MVP derived for a single reference picture and the MVP derived for the combined mode.

[0069]

[0097] Single inter prediction

[0098] If the reference frames of neighboring blocks are different from the reference frame of the current block but in the same direction, a temporal scaling algorithm can be used to scale the MV to match that reference frame in order to form the MVP of the motion vector of the current block. As shown in Figure 1H, the MVP of the motion vector mv0 1850 of the current block with temporal scaling can be derived using mv1 1855 from a neighboring block (the shaded block).

[0070]

[0099] Combined inter prediction

[0100] To derive the MVP of the current block, composite MVs from different neighboring blocks are utilized, but the reference frames of the composite MVs need to be the same as that of the current block. As shown in Figure 1I, the composite MVs (mv2 1960, mv3 1965) have the same reference frame as the current block, but these reference frames are from different neighboring blocks.

[0071]

[0101] Reference motion vector candidate bank

[0102] Each buffer corresponds to a unique reference frame type that covers a single reference frame or a pair of reference frames for single inter mode and combined inter mode respectively. All buffers are of the same size. When a new MV is added to a full buffer, existing MVs may be removed to make space for the new MV.

[0072]

[0103] For collecting reference MV candidates, the coding block may refer to the MV candidate bank in addition to the reference MV candidates obtained in the conventional AV1 reference MV list generation. After coding the superblock, the MV bank may be updated with the MVs used by the coding blocks of the superblock.

[0073]

[0104] Each tile may have an independent MV reference bank that can be utilized by all superblocks within the tile. At the start of encoding for each tile, the corresponding bank may be emptied. Thereafter, when coding each superblock within that tile, MVs from the bank may be used as MV reference candidates. At the end of encoding the superblock, the bank may be updated.

[0074]

[0105] Bank update

[0106] As shown in the graphic 200 of FIG. 2, the bank update process may be based on the superblock. That is, after the superblock is coded, the first (up to 64) candidate MVs used by each coding block within the superblock are added to the bank. During the update, the pruning process may also be involved during the update.

[0075]

[0107] Bank reference

[0108] After scanning the reference MV candidates of the conventional AV1 or the new AV2, if there are empty slots in the candidate list, the codec may refer to the MV candidate bank for additional MV candidates (within the buffer with matching reference frame types). While proceeding in the reverse direction from the end to the beginning of the buffer, if the MV in the bank buffer does not yet exist in the candidate list, that MV may be added to the candidate list.

[0076]

[0109] MVP list construction process in the state-of-the-art design

[0110] In the related art, as shown in flowchart 300 of FIG. 3, the MVP list may be constructed using complete pruning. One example may include operation 305 of adding adjacent SMVPs, operation 310 of adding TMVP, operation 315 of adding non-adjacent SMVPs, operation 320 of adding a sorting process for existing candidates, operation 325 of adding derived candidates, operation 330 of adding additional MVPs, and finally, operation 355 of adding candidates from the reference MV candidate bank.

[0077]

[0111] FIG. 4 shows a simplified block diagram of a communication system 400 according to an embodiment of the present disclosure. The communication system 400 may include at least two terminals 410 to 420 interconnected via a network 450. In the case of unidirectional data transmission, the first terminal 410 may code video data at a local location for transmission to the other terminal 420 via the network 450. The second terminal 420 may receive the coded video data of the other terminal from the network 450, decode the coded data, and display the restored video data. Unidirectional data transmission may be common in media serving applications and the like.

[0078]

[0112] FIG. 4 shows a second pair of terminals 430, 440 provided to support bidirectional transmission of coded video that may occur, for example, during a video conference. In the case of bidirectional data transmission, each terminal 430, 440 may code the captured video data at a local location for transmission to the other terminal via the network 450. Each terminal 430, 440 may also receive the coded video data transmitted by the other terminal, decode the coded data, and display the restored video data on a local display device.

[0079]

[0113] In FIG. 4, terminals 410 to 440 are shown as a server, a personal computer, and a smartphone, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure find use in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 450 represents any number of networks that transfer encoded video data between terminals 410 to 440, including, for example, wired and / or wireless communication networks. Communication network 450 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of network 450 may not be important for the operation of the present disclosure, unless otherwise described below.

[0080]

[0114] FIG. 5 shows the arrangement of a video encoder and a video decoder in a streaming environment, such as a streaming system 500, as an example of the use of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storage of compressed video on digital media such as CDs, DVDs, memory sticks, and the like.

[0081]

[0115] The streaming system can include a capture subsystem 513, which can include a video source 501, such as a digital camera, that creates, for example, an uncompressed video sample stream 502. The sample stream 502 is illustrated as a thick line to emphasize that it has a large amount of data when compared to an encoded video bit stream and can be processed by an encoder 503 coupled to the camera 501. The encoder 503 can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video bit stream 504 is illustrated as a thin line to emphasize that it has a small amount of data when compared to the sample stream and can be stored on a streaming server 505 for future use. One or more streaming clients 506, 508 can access the streaming server 505 to obtain copies 507, 509 of the encoded video bit stream 504. The client 506 can include a video decoder 510 that decodes a copy of the incoming encoded video bit stream 507 and creates an outgoing video sample stream 511 that can be rendered on a display 512 or other rendering device not shown. In some streaming systems, the video bit streams 504, 507, 509 can be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. A video coding standard, informally known as Versatile Video Coding (VVC), is under development. The disclosed subject matter may be used in the context of VVC.

[0082]

[0116] FIG. 6 can be a functional block diagram of a video decoder 510 according to an embodiment of the present invention.

[0083]

[0117] Receiver 610 may receive one or more codec video sequences decoded by decoder 510. In the same or another embodiment, one coded video sequence is decoded at a time, in which case the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequence may be received from channel 612, which may be a hardware / software link to a storage device storing the encoded video data. Receiver 610 may receive the encoded video data along with other data that may be transferred to respective consuming entities (not shown), such as coded audio data and / or auxiliary data streams. Receiver 610 may separate the coded video sequence from other data. To handle network jitter, buffer memory 615 may be coupled between receiver 610 and entropy decoder / parser 620, hereinafter “parser”. If receiver 610 is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous network, buffer 615 may not be needed or may be small. When used in a best-effort packet network such as the Internet, buffer 615 may be needed, may be relatively large, and advantageously may be of an adaptable size.

[0084]

[0118] Video decoder 510 may include an analyzer 620 for reconstructing symbol 621 from an entropy-coded video sequence. The categories of these symbols include information used to manage the operation of decoder 510 and, in some cases, information for controlling a rendering device such as display 512, which is not an essential part of the decoder but can be coupled to the decoder as shown in FIG. 6. The control information for the rendering device may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment not shown. Analyzer 620 may analyze / entropy-decode the received coded video sequence. The coding of the coded video sequence can conform to a video coding technology or standard and can follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding, etc., regardless of the presence or absence of context dependence. Analyzer 620 may extract a set of at least one subgroup parameter of at least one subgroup of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to a group. The subgroups can include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / analyzer may also extract information such as transform coefficients, quantization parameter (QP) values, motion vectors, etc. from the coded video sequence.

[0085]

[0119] The parser 620 may perform entropy decoding / parsing operations on the video sequence received from the buffer 615 to create the symbol 621. The parser 620 may receive the encoded data and selectively decode a specific symbol 621. Further, the parser 620 may determine whether a specific symbol 621 should be provided to the motion compensation prediction unit 653, the scaler / inverse transform unit 651, the intra prediction unit 652, or the loop filter 656.

[0086]

[0120] For the reconstruction of the symbol 621, multiple different units may be involved depending on the type of the coded video picture or a part thereof, such as inter pictures and intra pictures, inter blocks and intra blocks, and other factors. How each unit is involved may be controlled by subgroup control information parsed by the parser 620 from the coded video sequence. For the sake of brevity, the flow of such subgroup control information between the parser 620 and the following multiple units is not illustrated.

[0087]

[0121] In addition to the functional blocks already described, the decoder 510 may be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate.

[0088]

[0122] The first unit is the scaler / inverse transform unit 651. The scaler / inverse transform unit 651 receives from the parser 620, as the symbol 621, the quantized transform coefficients and control information including which transform to use, block size, quantization coefficient, quantization scaling matrix, etc. The scaler / inverse transform unit 651 can output a block including sample values that can be input to the aggregator 655.

[0089]

[0123] In some cases, the output samples of the scaler / inverse transform 651 may relate to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures, but can use prediction information from previously reconstructed parts of the current picture. The intra-picture prediction unit 652 can provide such prediction information. In some cases, the intra-picture prediction unit 652 uses the surrounding already reconstructed information fetched from the current partially reconstructed picture 658 to generate a block of the same size and shape as the block being reconstructed. The aggregator 655, in some cases, adds, sample by sample, the prediction information generated by the intra prediction unit 652 to the output sample information provided by the scaler / inverse transform unit 651.

[0090]

[0124] In other cases, the output samples of the scaler / inverse transform unit 651 may relate to inter-coded, and possibly motion-compensated, blocks. In such cases, the motion-compensation prediction unit 653 can access the reference picture memory 657 to fetch the samples used for prediction. After motion-compensating the fetched samples according to the symbols 621 related to the block, the aggregator 655 can, in this case, add these samples, called residual samples or residual signals in this case, to the output of the scaler / inverse transform unit to generate the output sample information. The address in the reference picture memory from which the motion-compensation unit fetches the prediction samples can be controlled by a motion vector, in a form, for example, of the symbol 621 that can have X, Y, and reference picture components and that is available to the motion-compensation unit. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory when sub-sample accurate motion vectors are used, a motion vector prediction mechanism, etc.

[0091]

[0125] The output samples of the aggregator 655 can be subject to various loop filtering techniques in the loop filter unit 656. The video compression technology can include in-loop filter technology, and the in-loop filter technique is controlled by parameters that the loop filter unit 656 can use as symbols 621 from the parser 620 included in the coded video bitstream, but also corresponds to meta information obtained during the decoding of the previous part in the decoding order of the coded picture or the coded video sequence, and can also correspond to previously reconstructed and loop filter processed sample values.

[0092]

[0126] The output of the loop filter unit 656 can be a sample stream that is output to the render device 512 and stored in the reference picture memory 658 for use in future inter-picture prediction.

[0093]

[0127] When a specific coded picture is completely reconstructed, it can be used as a reference picture for future prediction. When the coded picture is completely reconstructed and the coded picture is identified as a reference picture by, for example, the parser 620, the current reference picture 658 can become part of the reference picture buffer 657, and a new current picture memory can be reallocated before starting the reconstruction of the next coded picture.

[0094]

[0128] The video decoder 510 may perform a decoding operation according to a predetermined video compression technique that can be documented in a standard such as ITU-T Rec.H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence conforms to the syntax of the video compression technique or standard specified in the document or standard of the video compression technique, particularly the profile document therein. Also, for compliance, it may be necessary that the complexity of the coded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, for example, the maximum reconstructed sample rate measured in megasamples per second, the maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted through the specifications of a Hypothetical Reference Decoder (HRD) and the metadata for HRD buffer management signaled in the coded video sequence.

[0095]

[0129] In one embodiment, the receiver 610 may receive additional redundant data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder 510 to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0096]

[0130] FIG. 7 may be a functional block diagram of the video encoder 503 according to an embodiment of the present invention.

[0097]

[0131] The encoder 503 may receive video samples from a video source 501 that can capture a video image coded by the encoder 503 and that is not part of the encoder.

[0098]

[0132] The video source 501 may provide a source video sequence coded by the encoder 503 in the form of a digital video sample stream, and the form of the digital video sample stream can be any suitable bit depth, for example, 8 bits, 10 bits, 12 bits,..., any color space, for example, BT.601 Y CrCB, RGB,..., and any suitable sampling structure, for example, Y CrCb 4:2:0, Y CrCb 4:4:4. In a media serving system, the video source 501 may be a storage device that stores previously prepared video. In a video conferencing system, the video source 503 may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that give movement when viewed in sequence. Each picture itself may be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. A person skilled in the art can easily understand the relationship between pixels and samples. In the following description, samples are focused on.

[0099]

[0133] According to one embodiment, encoder 503 may code and compress pictures of a source video sequence in real time or under any other time constraints in response to requests by an application to obtain a coded video sequence 743. Enforcing an appropriate coding speed is one of the functions of controller 750. The controller controls other functional units and is functionally coupled to these units as will be described below. For simplicity, the couplings are not shown. Parameters set by the controller may include rate control related parameters such as picture skip, quantizer, lambda value of rate distortion optimization techniques, picture size, group of pictures (GOP) layout of pictures, maximum motion vector search range, and the like. Those skilled in the art can readily identify other functions of controller 750 that may be relevant to a video encoder 503 optimized for a particular system design.

[0100]

[0134] Some video encoders operate in what those skilled in the art would readily recognize as a "coding loop." As a very simplified explanation, the coding loop is part of encoder 730, hereinafter the "source coder," which is responsible for creating symbols based on the input picture and reference pictures to be coded, and local decoder 733 incorporated into encoder 503. Local decoder 733 reconstructs the symbols to create sample data, and since in the video compression technology contemplated by the disclosed subject matter the compression between the symbols and the coded video bitstream is reversible, the remote decoder also creates this sample data. The reconstructed sample stream is input into reference picture memory 734. Since decoding the symbol stream results in bit-exact results regardless of whether the decoder is local or remote, the contents of the reference picture buffer are also bit-exact between the local encoder and the remote encoder. In other words, the prediction portion of the encoder "sees" the same sample values as the sample values that the decoder will "see" when using the prediction during decoding as the reference picture samples. This basic principle of reference picture synchronization and the drift that results when synchronization cannot be maintained, for example due to channel errors, is well known to those skilled in the art.

[0101]

[0135] The operation of the "local" decoder 733 can be considered the same as the operation of the "remote" decoder 510 already described in detail in connection with FIG. 6. However, referring briefly to FIG. 7, since symbols are available and the encoding / decoding of symbols into the coded video sequence by entropy encoder 745 and parser 620 can be reversible, the entropy decoding portion of decoder 510, including channel 612, receiver 610, buffer 615, and parser 620, may not be fully implemented in local decoder 733.

[0102]

[0136] At this point, it can be said that any decoder technology existing within the decoder, except for parsing / entropy decoding, must similarly necessarily exist in substantially the same functional form within the corresponding encoder. Since the description regarding encoder technology is the reverse of the comprehensively described decoder technology, it can be omitted. More detailed explanations are required only in specific areas, which are provided below.

[0103]

[0137] As part of its operation, the source coder 730 may perform motion compensation prediction coding, which predictively codes an input frame by referring to one or more previously coded frames from the video sequence designated as the "reference frame". In this way, the coding engine 732 codes the difference between a pixel block of the input frame and a pixel block of the reference frame that can be selected as a prediction reference for the input frame.

[0104]

[0138] The local video decoder 733 may decode the coded video data of a frame that can be designated as a reference frame based on the symbols created by the source coder 730. The operation of the coding engine 732 is, advantageously, an irreversible process. When the coded video data can be decoded in a video decoder not shown in FIG. 7, the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder 733 may replicate the decoding process that can be performed by the video decoder for the reference frame and store the reconstructed reference frame in the reference picture cache 734. In this way, the encoder 503 may locally store a copy of the reconstructed reference frame having common content as the reconstructed reference frame obtained without transmission error by the remote video decoder.

[0105]

[0139] Predictor 735 may perform a prediction search for the coding engine 732. That is, for a new frame to be coded, predictor 735 may search the reference picture memory 734 for sample data as candidate reference pixel blocks, or specific metadata such as reference picture motion vectors and block shapes, that may function as appropriate prediction references for the new picture. Predictor 735 may operate on a per sample block x pixel block basis to find an appropriate prediction reference. In some cases, the input picture may have prediction references drawn from a plurality of reference pictures stored in the reference picture memory 734, as determined by the search results obtained by predictor 735.

[0106]

[0140] Controller 750 may manage the coding operations of video coder 730, including setting parameters and subgroup parameters used to encode video data.

[0107]

[0141] The output of all of the foregoing functional units may be subject to entropy coding in entropy coder 745. The entropy coder converts the symbols generated by the various functional units into a coded video sequence by reversibly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, and the like.

[0108]

[0142] The transmitter 740 may buffer the coded video sequence created by the entropy coder 745 to prepare the coded video sequence for transmission via the communication channel 760, which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter 740 may merge the coded video data from the video coder 730 with other data to be transmitted, such as coded audio data and / or an auxiliary data stream source (not shown).

[0109]

[0143] The controller 750 may manage the operation of the encoder 503. During coding, the controller 750 may assign a specific coded picture type to each coded picture, which may affect the coding technique applicable to each picture. For example, often a picture may be assigned as one of the following frame types.

[0110]

[0144] An intra picture, an I picture, may be a picture that can be coded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of these variations of I pictures, as well as their respective uses and characteristics.

[0111]

[0145] A predicted picture, a P picture, may be a picture that can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0112]

[0146] Bidirectional predicted pictures, B pictures, may be pictures that can be coded and decoded using intra prediction or inter prediction that use up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use three or more reference pictures and associated metadata for the reconstruction of a single block.

[0113]

[0147] The source picture may generally be spatially subdivided into a plurality of sample blocks, for example blocks of 4×4, 8×8, 4×8, or 16×16 samples each, and coded in block units. The blocks may be coded predictively with reference to other already-coded blocks as determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively, or blocks of an I picture may be coded predictively with reference to already-coded blocks of spatial prediction or intra prediction of the same picture. Pixel blocks of a P picture may be coded non-predictively via spatial prediction or temporal prediction with reference to one previously-coded reference picture. Blocks of a B picture may be coded non-predictively via spatial prediction or temporal prediction with reference to one or two previously-coded reference pictures.

[0114]

[0148] The video coder 503 may perform a coding operation in accordance with a predetermined video coding technology or standard such as ITU-T Rec.H.265. In that operation, the video coder 503 may perform various compression operations including a predictive coding operation that exploits temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to the syntax specified by the video coding technology or standard being used.

[0115]

[0149] In one embodiment, transmitter 740 may transmit additional data along with the encoded video. Video coder 730 may include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and redundant slices, supplementary enhancement information (SEI) messages, visual user usability information (VUI) parameter set fragments, and the like.

[0116]

[0150] FIG. 8 shows an exemplary process 800 for updating a bank of reference motion vector candidates.

[0117]

[0151] In operation 805, one or more motion vector predictors associated with the current block may be obtained from the reference motion vector candidate bank. The one or more motion vector predictors associated with the current block obtained from the reference motion vector candidate bank may include at least one or more motion vectors associated with one or more already decoded blocks, and the one or more already decoded blocks may belong to the same superblock as the current block. In some embodiments, the one or more motion vector predictors associated with the current block obtained from the reference motion vector candidate bank are obtained based on a reversed reference order. This indication of the reference order may be based on a condition being satisfied. The condition may be based on one of the size of the current block, one or more quantization parameter values, or the hash hit rate of the screen content.

[0118]

[0152] In operation 810, based on one or more motion vector predictors obtained, a motion vector associated with the current block may be determined. In operation 815, based on the determined motion vector, the current block may be decoded. In some embodiments, after decoding of the current block, another block within the same super block may be decoded based on the motion vector associated with the current block inserted into the reference motion vector candidate bank. In operation 820, the reference motion vector candidate bank may be updated by inserting the motion vector associated with the current block into the reference motion vector candidate bank.

[0119]

[0153] In some embodiments, in operation 825, updating the reference motion vector candidate bank may include inserting the motion vector associated with the current block into the reference motion vector candidate bank after decoding all blocks within the same super block.

[0120]

[0154] In operation 830, in some embodiments, updating the reference motion vector candidate bank may include inserting the motion vector associated with the current block into the reference motion vector candidate bank after decoding all blocks within a predefined level of the quadtree associated with the current block. The predefined level of the quadtree may be the first level of the quadtree after the super block level. The predefined level of the quadtree may be signaled in high level syntax.

[0121]

[0155] According to one embodiment, in operation 835, updating the reference motion vector candidate bank may include inserting the motion vector associated with the current block into the reference motion vector candidate bank after decoding all blocks within a predefined region around the current block. The predefined region around the current block may include a block size or a fixed region such as 64×64, 32×64.

[0122]

[0156] In some embodiments, operations 805-835 may include a reference motion vector candidate bank that is a first reference motion vector candidate bank and another reference motion vector candidate bank that is a second reference motion vector candidate bank. Updating the reference motion vector may include updating the second reference motion vector candidate bank by inserting the motion vector associated with the current block into the second reference motion vector candidate bank after decoding all the blocks within a predefined level of the quadtree associated with the current block.

[0123]

[0157] In some embodiments, based on the reference motion vector candidate bank being the first reference motion vector candidate bank and another reference motion vector candidate bank being the second reference motion vector candidate bank, in operation 815, if the size of the current block being decoded may be greater than a threshold, the current block may be decoded based on one or more first motion vectors within the first reference motion vector candidate bank. In the same or another embodiment, in operation 815, if the size of the current block may be less than or equal to the threshold, the current block may be decoded based on one or more second motion vectors within the second reference motion vector candidate bank.

[0124]

[0158] According to one embodiment, when performing an operation among operations 825-835, a flag that may be associated with the current block indicating whether a motion vector from the first reference motion vector candidate bank or a motion vector from the second reference motion vector candidate bank should be used may be used.

[0125]

[0159] Figure 8 shows exemplary blocks of process 800, but in some implementations, process 800 may include additional blocks, fewer blocks than those illustrated in Figure 8, blocks different from those illustrated in Figure 8, or blocks arranged in a different way from those illustrated in Figure 8. Additionally or alternatively, two or more of the blocks of process 800 may be executed in parallel.

[0126]

[0160] Further, the proposed method may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium to execute one or more of the proposed methods.

[0127]

[0161] The techniques described above may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media. For example, Figure 9 shows a computer system 900 suitable for implementing a particular embodiment of the disclosed subject matter.

[0128]

[0162] The computer software may include code that contains instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or that can be executed through interpretation, microcode execution, and can be coded using any suitable machine code or computer language that can be the subject of an assembly, compilation, linking, or similar mechanism.

[0129]

[0163] The instructions may be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.

[0130]

[0164] The components shown in FIG. 9 for computer system 900 are essentially exemplary and do not imply any limitations regarding the use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of computer system 900.

[0131]

[0165] Computer system 900 may include specific human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, through tactile input (keystrokes, swipes, movement of a data glove, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). The human interface device may also be used to capture certain media that is not necessarily directly related to conscious human input, such as audio (speech, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.).

[0132]

[0166] The input human interface device may include one or more of keyboard 901, mouse 902, trackpad 903, touch screen 910, data glove 1204, joystick 905, microphone 906, scanner 907, camera 908 (only one of each is shown).

[0133]

[0167] The computer system 900 may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback by a touch screen 910, a data glove 1204, or a joystick 905, although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers 909, headphones (not shown), etc.), visual output devices (such as screens 910, including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light emitting diode (OLED) screens, regardless of whether each has touch screen input capabilities and regardless of whether each has tactile feedback capabilities, and some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher output through means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown) may also be included.

[0134]

[0168] The computer system 900 can also include human-accessible storage devices and associated media such as an optical medium including a CD / DVD ROM / RW 920 having a CD / DVD or similar medium 921, a thumb drive 922, a removable hard drive or solid state drive 923, legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.

[0135]

[0169] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.

[0136]

[0170] The computer system (900) can also include an interface to one or more communication networks (955). The network (955) can be, for example, a wireless network, a wired network, or an optical network. The network (955) can further be a local network, a wide area network, a metropolitan network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of the network (955) include local area networks such as Ethernet and wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital networks for TV including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial networks including CANBus, etc. A specific network (955) usually requires an external network interface adapter (954) attached to a specific general-purpose data port or peripheral bus (949) (such as a USB port of the computer system (900)), and other networks are usually integrated into the core of the computer system (900) by attaching to the system bus as described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). The computer system (900) can communicate with other entities using any of these networks (955). Such communication can be only unidirectional reception (such as broadcast TV), only unidirectional transmission (such as CANbus to a specific CANbus device), or bidirectional with other computer systems using local or wide area digital networks. For each of these networks (955) and network interfaces (954) described above, specific protocols and protocol stacks can be used.

[0137]

[0171] The foregoing human interface device, human-accessible memory device, and network interface can be attached to the core 940 of the computer system 900.

[0138]

[0172] The core 940 can include one or more central processing units (CPUs) 941, a graphics processing unit (GPU) 942, a specialized programmable processing unit 943 in the form of a field programmable gate array (FPGA), a hardware accelerator 944 for specific tasks, and the like. These devices may be connected through a system bus 1248 together with a read-only memory (ROM) 945, a random access memory (RAM) 946, an internal hard drive that is not accessible to the user, an internal mass storage 947 such as a solid state drive (SSD), and the like. In some computer systems, the system bus 1248 can be made accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, and the like. Peripheral devices can be attached directly to the system bus 1248 of the core or through a peripheral bus 949. Architectures for the peripheral bus include Peripheral Component Interconnect (PCI), USB, and the like.

[0139]

[0173] The CPU 941, GPU 942, FPGA 943, and accelerator 944 can execute specific instructions that can together constitute the foregoing computer code. The computer code can be stored in the ROM 945 or the RAM 946. Migration data can also be stored in the RAM 946, while persistent data can be stored, for example, in the internal mass storage 947. Fast storage and retrieval for any memory device can be enabled through the use of a cache memory that can be closely associated with one or more CPUs 941, GPUs 942, mass storage 947, ROM 945, RAM 946, and the like.

[0140]

[0174] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or can be of the kind well known and available to those having skill in the computer software arts.

[0141]

[0175] By way of example and not limitation, a computer system 900 having an architecture, specifically a core 940, can provide functionality as a result of software embodied in one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be the user-accessible mass storage introduced above, and media associated with specific storage of the core 940 that is non-transitory in nature, such as the on-core mass storage 947 or ROM 945. The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core 940. The computer-readable media can include one or more memory devices or chips depending on specific needs. The software can cause the core 940, specifically the processor (including a CPU, GPU, FPGA, etc.) therein, to define data structures stored in the RAM 946 and modify such data structures according to processes defined by the software, thereby executing a specific process or a specific portion of a specific process described herein. Additionally, or alternatively, the computer system can provide functionality as a result of logic being hardwired or otherwise embodied within a circuit (e.g., an accelerator 944) that operates instead of or in conjunction with software to execute a specific process or a specific portion of a specific process described herein. References to software can, as necessary, include logic, and vice versa. References to computer-readable media can, as necessary, include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody execution logic, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0142]

[0176] Although the present disclosure has described several exemplary embodiments, there are changes, substitutions, and various alternative equivalents that are included within the scope of the present disclosure. Thus, it will be understood that, although not explicitly illustrated or described herein, many systems and methods that embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure can be devised by those skilled in the art.

Claims

1. A method for decoding video data, which is executed by at least one processor, obtaining a plurality of motion vector predictors associated with a current block from a reference motion vector candidate bank, wherein the plurality of obtained motion vector predictors are associated with a plurality of already decoded blocks, and the plurality of already decoded blocks belong to the same super block as the current block; determining a motion vector associated with the current block based on the plurality of obtained motion vector predictors; updating the reference motion vector candidate bank by inserting the motion vector associated with the current block into the reference motion vector candidate bank; decoding the current block based on the determined motion vector associated with the current block and a method comprising.

2. The method is further comprising decoding another block within the same super block based on the motion vector associated with the current block inserted into the reference motion vector candidate bank, according to the method of Claim 1.

3. The step of updating the reference motion vector candidate bank is including inserting the motion vector associated with the current block into the reference motion vector candidate bank after decoding all blocks within the same super block, according to the method of Claim 1.

4. The step of updating the reference motion vector candidate bank is After decoding all blocks within a pre - defined level of the quadtree associated with the current block, inserting the motion vector associated with the current block into the reference motion vector candidate bank, the method according to claim 1.

5. The method according to claim 4, wherein the pre - defined level of the quadtree is the first level of the quadtree after the super - block level, and the pre - defined level of the quadtree is signaled in high - level syntax.

6. The step of updating the reference motion vector candidate bank comprises After decoding all blocks within a pre - defined area around the current block, inserting the motion vector associated with the current block into the reference motion vector candidate bank, the method according to claim 1.

7. The reference motion vector candidate bank is a first reference motion vector candidate bank, and the method comprises After decoding all blocks within a pre - defined level of the quadtree associated with the current block, inserting the motion vector associated with the current block into a second reference motion vector candidate bank, thereby further comprising the step of updating the second reference motion vector candidate bank, the method according to claim 1.

8. The method according to claim 7, wherein the pre - defined level of the quadtree is the first level of the quadtree after the super - block level.

9. Based on the size of the current block during decoding being larger than a threshold, decoding the current block based on a plurality of first motion vectors within the first reference motion vector candidate bank, or Based on the size of the current block being less than or equal to the threshold, decoding the current block based on a plurality of second motion vectors in the second reference motion vector candidate bank The method according to claim 7, further comprising.

10. The method according to claim 7, further comprising decoding a flag associated with the current block indicating whether a motion vector from the first reference motion vector candidate bank or a motion vector from the second reference motion vector candidate bank should be used.

11. The method according to claim 1, wherein the plurality of obtained motion vector predictors are obtained based on a reversed reference order.

12. The method according to claim 11, wherein the reversed reference order is used to obtain the plurality of obtained motion vector predictors associated with the current block from the reference motion vector candidate bank based on a condition being satisfied.

13. The condition is the size of the current block, a quantization parameter value, or a hash hit rate of screen content The method according to claim 12, based on one of.

14. A device for decoding video data, the device comprising at least one memory configured to store program code, and at least one processor configured to read the program code and operate as commanded by the program code and the program code is Acquisition code configured to cause the at least one processor to acquire a plurality of motion vector predictors associated with a current block from a reference motion vector candidate bank, wherein the plurality of acquired motion vector predictors are associated with a plurality of already decoded blocks, and the plurality of already decoded blocks belong to the same superblock as the current block, and Determination code configured to cause the at least one processor to determine a motion vector associated with the current block based on the plurality of acquired motion vector predictors; and First update code configured to cause the at least one processor to update the reference motion vector candidate bank by inserting the motion vector associated with the current block into the reference motion vector candidate bank; and First decoding code configured to cause the at least one processor to decode the current block based on the determined motion vector associated with the current block A device comprising.

15. The program code is The device according to claim 14, wherein the program code includes first insertion code configured to cause the at least one processor to insert the motion vector associated with the current block into the reference motion vector candidate bank after decoding all blocks within the same superblock.

16. The program code is The device according to claim 14, wherein the program code includes second insertion code configured to cause the at least one processor to insert the motion vector associated with the current block into the reference motion vector candidate bank after decoding all blocks within a predefined level of the quadtree associated with the current block.

17. The program code is The device according to claim 14, further comprising third insertion code configured to cause the at least one processor to insert the motion vector associated with the current block into the reference motion vector candidate bank after decoding all blocks within a predefined area around the current block.

18. where the reference motion vector candidate bank is a first reference motion vector candidate bank, and the program code further comprises second update code configured to cause the at least one processor to update the second reference motion vector candidate bank by inserting the motion vector associated with the current block into the second reference motion vector candidate bank after decoding all blocks within a predefined level of the quadtree associated with the current block.

19. The program code second decoding code configured to cause the at least one processor to decode the current block based on a plurality of first motion vectors within the first reference motion vector candidate bank based on the size of the current block being greater than a threshold during decoding, or third decoding code configured to cause the at least one processor to decode the current block based on a plurality of second motion vectors within the second reference motion vector candidate bank based on the size of the current block being less than or equal to the threshold. The device according to claim 18, further comprising the same.

20. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a device for decoding video data, cause the one or more processors to Obtaining a plurality of motion vector predictors associated with a current block from a reference motion vector candidate bank, wherein the plurality of obtained motion vector predictors are associated with a plurality of already decoded blocks, and the plurality of already decoded blocks belong to the same super block as the current block; Determining a motion vector associated with the current block based on the plurality of obtained motion vector predictors; Updating the reference motion vector candidate bank by inserting the motion vector associated with the current block into the reference motion vector candidate bank; Decoding the current block based on the determined motion vector associated with the current block; A non-transitory computer-readable medium including one or more instructions to cause the above to be performed.

Citation Information

Patent Citations

  • History-based motion vector predictor

    US20200112741A1

  • Inherited motion information for decoding a current coding unit in a video coding system

    WO2020007362A1