Any LUT that should be updated or not updated

By utilizing a method that maintains tables of motion candidates and updates them based on specific rules, the video coding technology addresses the challenge of efficiently managing motion vectors, resulting in improved compression ratios and reduced bandwidth demand.

JP7693632B2Active Publication Date: 2025-06-17DOUYIN VISION CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022195878
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-06-29
Filing Date
2022-12-07
Publication Date
2025-06-17
Estimated Expiration
2039-07-01

AI Technical Summary

Technical Problem

Current video coding technologies face challenges in efficiently managing motion vectors, leading to increased bandwidth demand and reduced coding efficiency, especially as the number of connected user devices increases.

Method used

The proposed method involves maintaining a set of tables containing motion candidates and their associated motion information, which are used for encoding and decoding digital video. This method includes converting video blocks into bit-stream representations and updating tables based on specific update rules to optimize motion vector prediction.

Benefits of technology

This approach enhances the compression ratio of videos by improving the prediction and encoding of motion information, thereby reducing bandwidth demand and increasing coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693632000006
    Figure 0007693632000006
  • Figure 0007693632000007
    Figure 0007693632000007
  • Figure 0007693632000008
    Figure 0007693632000008
Patent Text Reader

Abstract

A method, system, and device for encoding and decoding digital video using merged lists of motion vectors are provided. [Solution] A video decoding method includes maintaining a number of tables, each table containing a set of motion candidates, each motion candidate associated with corresponding motion information derived from a previously coded video block; converting between a current video block and a bitstream representation of the current video block in the video domain; and updating one or more tables based on an update rule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) Based on the applicable patent laws and / or regulations under the Paris Convention, this application aims to timely claim the priority and benefits of International Patent Application No. PCT / CN2018 / 093663 filed on June 29, 2018. Under the laws of the United States, for all purposes, the entire disclosure of International Patent Application No. PCT / CN2018 / 093663 is incorporated by reference as part of the disclosure of this application.

[0002] This patent specification relates to video coding techniques, devices, and systems.

Background Art

[0003] Despite the progress of video compression, digital video still occupies the largest bandwidth usage in the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for the use of digital video is expected to continue to grow.

Summary of the Invention

[0004] This specification discloses a method, system, and device for encoding and decoding digital video using a merge list of motion vectors.

[0005] In one exemplary aspect, a video decoding method is disclosed. The method includes maintaining a number of tables, each table including a set of motion candidates, each motion candidate being associated with corresponding motion information derived from a previously - coded video block; performing a conversion between a current video block and the bit - stream representation of the current video block in a video region; and updating one or more tables based on an update rule.

[0006] In another exemplary aspect, another video decoding method is disclosed. This method includes checking a set of tables, each table including one or more motion candidates, each motion candidate being associated with motion information of the motion candidate, processing motion information of this video block based on the one or more tables, and updating the one or more tables based on the video block generated by the processing.

[0007] In yet another exemplary aspect, another video decoding method is disclosed. This method includes checking a set of tables, each table including one or more motion candidates, each motion candidate being associated with motion information of the motion candidate, selecting one or more tables based on the position of the video block within the picture, processing motion information of this video block based on the selected one or more tables, and updating the selected tables based on the video block generated by the processing.

[0008] In yet another exemplary aspect, another video decoding method is disclosed. This method includes checking a set of tables, each table including one or more motion candidates, each motion candidate being associated with motion information of the motion candidate, selecting one or more tables based on the distance between the video block and one of the motion candidates within the one or more tables, processing motion information of the picture block based on the selected one or more tables, and updating the selected tables based on the video block generated by the processing.

[0009] In yet another exemplary aspect, a video encoder device implementing the video encoding method described herein is disclosed.

[0010] In yet another representative aspect, the various techniques described herein may be implemented as a computer program product stored on a non-transitory computer-readable medium. This computer program product includes program code for executing the methods described herein.

[0011] In yet another representative embodiment, the video decoder device may implement the methods as described herein.

[0012] Details of one or more implementations are set forth in the accompanying appendices, drawings, and the following description. Other features will be apparent from the description and drawings, and from the claims.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

[0014] To improve the compression ratio of videos, researchers are constantly seeking new techniques for coding videos.

[0015] 1. Introduction

[0016] This specification relates to video coding techniques. Specifically, it relates to the coding of motion information in video coding (e.g., merge mode, AMVP mode). It may be applied to existing video coding standards such as HEVC, or it may be applied to finalize the Versatile Video Coding (VVC) standard. The present invention is also applicable to future video coding standards or video codecs.

[0017] Brief Description

[0018] Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and both organizations jointly created H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. The video coding standard, H.262, is based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. An example of a typical HEVC encoder framework is shown in FIG. 1.

[0019] 2.1 Partition Structure

[0020] 2.1.1 Partition Tree Structure in H.264 / AVC

[0021] The core of the coding layer in the previous standard was a macroblock that included luminance samples of a 16×16 block and, in the case of normal 4:2:0 color sampling, chroma samples of two corresponding 8×8 blocks.

[0022] Intra-coded blocks use spatial prediction to utilize the spatial correlation between pixels. Two partitions are defined, which are 16×16 and 4×4.

[0023] Inter-coded blocks use temporal prediction instead of spatial prediction by estimating the motion between pictures. The motion can be estimated independently for either a 16×16 macroblock or its sub-macroblock partitions: 16×8, 8×16, 8×8, 8×4, 4×8, 4×4 (see Figure 2). Only one motion vector (MV) is permitted per sub-macroblock partition.

[0024] 2.1.2 Partition Tree Structure in HEVC

[0025] In HEVC, a CTU is divided into CUs using a quadtree structure called a coding tree so as to adapt to various local features. The decision of whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU can be further divided into one, two, or four PUs depending on the PU partition type. Inside one PU, the same prediction process is applied and the relevant information is sent to the decoder in PU units. After obtaining a residual block by applying the prediction process based on the PU partition type, the CU can be divided into transform units (TUs) based on another quadtree structure similar to the coding tree for the CU. One of the important features of the HEVC structure is having multiple partition concepts including CUs, PUs, and TUs.

[0026] Next, focus is placed on various features related to hybrid video coding using HEVC.

[0027] 1) Coding tree unit and coding tree block (CTB) structure: A similar structure in HEVC is the coding tree unit (CTU), which has a size selected by the encoder and may be larger than a conventional macroblock. A CTU consists of a luminance CTB, corresponding chroma CTBs, and syntax elements. The size L×L of the luminance CTB can be selected as L = 16, 32, or 64 samples, and larger sizes generally enable better compression. HEVC then supports dividing the CTB into smaller blocks using a tree structure and signaling similar to a quadtree.

[0028] 2) Coding unit (CU) and coding block (CB): The quadtree syntax of the CTU specifies the size and position of its luminance and chroma CBs. The root of the quadtree is associated with the CTU. Thus, the size of the luminance CTB is the largest size supported for the luminance CB. Dividing the CTU into a luminance CB and chroma CBs is signaled together. One luminance CB and usually two chroma CBs, along with the associated syntax, form one coding unit (CU). A CTB may contain only one CU or may be divided to form multiple CUs, and each CU has a division into prediction units (PUs) associated with it and a tree of one transform unit (TU).

[0029] 3) Prediction Unit and Prediction Block (PB): The decision of whether to code a picture area using inter-picture or intra-picture prediction is made at the CU level. The partitioning structure of the PU has its root at the CU level. Based on the determination of the basic prediction type, next, the luma and chroma CB sizes can be further partitioned and predicted from the luma and chroma prediction blocks (PBs). HEVC supports variable PB sizes from 64×64 samples to 4×4 samples. Figure 3 shows an example of the permitted PBs for an MxM CU.

[0030] 4) TU and Transform Block: The prediction residual is coded using block transform. The TU tree structure has its root at the CU level. This luma CB residual may be the same as the luma transform block (TB), or may be further partitioned into smaller luma TBs. The same applies to the chroma TB. For square TB sizes of 4×4, 8×8, 16×16, and 32×32, integer basis functions similar to the integer basis functions of the discrete cosine transform (DCT) are defined. For the 4×4 transform of the luma intra-picture prediction residual, an integer transform derived from the form of the discrete sine transform (DST) is alternatively specified.

[0031] Figure 4 shows an example of subdividing the CTB into CBs [and transform blocks (TBs)]. The solid lines indicate the CB boundaries, and the dotted lines indicate the TB boundaries. (a) CTB and its partitioning (b) The corresponding quadtree.

[0032] 2.1.2.1 Tree Structure Partitioning into Transform Blocks and Units

[0033] In the case of residual coding, the CB can be recursively divided into transform blocks (TBs). This division is signaled by a residual quadtree. As shown in Figure 4, only the division of square CBs and TBs is specified so that one block can be recursively divided into quadrants. For a given luminance CB of size M×M, a flag signals whether it is divided into four blocks of size M / 2×M / 2. If further division is possible, each quadrant is assigned a flag indicating whether it is divided into four quadrants, as signaled by the maximum depth of the residual quadtree shown in the SPS. The resulting leaf node blocks of the residual quadtree are transform blocks that are further processed by transform coding. The encoder indicates the maximum luminance TB size and the minimum luminance TB size that it will use. If the CB size is larger than the maximum TB size, the division is implicitly performed. If the division results in a luminance TB size smaller than the indicated minimum value, it is implicitly not performed. Except when the luminance TB size is 4×4, the chroma TB size is half of the luminance TB size in each dimension, and in this case, one 4×4 chroma TB is used for the area covered by four 4×4 luminance TBs. For intra-picture prediction CUs, the decoded samples of the nearest neighboring TB (inside or outside the CB) are used as reference data for intra-picture prediction.

[0034] In contrast to conventional standards, the HEVC design allows one TB to span multiple PBs for inter-picture prediction CUs, maximizing the potential coding efficiency benefits of the quadtree-structured TB division.

[0035] 2.1.2.2 Parent and Child Nodes

[0036] The CTB is divided based on a quadtree structure, and its nodes are coding units. Multiple nodes in the quadtree structure include leaf nodes and non-leaf nodes. A leaf node has no child nodes within the tree structure (i.e., a leaf node is not further divided). Non-leaf nodes include the root node of the tree structure. The root node corresponds to the first video block (e.g., CTB) of the video data. For each non-root node among the multiple nodes, each non-root node corresponds to a video block that is a sub-block of the video block corresponding to the parent node of the non-root node in the tree structure. Each non-leaf node among the multiple non-leaf nodes has one or more child nodes in the tree structure.

[0037] 2.1.3 Quadtree + Binary Tree Block Structure with Larger CTUs in JEM

[0038] In order to explore future video coding technologies beyond HEVC, in 2015, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET). Since then, many new methods have been adopted by JVET and incorporated into the reference software called the Joint Exploration Model (JEM).

[0039] 2.1.3.1 QTBT Block Partition Structure

[0040] Unlike HEVC, the QTBT structure eliminates the concept of multiple partition types. That is, it removes the separation of the concepts of CU, PU, and TU, and improves the flexibility of the shape of the CU partition. In the QTBT block structure, the CU can have either a square or a rectangular shape. As shown in Figure 5, first, the coding tree unit (CTU) is divided in a quadtree structure. The leaf nodes of the quadtree are further divided by a binary tree structure. There are two types of binary tree divisions: symmetric horizontal division and symmetric vertical division. The leaf nodes of the binary tree are called coding units (CUs), and this segmentation is used for prediction and transformation processing without further division. This means that in the QTBT-coded block structure, the CU, PU, and TU have the same block size. In JEM, a CU may consist of coded blocks (CBs) of different color components. For example, in the case of P and B slices in 4:2:0 chroma format, one CU contains one luma CB and two chroma CBs, and sometimes consists of a single-component CB. For example, one CU contains only one luma CB, or in the case of I slices, only two chroma CBs.

[0041] Specify the following parameters for the QTBT partitioning scheme. - CTU size: The size of the root node of one quadtree, the same concept as in HEVC - MinQTSize: The minimum allowable leaf node size of the quadtree - MaxBTSize: The maximum allowable binary tree root node size - Ma×BTDepth: The maximum allowable depth of the binary tree - MinBTSize: The minimum allowable leaf node size of the binary tree

[0042] In an example of the split structure of QTBT, set the CTU size as a 128×128 luma sample having chroma samples of two corresponding 64×64 blocks, set MinQTSize as 16×16, set Ma×BTSize as 64×64, set MinBTSize (for both width and height) as 4×4, and set Ma×BTDepth as 4. The quadtree split is first applied to the CTU to generate leaf nodes of the quadtree. The size of the leaf nodes of the quadtree can range from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the leaf quadtree node is 128×128, since its size exceeds Ma×BTSize (i.e., 64×64), it is not further split by the binary tree. Otherwise, the leaf quadtree node may be further split by the binary tree. Therefore, this leaf node of the quadtree is also the root node of the binary tree, and the depth of that binary tree is 0. When the depth of the binary tree reaches Ma×BTDepth (i.e., 4), no further splitting is considered. If the width of the binary tree node is equal to MinBTSize (i.e., 4), no further horizontal splitting is considered. Similarly, if the height of the binary tree node is equal to MinBTSize, no further vertical splitting is considered. The leaf nodes of the binary tree are further processed by prediction and transformation processing without further splitting. In JEM, the maximum CTU size is 256×256 luma samples.

[0043] Figure 5 (left) shows an example of block splitting using QTBT, and Figure 5 (right) shows the corresponding tree representation. The solid line represents the quadtree split, and the dotted line represents the binary tree split. In each split (i.e., non-leaf) node of the binary tree, one flag is signaled to indicate which split type (i.e., horizontal or vertical) is used. Here, 0 represents a horizontal split, and 1 represents a vertical split. In the case of the quadtree split, since the quadtree split always splits the block horizontally and vertically to generate four sub-blocks of equal size, there is no need to indicate the split type.

[0044] Furthermore, the QTBT scheme supports the ability to have separate QTBT structures for luminance and chroma. Currently, for P and B slices, the luminance and chroma CTBs in one CTU share the same QTBT structure. However, for I slices, the luminance CTB is divided into CUs by the QTBT structure, and the chroma CTB is divided into chroma CUs by a different QTBT structure. This means that one CU in one I slice consists of one coded block of one luminance component or one coded block of two chroma components, while one CU in one P or B slice consists of coded blocks of all three color components.

[0045] In HEVC, the inter prediction for small blocks is restricted to reduce the memory access of motion compensation. As a result, bi-prediction is not supported for 4×8 and 8×4 blocks, and inter prediction is not supported for 4×4 blocks. In the QTBT of JEM, these restrictions are removed.

[0046] 2.1.4 Ternary Trees in VVC

[0047] In some embodiments, tree types other than quadtree and binary tree are supported. In this implementation, as shown in FIGS. 6(d) and (e), two or more ternary tree (TT) partitions are introduced, namely, horizontal and vertical center-side ternary trees.

[0048] FIG. 6 shows (a) quadtree partitioning, (b) vertical binary tree partitioning, (c) horizontal binary tree partitioning, (d) vertical center-side ternary tree partitioning, (e) horizontal center-side ternary tree partitioning.

[0049] In some implementations, there are two levels of trees, namely, a region tree (quad-tree) and a prediction tree (binary tree or ternary tree). A CTU is first divided by a region tree (RT). The RT leaves may be further divided by a prediction tree (PT). The PT leaves may also be further divided by the PT until the maximum PT depth is reached. The PT leaf is the basic coding unit, which is also called a CU for convenience. A CU cannot be further divided. Both prediction and transformation are applied to the CU in the same way as in JEM. The entire partition structure is called a "multi-type tree".

[0050] 2.1.5 Partition Structure

[0051] The tree structure used in this response is called a multi-tree type (MTT), which is a generalization of QTBT. In QTBT, as shown in FIG. 5, first, a coding tree unit (CTU) is divided into a quad-tree structure. The leaf nodes of the quad-tree are further divided by a binary tree structure.

[0052] The basic structure of MTT consists of two types of tree nodes. As shown in FIG. 7, the region tree (RT) and the prediction tree (PT) support nine types of partitions.

[0053] FIG. 7 shows (a) quadtree partitioning (b) vertical binary tree partitioning (c) horizontal binary tree partitioning (d) vertical ternary tree partitioning (e) horizontal ternary tree partitioning (f) horizontal upper asymmetric binary tree partitioning (g) horizontal lower asymmetric binary tree partitioning (h) vertical left asymmetric binary tree partitioning (i) vertical right asymmetric binary tree partitioning.

[0054] One region tree can recursively divide one CTU into square blocks so that each leaf node of the region tree is a 4×4-sized block. At each node in the region tree, the prediction tree can be a binary tree (BT), a ternary tree (TT), and an asymmetric binary tree (ABT). In the PT split, it is prohibited to have a quadtree partition on the branches of the prediction tree. Similar to JEM, the luminance tree and the chroma tree are divided into I slices. The signaling methods for RT and PT are shown in FIG. 8.

[0055] 2.2 Inter Prediction in HEVC / H.265

[0056] Each inter prediction PU has motion parameters for one or two reference picture lists. The motion parameters include motion vectors and reference picture indices. The use of one of the two reference picture lists may be signaled using inter_pred_idc. The motion vector may be explicitly coded as a delta with respect to the predictor, and such a coding mode is called the AMVP mode.

[0057] When one CU is coded in skip mode, one PU is associated with this CU, there are no significant residual coefficients, and there are no coding motion vector deltas or reference picture indices. The merge mode is specified, and thereby the motion parameters for the current PU are obtained from neighboring PUs including spatial and temporal candidates. The merge mode can be applied not only for skip mode but also for any inter-predicted PU. As an alternative to the merge mode, there is an explicit transmission of the motion parameters, and for each PU, the motion vectors, which are the reference picture indices corresponding to each reference picture list and the use of the reference picture list, are explicitly signaled.

[0058] If signaling indicates using one of the two reference picture lists, a PU is generated from a block of one sample. This is called "single prediction". Single prediction is available for both P slices and B slices.

[0059] If signaling indicates using both reference picture lists, a PU is generated from a block of two samples. This is called "dual prediction". Dual prediction is available only for B slices.

[0060] Hereinafter, the inter prediction modes defined in HEVC will be described in detail. First, the merge mode will be described.

[0061] 2.2.1 Merge Mode

[0062] 2.2.1.1 Derivation of Merge Mode Candidates

[0063] When predicting a PU using the merge mode, an index pointing to an entry in the merge candidate list is parsed from the bitstream and used to retrieve the motion information. The construction of this list is defined in the HEVC standard and can be summarized based on the following sequence of steps. · Step 1: Initial candidate derivation o Step 1.1: Spatial candidate derivation o Step 1.2: Redundancy check of spatial candidates o Step 1.3: Temporal candidate derivation · Step 2: Insertion of additional candidates o Step 2.1: Creation of dual prediction candidates o Step 2.2: Insertion of motion zero candidates

[0064] These steps are also schematically shown in FIG. 9. For spatial merge candidate derivation, up to four merge candidates are selected from candidates at five different positions. For temporal merge candidate derivation, up to one merge candidate is selected from two candidates. Since a certain number of candidates are assumed for each PU on the decoder side, additional candidates are generated if the number of candidates does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is fixed, the index of the best merge candidate is encoded using a shortened unary binary (TU). When the size of the CU is equal to 8, all PUs of the current CU share the same one merge candidate list as the merge candidate list of the 2N×2N prediction unit.

[0065] Hereinafter, the operations associated with the above-described steps will be described in detail.

[0066] 2.2.1.2 Spatial Candidate Derivation

[0067] In deriving spatial merge candidates, up to four merge candidates are selected from the candidates at the positions shown in FIG. 10. The derivation order is A1, B1, B0, A0, B2. Position B2 is considered only if any of the PUs at positions A1, B1, B0, A0 are not available (e.g., because they belong to another slice or tile) or if they are intra-coded. After adding the candidate at position A1, when adding the remaining candidates, a redundancy check is performed, which can reliably exclude candidates with the same motion information from the list and improve coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only the pairs linked by arrows in FIG. 11 are considered, and the candidate is added to the list only if the corresponding candidates used for the redundancy check do not have the same motion information. Another source of overlapping motion information is the "second PU" associated with a partition different from 2N×2N. As an example, FIG. 12 shows the second PU for the cases of N×2N and 2N×N respectively. When dividing the current PU into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would make the dual prediction units have the same motion information, and it would be redundant to have only one PU per coding unit. Similarly, when dividing the current PU into 2N×N, position B1 is not considered.

[0068] 2.2.1.3 Temporal Candidate Derivation

[0069] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, based on the collocated PU belonging to the picture having the minimum POC difference from the current picture in a given reference picture list, a scaled motion vector is derived. In the slice header, the reference picture list used for the derivation of the collocated PU is clearly signaled. As shown by the dotted line in FIG. 13, the scaled motion vector of the temporal merge candidate is obtained. This is scaled from the motion vector of the collocated PU using the POC distances tb and td. tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated PU and the collocated picture. The reference picture index of the temporal merge candidate is set equal to zero. The actual implementation of this scaling process is described in the HEVC specification. In the case of a B slice, two motion vectors, i.e., one for reference picture list 0 and the other for reference picture list 1, are obtained and combined to form a bi-prediction merge candidate. Explanation of the scaling of the motion vector for the temporal merge candidate.

[0070] In the collocated PU (Y) belonging to the reference frame, as shown in FIG. 14, the position of the temporal candidate is selected between candidate C0 and candidate C1. If the PU at position C0 is not available, is intra-coded, or is outside the current CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate.

[0071] 2.2.1.4 Additional Candidate Insertion

[0072] In addition to the spatio-temporal merge candidates, there are two additional types of merge candidates, namely, combined bi-prediction merge candidates and zero merge candidates. The combined bi-prediction merge candidates are generated using the spatio-temporal merge candidates. The combined bi-prediction merge candidates are used only for B slices. The combined bi-prediction candidates are generated by combining the first reference picture list motion parameter of the first candidate and the second reference picture list motion parameter of another candidate. If these two tuples provide different motion hypotheses, these tuples form a new bi-prediction candidate. As an example, FIG. 15 shows a case where combined bi-prediction merge candidates to be added to the final list (right side) are generated using two candidates having mvL0, refIdxL0 or mvL1, refIdxL1 in the original list (left side). There are various rules regarding the combinations considered for generating these additional merge candidates.

[0073] Motion zero candidates are inserted and the remaining entries in the merge candidate list are filled to hit the MaxNumMergeCand capacity. These candidates have a spatial displacement of zero and a reference picture index that starts from zero and increases each time a new zero motion candidate is added to the list. The number of reference frames used by these candidates is 1 for uni-directional prediction and 2 for bi-directional prediction, respectively. Finally, no redundancy check is performed on these candidates.

[0074] 2.2.1.5 Motion Estimation Region for Parallel Processing

[0075] To speed up the encoding process, motion estimation can be performed in parallel, thereby simultaneously deriving the motion vectors of all prediction units within a given region. Since one prediction unit cannot derive motion parameters from adjacent PUs until the associated motion estimation is completed, the derivation of merge candidates from the spatial neighborhood may interfere with parallel processing. To mitigate the trade-off between coding efficiency and processing latency, HEVC defines a motion estimation region (MER), the size of which is signaled in the picture parameter set using the "log2_parallel_merge_level_minus2" syntax element. When defining one MER, merge candidates in the same region are marked as unavailable and thus not considered in list construction.

[0076] 7.3.2.3 Picture Parameter Set RBSP Syntax

[0077] 7.3.2.3.1 General Picture Parameter Set RBSP Syntax

[0078]

Table 1

[0079] log2_parallel_merge_level_minus2 plus 2 specifies the value of the variable Log2ParMrgLevel used in the derivation process of the luminance motion vectors in the merge mode specified in Section 8.5.3.2.2.2 and the derivation process of the spatial merge candidates specified in Section 8.5.3.2.3. The value of log2_parallel_merge_level_minus2 shall be in the range including 0 to CtbLog2SizeY_2. The variable Log2ParMrgLevel is derived as follows. Log2ParMrgLevel = log2_parallel_merge_level_minus2 + 2 (7-37) Note 3: The value of Log2ParMrgLevel indicates the built-in ability to derive merge candidate lists in parallel. For example, when Log2ParMrgLevel is equal to 6, the merge candidate lists for all prediction units (PUs) and coding units (CUs) included in a 64×64 block can be derived in parallel.

[0080] 2.2.2 Motion Vector Prediction in AMVP Mode

[0081] Motion vector prediction utilizes the spatio-temporal correlation between the motion vector and neighboring PUs and uses this for explicit transmission of motion parameters. First, the availability of the left and upper temporally neighboring PU positions is checked, redundant candidates are removed, and the zero vector is added to make the length of the candidate list constant, thereby constructing a motion vector candidate list. Next, the encoder can select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to the signaling of the merge index, the index of the best motion vector candidate is encoded using a shortened unary term. The maximum value to be encoded in this case is 2 (e.g., Figures 2 to 8). The following section will explain the details of the derivation process of motion vector prediction candidates.

[0082] 2.2.2.1 Derivation of Motion Vector Prediction Candidates

[0083] Figure 16 summarizes the derivation process of motion vector prediction candidates.

[0084] In motion vector prediction, two types of motion vector candidates, namely spatial motion vector candidates and temporal motion vector candidates, are considered. To derive spatial motion vector candidates, as shown in Figure 11, based on the motion vectors of each PU at five different positions, ultimately two motion vector candidates are derived.

[0085] To derive a temporal motion vector candidate, one motion vector candidate is selected from two candidates derived based on two different positions arranged at the same location. After creating the first spatio-temporal candidate list, duplicate motion vector candidates in the list are removed. If the number of candidates is more than two, motion vector candidates with a reference picture index greater than 1 in the associated reference picture list are removed from the list. If the number of spatio-temporal motion vector candidates is less than two, additional zero motion vector candidates are added to the list.

[0086] 2.2.2.2 Spatial Motion Vector Candidates

[0087] In the derivation of spatial motion vector candidates, among the five potential candidates derived from the PUs at the positions as shown in FIG. 11, those at the same position as motion merge are considered as up to two candidates. The derivation order for the left side of the current PU is defined as A0, A1, scaled A0, scaled A1. The derivation order for the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases where it can be used as a motion vector candidate, namely two cases where spatial scaling is not required and two cases where spatial scaling is used. Summarizing the four different cases, it is as follows. · Without Spatial Scaling - (1) The same reference picture list and the same reference picture index (same POC) - (2) Different reference picture lists but the same reference picture (same POC) · Spatial Scaling - (3) The same reference picture list but different reference pictures (different POCs) - (4) Different reference picture lists and different reference pictures (different POCs)

[0088] First, check the case of non-spatial scaling, and then perform spatial scaling. Regardless of the reference picture list, if the POC is different between the reference picture of the neighboring PU and the reference picture of the current PU, consider spatial scaling. If all PUs of the left candidate are not available or are intra-coded, the scaling of the upper motion vector is useful for the parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling is not permitted for the upper motion vector.

[0089] In the spatial scaling process, as shown in Figure 17, similar to the temporal scaling, scale the motion vectors of the neighboring PUs. The main difference is that the reference picture list and index of the current PU are given as inputs, and the actual scaling process is the same as the temporal scaling.

[0090] 2.2.2.3 Temporal Motion Vector Candidates

[0091] Except for deriving the reference picture index, all the processes for deriving the temporal merge candidates are the same as those for deriving the spatial motion vector candidates (see Figure 6). The reference picture index is signaled to the decoder.

[0092] 2.2.2.4 Signaling of AMVP Information

[0093] In the case of the AMVP mode, in the bitstream, four parts, namely, the prediction direction, the reference index, the MVD, and the mv predictor candidate index, can be signaled.

[0094] Syntax Table:

[0095] [Table 2]

[0096] 7.3.8.9 Motion Vector Difference Syntax

[0097]

Table 3

[0098] 2.3 New Inter - prediction Method in Joint Exploration Model (JEM)

[0099] 2.3.1 Motion Vector Prediction Based on Sub - CUs

[0100] In JEM with QTBT, each CU can have at most one set of motion parameters for each prediction direction. In the encoder, by splitting the large CU into sub - CUs and deriving the motion information of all sub - CUs of the large CU, two sub - CU level motion vector prediction methods are considered. By the alternative temporal motion vector prediction (ATMVP) method, each CU can fetch a set of multiple motion information from multiple blocks smaller than the current CU in the arranged reference pictures. In the spatio - temporal motion vector prediction (STMVP) method, the motion vector of the sub - CU is recursively derived using the temporal motion vector predictor and the spatial neighboring motion vectors.

[0101] To maintain a more accurate motion field for sub - CU motion prediction, the motion compression of the reference frames is currently disabled.

[0102] 2.3.1.1 Alternative Temporal Motion Vector Prediction

[0103] In the alternative temporal motion vector prediction (ATMVP), the motion vector temporal motion vector prediction (TMVP) method is modified by taking out multiple sets of motion information (including motion vectors and reference indices) from blocks smaller than the current CU. As shown in Figure 18, the sub - CU is a square of N×N blocks (by default, N is set to 4).

[0104] ATMVP predicts the motion vectors of sub-CUs within a CU in two steps. The first step is to identify the corresponding block in reference picture 1 with a so-called temporal vector. This reference picture is called the motion source picture. The second step is to divide the current CU into sub-CUs as shown in FIG. 18, and obtain the motion vectors and reference indices of each sub-CU from the blocks corresponding to each sub-CU.

[0105] In the first step, the reference picture and the corresponding block are determined by the motion information of spatially neighboring blocks of the current CU. To avoid repeated scanning of neighboring blocks, the first merge candidate in the merge candidate list of the current CU is used. The first available motion vector and its related reference index are set to the temporal vector and the index of the motion source picture. Thus, in ATMVP, compared with TMVP, the corresponding block can be identified more accurately, and the corresponding block (sometimes called the arranged block) is always at the lower right or center position relative to the current CU. In one example, if the first merge candidate is from a neighboring block on the left (i.e., A1 in FIG. 19), the related MV and reference picture are used to identify the source block and the source picture.

[0106] FIG. 19 shows an example of the identification of the source block and the source picture.

[0107] In the second step, by adding the temporal vector to the coordinates of the current CU, the corresponding block of the sub-CU is identified by the temporal vector in the motion source picture. For each sub-CU, the motion information of the sub-CU is derived using the motion information (the smallest motion grid covering the central sample) of its corresponding block. After identifying the motion information of the corresponding N×N block, it is converted into the motion vector and reference index of the current sub-CU, similar to TMVP in HEVC, and motion scaling and other procedures are applied. For example, the decoder checks whether the low-delay condition (the POC of all reference pictures of the current picture is smaller than the POC of the current picture) is satisfied, and in some cases, uses the motion vector MVx (the motion vector corresponding to reference picture list x) to predict the motion vector MVy of each sub-CU (where X is equal to 0 or 1, and Y is equal to 1 - X).

[0108] 2.3.1.2 Spatial-Temporal Motion Vector Prediction

[0109] In this method, the motion vectors of the sub-CUs are derived recursively along the raster scan order. This concept is shown in Figure 20. Consider an 8×8 CU containing four 4×4 sub-CUs, A, B, C, and D. The 4×4 blocks in the neighborhood of the current frame are labeled a, b, c, and d.

[0110] The derivation of the motion of Sub - CU A starts by identifying its two spatial neighborhoods. The first neighborhood is the N×N block above Sub - CU A (block c). If this block c is not available or is intra - coded, then check other N×N blocks above Sub - CU A (starting from block c and moving from left to right). The second neighborhood is the block to the left of Sub - CU A (block b). If block b is not available or is intra - coded, then check other blocks to the left of Sub - CU A (starting from block b and moving from top to bottom). The motion information obtained from the neighboring blocks of each list is scaled to the first reference frame of the given list. Next, a temporal motion vector predictor (TMVP) for sub - block A is derived following a procedure similar to that specified for TMVP derivation in HEVC. Fetch the motion information of the arranged blocks at location D and scale it accordingly. Finally, after searching and scaling the motion information, all available motion vectors (up to 3) for each reference list are averaged separately. This averaged motion vector is taken as the motion vector of the current Sub - CU.

[0111] Figure 20 shows an example of one CU having four sub - blocks (A - D) and their neighboring blocks.

[0112] 2.3.1.3 Sub - CU Motion Prediction Mode Signaling

[0113] The Sub - CU mode is made valid as an additional merge candidate and no additional syntax elements are required to signal the mode. Add two additional merge candidates to the merge candidate list of each CU, as represented by the ATMVP mode and the STMVP mode. If the sequence parameter set indicates that ATMVP and STMVP are valid, use up to 7 merge candidates. The encoding logic for the additional merge candidates is the same as that for the merge candidates in HM, i.e., for each CU in a P or B slice, more than two RD checks are required for the two additional merge candidates.

[0114] In JEM, all bins of the merge index are context-coded by CABAC. On the other hand, in HEVC, only the first bin is context-coded and the remaining bins are context-bypass-coded.

[0115] 2.3.2 Adaptive Motion Vector Difference Resolution

[0116] In HEVC, when use_integer_mv_flag is 0 in the slice header, the motion vector difference (MVD) (the difference between the motion vector and the predicted motion vector of the PU) is signaled in units of quarter-luminance samples. In JEM, local adaptive motion vector resolution (LAMVR) is introduced. In JEM, the MVD can be coded in units of 1 / 4 luminance samples, integer luminance samples, or four luminance samples. The MVD resolution is controlled at the coding unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU having at least one non-zero MVD module.

[0117] For a CU having at least one non-zero MVD module, a first flag is signaled to indicate whether 1 / 4 luminance sample MV accuracy is used in the CU. If the first flag (equal to 1) indicates that 1 / 4 luminance sample MV accuracy is not used, another flag is signaled to indicate whether integer luminance sample MV accuracy or 4 luminance sample MV accuracy is used.

[0118] If the first MVD resolution flag of the CU is zero, or not coded for the CU (i.e., all MVDs in the CU are zero), 1 / 4 luminance sample MV resolution is used for the CU. When the CU uses integer luminance sample MV accuracy or 4 luminance sample MV accuracy, the MVP in the AMVP candidate list of the CU is rounded to the corresponding accuracy.

[0119] In the encoder, the RD check at the CU level is used to determine which MVD resolution to use for the CU. That is, the RD check at the CU level is performed three times for each MVD resolution. In JEM, the following encoding method is applied to increase the encoder speed.

[0120] During the RD check of a CU with a normal 1 / 4 luminance sample MVD resolution, the motion information (integer luminance sample accuracy) of the current CU is stored. During the RD check of the same CU with integer luminance samples and 4 luminance sample MVD resolutions, the stored motion information (after rounding) is used as the starting point for further small-range motion vector refinement, so that the time-consuming motion estimation process does not repeat three times.

[0121] The RD check of a CU with a 4 luminance sample MVD resolution is called conditionally. In the case of a CU, if the RD cost of the integer luminance sample MVD resolution is much larger than that of the 1 / 4 luminance sample MVD resolution, the RD check of the 4 luminance sample MVD resolution for the CU is omitted.

[0122] 2.3.3 Pattern Matching Motion Vector Derivation

[0123] The Pattern Matching Motion Vector Derivation (PMMVD) mode is a special merge mode based on frame rate up-conversion (FRUC) technology. In this mode, the motion information of the block is not signaled and is derived on the decoder side.

[0124] The FRUC flag is signaled to the CU if the merge flag is true. If the FRUC flag is false, the merge index can be signaled and the normal merge mode is used. If the FRUC flag is true, an additional FRUC mode flag is signaled to indicate which method (bilateral matching or template matching) is used to derive the motion information of the block.

[0125] On the encoder side, the decision of whether to use the FRUC merge mode for a CU is based on the selection of the RD cost, in the same way as for normal merge candidates. That is, the RD cost selection is used to check both of the two matching modes (e.g., bilateral matching and template matching) for one CU. What leads to the minimum cost is compared with other CU modes. If the FRUC matching mode is the most efficient, the FRUC flag is set to true for the CU, and the relevant matching mode is used.

[0126] The motion derivation process in the FRUC merge mode has two steps. First, motion search at the CU level is performed, and then motion refinement at the sub-CU level is performed. At the CU level, an initial motion vector for the entire CU is derived based on bilateral matching or template matching. First, a list of MV candidates is generated, and the candidate that leads to the minimum matching cost is selected as the starting point for further CU level improvement. Then, a local search based on bilateral matching or template matching near the starting point is performed, and the MV result with the minimum matching cost is taken as the MV for the entire CU. Subsequently, with the derived CU motion vector as the starting point, the motion information at the sub-CU level is further refined.

[0127] For example, for the derivation of W×H CU motion information, the following derivation process is performed. In the first stage, an MV for the entire W×H CU is derived. In the second stage, the CU is further divided into M×M sub-CUs. The value of M is calculated as (16), where D is a predefined division depth, which is set to 3 by default in JEM. Then, the MV for each sub-CU is derived.

[0128]

Number

[0129] As shown in FIG. 21, this bilateral matching is used to derive the motion information of the current CU by finding the closest matching between two blocks along the motion trajectory of the current CU in two different reference pictures. Assuming a continuous motion trajectory, the motion vectors MV0 and MV1 pointing to the two reference blocks are proportional to the temporal distances, e.g., TD0 and TD1, between the current picture and the two reference pictures. As a special case, when the current picture is temporally between the two reference pictures and the temporal distances between the current picture and the two reference pictures are the same, the bilateral matching results in a mirror-based bidirectional MV.

[0130] As shown in FIG. 22, the motion information of the current CU is derived using template matching by finding the closest matching between a template (blocks in the vicinity of the upper side and / or left side of the current CU) in the current picture and a block (of the same size as the template) in the reference picture. In addition to the aforementioned FRUC merge mode, template matching is also applied to the AMVP mode. In JEM, similar to HEVC, AMVP has two candidates. By using the template matching method, new candidates are derived. If the candidate newly derived by template matching is different from the first existing AMVP candidate, it is inserted at the beginning of the AMVP candidate list, and then the list size (which means removing the second existing AMVP candidate) is set to 2. When applied to the AMVP mode, only CU-level search is applied.

[0131] 2.3.3.1 CU-Level MV Candidate Set

[0132] The CU-level MV candidate set consists of the following. (i) The original AMVP candidates when the current CU is in the AMVP mode (ii) All merge candidates, (iii) Multiple MVs in the interpolation MV field. (iv) The motion vector in the upper left neighborhood

[0133] When using bilateral matching, each valid MV of the merge candidates is used as an input to assume bilateral matching and generate MV pairs. For example, one valid MV of the merge candidates is (MVa, refa) in reference list A, and the reference picture refb of the bilateral MV forming the pair is found in another reference list B, and refa and refb are on different sides of the currently present picture in terms of time. If such refb is not available in reference list B, refb is determined as a reference different from refa, and the temporal distance from the current picture is the minimum value in list B. After determining refb, MVb is derived by scaling MVa based on the temporal distances between the current picture and refa, refb.

[0134] Four MVs from the interpolated MV field are also added to the CU-level candidate list. Specifically, the interpolated MVs at the positions of (0,0), (W / 2,0), (0,H / 2), and (W / 2,H / 2) of the current CU are added.

[0135] When applying FRUC in AMVP mode, the original AMVP candidates are also added to the CU-level MV candidate set.

[0136] At the CU level, 15 MVs for AMVP CUs and up to 13 MVs for merge CUs are added to the candidate list.

[0137] 2.3.3.2 Sub-CU Level MV Candidate Set

[0138] The MV candidate set at the sub-CU level consists of the following. (i) The MVs determined from the CU-level search, (ii) The MVs in the vicinity of the upper left and upper right, (iii) The scaled versions of the MVs juxtaposed from the reference pictures, (iv) Up to four ATMVP candidates, (v) Up to four STMVP candidates

[0139] The scaled MV from the reference picture is derived as follows. Traverse all the reference pictures in both lists. The MV at the array position of the sub-CU in the reference picture is scaled relative to the reference of the start CU level MV.

[0140] The candidates for ATMVP and STMVP are limited to the first four candidates.

[0141] At the sub-CU level, up to 17 MVs are added to the candidate list.

[0142] 2.3.3.3 Generation of Interpolation MV Field

[0143] Before coding a certain frame, an interpolation motion field is generated for the entire picture based on one-sided ME. And this motion field may be used later as an MV candidate at the CU level or sub-CU level.

[0144] First, the motion fields of each reference picture in both reference lists are traversed at the 4×4 block level. In each 4×4 block, for the motion related to the block passing through the 4×4 block of the current picture, if the interpolated motion has not been assigned yet, based on the temporal distances TD0 and TD1 (similar to the MV scaling of TMVP in HEVC), scale the motion of the reference block to the current picture (shown in Figure 23), and assign the scaled motion to the block of the current frame. If an MV scaled to a 4×4 block has not been assigned, the motion of the block is marked as unavailable in the interpolated motion field.

[0145] 2.3.3.4 Interpolation and Matching Cost

[0146] When one motion vector points to one fractional sample position, motion compensation interpolation is required. To reduce complexity, bilinear interpolation is used for both bilateral matching and template matching instead of normal 8-tap HEVC interpolation.

[0147] The calculation of the matching cost is slightly different in different steps. When selecting a candidate from the candidate set at the CU level, the matching cost is the sum of absolute differences (SAD) of bilateral matching or template matching. After determining the starting MV, the matching cost C of bilateral matching in the sub-CU level search is calculated as follows.

[0148]

Equation

[0149] Here, w is a weight coefficient empirically set to 4, and MV and MV s represent the current MV and the starting MV, respectively. SAD is still used as the matching cost of template matching in the sub-CU level search.

[0150] In the FRUC mode, the MV is derived by using only luminance samples. The derived motion is used for both luminance and chroma for MC inter prediction. After determining the MV, final MC is performed using an 8-tap interpolation filter for luminance and a 4-tap interpolation filter for chroma.

[0151] 2.3.3.5 Improvement of MV

[0152] MV refinement is MV search based on patterns having criteria of bilateral matching cost or template matching cost. In JEM, two search patterns, namely, unrestricted centered bias diamond search (UCBDS) and adaptive cross search for MV refinement at CU level and sub-CU level are supported respectively. For both CU and sub-CU level MV improvement, the MV is directly searched at the accuracy of 1 / 4 luminance sample MV, followed by improvement of 1 / 8 luminance sample MV. The search range of MV refinement for CU and sub-CU steps is set equal to 8 luminance samples.

[0153] 2.3.3.6 Selection of Prediction Direction in Template Matching FRUC Merge Mode

[0154] In bilateral matching merge mode, dual prediction is always applied. This is because the motion information of the CU is derived based on the closest matching between two blocks along the motion trajectory of the current CU in two different reference pictures. For template matching merge mode, there is no such limitation. In template matching merge mode, the encoder can choose from a single prediction from list0, a single prediction from list1, or dual prediction for the CU. The selection is made as follows based on the template matching cost. If costBi <= factor * min(cost0, cost1) Use dual prediction. Otherwise, if cost0 <= cost1, Use a single prediction from list0. Otherwise, Use a single prediction from list1.

[0155] Here, cost0 is the SAD of list0 template matching, cost1 is the SAD of list1 template matching, and costBi is the SAD of bi-prediction template matching. When the value of factor is 1.25, it means that the selection process is biased towards bi-prediction. This inter-prediction direction selection is only applied to the template matching process at the CU level.

[0156] 2.3.4 Decoder-side Motion Vector Improvement

[0157] In the bi-prediction operation, to predict one block area, a bi-prediction block composed of the motion vectors (MVs) of list0 and the MVs of list1 is combined respectively to form one prediction signal. In the decoder-side motion vector improvement (DMVR) method, the two motion vectors of bi-prediction are further improved by bilateral template matching processing. In order to obtain the improved MVs without transmitting additional motion information, bilateral template matching is applied in the decoder, and a distortion-based search is performed between the bilateral template and the reconstructed samples in the reference picture.

[0158] In DMVR, as shown in Fig. 23, a bilateral template is generated from the first MV0 of list0 and MV1 of list1 as the weighted combination (i.e., average) of the bi-prediction blocks. The template matching operation consists of calculating the cost metric between the generated template and the sample region (near the first prediction block) in the reference picture. For each of the two reference pictures, the MV with the minimum template cost is regarded as the updated MV of that list and replaces the original MV. In JEM, nine MV candidates are searched for each list. The nine MV candidates include the original MV and eight surrounding MVs that have one luminance sample offset from the original MV in either the horizontal or vertical direction or both directions. Finally, two new MVs, namely MV0’ and MV1’ as shown in Fig. 24, are used to generate the final bi-prediction result. The sum of absolute differences (SAD) is used as the cost metric.

[0159] DMVR is applied to the merge mode of bi-prediction between one MV from a past reference picture and one MV from a future reference picture without transmitting additional syntax elements. In JEM, if LIC, affine motion, FRUC, or sub-CU merge candidates are valid for a CU, DMVR is not applied.

[0160] 2.3.5 Examples of Merge / Skip Modes with Improvements to Bilateral Matching

[0161] First, the merge candidate list is constructed by inserting the motion vectors and reference indices of spatially neighboring blocks and temporally neighboring blocks into a candidate list with redundancy checking until the number of available candidates reaches the maximum candidate size of 19. The merge candidate list for the merge / skip mode is constructed by inserting the spatial candidates (Fig. 11), temporal candidates, affine candidates, advanced temporal MVPs (ATMVP) candidates, spatio-temporal MVPs (STMVP) candidates, and additional candidates used in HEVC (merge candidates and zero candidates) based on a predefined insertion order.

[0162] - Spatial candidates for blocks 1 to 4

[0163] - Extrapolated affine candidates for blocks 1 to 4

[0164] - ATMVP

[0165] - STMVP

[0166] - Virtual affine candidates

[0167] - Spatial candidates (block 5) (used only when the number of available candidates is less than 6)

[0168] - Extrapolated affine candidates (block 5)

[0169] - Temporal candidates (derived like HEVC)

[0170] - Extrapolate affine candidates after non - adjacent spatial candidates (blocks 6 to 49 shown in Figure 25)

[0171] - Combined candidates

[0172] - Zero candidates

[0173] Note that the IC flag is inherited from the merge candidates, except for STMVP and affine. Also, for the first four spatial candidates, those with dual prediction are inserted before those with single prediction.

[0174] In some implementations, it is possible to access blocks that are not connected to the current block. If a non - adjacent block is coded in a non - intra mode, the relevant motion information may be added as an additional merge candidate.

[0175] 3. Examples of problems to be solved by the embodiments disclosed in this specification

[0176] The current HEVC design can take advantage of the correlation of blocks in the vicinity of the current block (adjacent to the current block) to better code motion information. However, the neighboring blocks may correspond to different objects with different motion trajectories. In this case, the prediction from the neighboring blocks is not efficient.

[0177] Predicting from the motion information of non - adjacent blocks would incur the cost of storing all motion information (generally at the 4×4 level) in the cache, resulting in additional coding gain and significantly increasing the complexity of the hardware implementation.

[0178] 4. Some Examples

[0179] To overcome the drawbacks of existing implementations, in various embodiments, a lookup - table (LUT) - based motion vector prediction technique that uses one or more lookup tables storing at least one motion candidate for predicting the motion information of a block is implemented, and video coding with higher coding efficiency can be provided. Each LUT may include one or more motion candidates each associated with corresponding motion information. The motion information of the motion candidates may include prediction direction, reference index / picture, motion vector, LIC flag, affine flag, motion vector derivation (MVD) accuracy, and / or some or all of the MVD values. The motion information may further include block position information to indicate where the motion information comes from.

[0180] LUT-based motion vector prediction based on the disclosed technology can improve both existing and future video coding standards and is illustrated in the following examples for various implementations. Since the LUT enables encoding / decoding processing based on history data (e.g., already processed blocks), LUT-based motion vector prediction can also be referred to as the history-based motion vector prediction (HMVP) method. In the LUT-based motion vector prediction method, one or more tables with motion information from previously coded blocks are maintained during the encoding / decoding process. During the encoding / decoding of one block, the associated motion information in the LUT can be added to the motion candidate list, and after encoding / decoding one block, the LUT can be used. The following examples should be considered as examples for explaining general concepts. These examples should not be interpreted in a narrow sense. Furthermore, these examples can be combined in any way.

[0181] In some embodiments, to predict the motion information of one block, one or more lookup tables storing at least one motion candidate may be used. The embodiments can show a set of motion information stored in the lookup table using the motion candidates. In the case of conventional AMVP or merge mode, in the embodiments, AMVP or merge candidates may be used to store the motion information.

[0182] The following examples illustrate general concepts.

[0183] Examples of Lookup Tables

[0184] Example A1: Each lookup table may include one or more motion candidates where each candidate is associated with its motion information. a. The motion information of the motion candidate may include a prediction direction, a reference index / picture, a motion vector, a LIC flag, an affine flag, an MVD accuracy, and some or all of the MVD values. b. The motion information may further include block position information to indicate where the motion information comes from. c. One counter may be further assigned for each lookup table. i. The counter may be initialized to zero at the start of encoding / decoding of a picture / slice / LCU (CTU) row / tile. ii. In one example, the counter may be updated after encoding / decoding a CTU / CTB / CU / CB / PU / a certain region size (e.g., 8×8 or 16×16). iii. In one example, each time one candidate is added to the lookup table, the counter is incremented by one. iv. In one example, the counter should be less than or equal to the size of the table (the number of allowed motion candidates). v. Alternatively, this counter may be used to indicate how many motion candidates were attempted to be added to the lookup table (some of which may have been included in the lookup table but may later be removed from the table). In this case, the counter may be greater than the size of the table. d. The size of the table (the number of allowed motion candidates) and / or the number of tables may be fixed or adaptable. The size of the table may be the same for all tables or different for different tables. i. Alternatively, different sizes may be used for different lookup tables (e.g., 1 or 2). ii. In one example, the size of the table and / or the number of tables may be predefined. iii. In one example, the size of the table and / or the number of tables may be signaled in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), slice header, tile header, coding tree unit (CTU), coding tree block (CTB), coding unit (CU) or prediction unit (PU), a region including a plurality of CTU / CTB / CU / PU. iv. The size and / or number of the tables may further depend on the slice type, the temporal layer index of the picture, and the picture order count (POC) distance between one slice and the nearest intra-slice. e. There are N tables used for one coding thread, and N*P tables are required to code one slice, where P indicates the number of LCU rows or the number of tiles. i. Alternatively, only P tables may be required to code one slice, where P indicates the number of LCU rows, each LCU row uses only one lookup table, and when tiles are disabled, N may be greater than 1.

[0185] Selection of LUT

[0186] Example B1: When coding one block, some or all of the motion candidates from one lookup table can be checked in sequence. When checking one motion candidate during coding of one block, this motion candidate may be added to the motion candidate list (e.g., AMVP, merge candidate list). a. Alternatively, motion candidates from multiple lookup tables may be checked in sequence. b. The lookup table index may be signaled in a CTU, CTB, CU or PU, or a region including multiple CTUs / CTBs / CUs / PUs.

[0187] Example B2: The selection of the lookup table may depend on the position of the block. a. It may depend on the CTU address containing the block. Here, for the purpose of explaining the idea, for example, a dual lookup table (DLUT) is given. i. If the block is located in one of the first M CTUs in a CTU row, the first lookup table may be used for coding the block, and for blocks located in the remaining CTUs in the CTU row, the second lookup table may be used. ii. If the block is located in one of the first M CTUs within a CTU row, first check whether to code the block according to the motion candidates in the first look-up table. If there are not enough candidates in the first table, the second look-up table can be further used. On the other hand, for blocks located in the remaining CTUs of the CTU row, the second look-up table may be used. iii. Alternatively, for blocks located in the remaining CTUs of the CTU row, first check the motion candidates in the second look-up table for block coding. If there are not enough candidates in the second table, the first look-up table may be further used. b. It may depend on the distance between the position of the block in one or more look-up tables and the position associated with one motion candidate. iv. In one example, if one motion candidate is associated with a smaller distance to the block to be coded, it may be checked earlier compared to another motion candidate.

[0188] Usage of the look-up table

[0189] Example C1: The total number of motion candidates in the look-up table to be checked may be predefined. a. It may further depend on coding information, block size, block shape, etc. For example, in the case of the AMVP mode, only m motion candidates are checked, and in the case of the merge mode, n motion candidates may be checked (e.g., m = 2, n = 44). b. In one example, the total number of motion candidates to be checked may be signaled in a region including a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile header, a coding tree unit (CTU), a coding tree block (CTB), a coding unit (CU), or a prediction unit (PU), or a plurality of CTUs / CTBs / CUs / PUs.

[0190] Example C2: One or more motion candidates included in one look-up table may be directly inherited by one block. a. They may be used for merge mode coding. That is, the motion candidates may be checked in the merge candidate list derivation process. b. These may be used for affine merge mode coding. i. When the affine flag is 1, the motion candidates in the look-up table can be added as affine merge candidates. c. In the following cases, checking of the motion candidates in the look-up table may be enabled. i. After inserting the TMVP candidate, there is space in the merge candidate list. ii. After checking specific spatially neighboring blocks for spatial merge candidate derivation, there is space in the merge candidate list. iii. After all spatial merge candidates, there is space in the merge candidate list. iv. After the combined dual prediction merge candidates, there is space in the merge candidate list. v. Pruning may be applied before adding the motion candidates to the merge candidate list.

[0191] Example C3: The motion candidates included in the look-up table may be used as a prediction module for coding the motion information of the block. a. They may be used for AMVP mode coding. That is, the motion candidates may be checked in the AMVP candidate list derivation process. b. In the following cases, checking of the motion candidates in the look-up table may be enabled. i. After inserting the TMVP candidate, there is space in the AMVP candidate list. ii. The AMVP candidate list is selected from the spatial neighborhood, pruned, and then there is space immediately before inserting the TMVP candidate. iii. If there is no AMVP candidate from the upper adjacent block without scaling and / or no AMVP candidate from the left neighboring block without scaling. iv. Pruning may be applied before adding motion candidates to the AMVP candidate list. c. Check motion candidates having the same reference picture as the current reference picture. i. Alternatively, also check motion candidates having a reference picture different from the current reference picture (in the state of MV scaling). ii. Alternatively, first check all motion candidates having the same reference picture as the current reference picture, and then check motion candidates having a reference picture different from the current reference picture. iii. Alternatively, check the motion candidates after merging.

[0192] Example C4: The checking order of motion candidates in the lookup table is defined as follows (assuming that K (K>=1) motion candidates can be checked): a. The last K motion candidates in the lookup table are arranged, for example, in descending order of the entry index to the LUT. b. The first K%L candidates. L is the size of the lookup table when K>=L, and is arranged, for example, in descending order of the entry index to the LUT. c. When K>=L, all candidates (L candidates) in the lookup table. d. Alternatively, further based on the descending order of the motion candidate index. e. Alternatively, select K motion candidates based on candidate information, such as the distance of the position associated with the motion candidate and the current block. f. The checking order of different lookup tables is defined in the usage of the lookup tables in the following subsections. g. Once the merge / AMVP candidate list reaches the maximum allowable number of candidates, this checking process ends. h. Instead, when the number of added motion candidates reaches the maximum allowable number of motion candidates, it ends. i. One syntax element indicating the table size and the number of motion candidates to be checked (i.e., K = L) may be signaled in the SPS, PPS, slice header, and tile header.

[0193] Example C5: Whether to use a lookup table for motion information coding of one block may be signaled in the SPS, PPS, slice header, tile header, CTU, CTB, CU, or PU, or in a region including a plurality of CTUs / CTBs / CUs / PUs.

[0194] Example C6: Whether to apply prediction from a lookup table may further depend on coding information. If it is inferred that prediction is not to be applied to one block, additional signaling of prediction instructions is skipped. Alternatively, if it is inferred that prediction is not to be applied to one block, there is no need to access the motion candidates of the lookup table, and the check of the relevant motion candidates is omitted. a. Whether to apply prediction from a lookup table may depend on the block size / block shape. In one example, for smaller blocks, such as 4×4, 8×4, or 4×8 blocks, execution of prediction from a lookup table is not permitted. b. Whether to apply prediction from a lookup table may depend on whether the block is coded in AMVP mode or merge mode. In one example, in AMVP mode, prediction from a lookup table is not permitted. c. Whether to apply prediction from a lookup table may depend on whether the block is coded with affine motion or other types of motion (e.g., translational motion). In one example, in affine mode, prediction from a lookup table is not permitted.

[0195] Example C7: Motion information of blocks in different frames / slices / tiles may be predicted using motion candidates of the look-up table in previously coded frames / slices / tiles. a. In one example, only the look-up table associated with the reference picture of the current block may be utilized to code the current block. b. In one example, only the look-up table associated with pictures having the same slice type and / or the same quantization parameter of the current block may be utilized to code the current block.

[0196] Update of the look-up table

[0197] Example D1: One or more look-up tables may be updated after coding a block with motion information (i.e., IntraBC mode, inter-coding mode). a. In one example, the rule for selecting the look-up table may be reused to determine whether to update the look-up table. b. The look-up table may be updated based on coding information and / or the position of the block / LCU. c. When a block is coded with directly signaled motion information (e.g., AMVP mode), the motion information of the block may be added to the look-up table. i. Alternatively, when a block is coded with motion information directly inherited from a spatially neighboring block without any improvement (e.g., spatial merge candidate without improvement), the motion information of the block should not be added to the look-up table. ii. Alternatively, when a block is coded with motion information directly inherited from a spatially neighboring block with improvement (such as DMVR, FRUC, etc.), the motion information of the block should not be added to any look-up table. iii. Alternatively, if the block is coded with motion information directly inherited from motion candidates stored in a lookup table, the motion information of the block should not be added to any lookup table. d. Select M (M >= 1) representative positions within the block and update the lookup table using the motion information associated with these representative positions. i. In one example, the representative position is defined as one of the four corner positions (e.g., C0 to C3 in FIG. 26) within the block. ii. In one example, the representative position is defined as the center position (e.g., Ca_Cd in FIG. 26) within the block. iii. If sub-block prediction is not permitted for the block, M is set to 1. iv. If sub-block prediction is permitted for the block, M can be exclusively set to 1 or the total number of sub-blocks, or a unique value between 1 and the number of sub-blocks. v. Alternatively, if sub-block prediction is permitted for the block, M can be set to 1, and the selection of representative sub-blocks is based on... 1. Frequency of the motion information utilized 2. Whether it is a bi-predicted block 3. Based on the reference picture index / reference picture 4. Difference of the motion vector compared to other motion vectors (e.g., select the maximum MV difference) 5. Other coding information e. When selecting a set of M (M >= 1) representative positions to update the lookup table, after checking further conditions, these may be added to the lookup table as further motion candidates. i. The set of new motion information may be pruned for the existing motion candidates in the lookup table. ii. In one example, the set of new motion information should not be the same as any or part of the existing motion candidates in the lookup table. iii. Alternatively, in the case of a new set of motion information and the same reference picture from one of the existing motion candidates, the MV difference must be less than 1 / a plurality of thresholds. For example, the horizontal and / or vertical modules of the MV difference must be greater than a distance of 1 pixel. iv. Alternatively, if K > L, the new set of motion information can be pruned only by the last K candidates or the first K % L of the existing motion candidates, and the old motion candidates can be made active again. v. Alternatively, no pruning is applied. f. If a set of M motion information is used to update the look-up table, the corresponding counter should be incremented by M. g. Before coding the current block, for one set of motion information selected (in the above manner), set the counter of the look-up table to be updated to K for this set of motion information. After coding this block, add it as an additional motion candidate with an index equal to K % L (where L is the size of the look-up table). An example is shown in Figure 27. i. Alternatively, it is added as an additional motion candidate with an index equal to min(K + 1, L - 1). Alternatively, if K >= L, remove the first motion candidate (index = 0) from the look-up table and decrement the subsequent K candidate indices by 1. h. After coding one intra constraint block, the look-up table may be emptied. i. When adding an entry of motion information to the look-up table, more entries of motion information may be added to the table by deriving from the motion information. In this case, the counter associated with the look-up table can be incremented by a number greater than 1. i. In one example, the MV of the entry of motion information is scaled and put into the table. ii. In one example, the MV of the entry of motion information is added by (dx, dy) and put into the table. iii. In one example, calculate the average of the MVs of the entries of two or more motion information and put it in a table.

[0198] Example D2: When one block is located at the boundary of one picture / slice / tile, updating the lookup table is not always permitted.

[0199] Example D3: To code the current LCU row, the motion information of the above LCU row may be invalidated. a. In this case, at the start of a new slice / tile / LCU row, the number of available motion candidates may be reset to 0.

[0200] Example D4: At the start of coding a slice / tile using a new temporal layer index, the number of available motion candidates can be reset to 0.

[0201] Example D5: The lookup table may be updated continuously in the rows / slices of one slice / tile / LCU having the same temporal layer index. a. Alternatively, the lookup table may be updated only after coding / decoding each S (S >= 1) CTU / CTB / CU / CB, or after coding / decoding a specific region (for example, a size equal to 8×8 or 16×16). b. Alternatively, one lookup table may stop updating when the maximum allowable counter is reached. c. In one example, the counter may be predefined. Alternatively, it is signaled in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), slice header, tile header, coding tree unit (CTU), coding tree block (CTB), coding unit (CU) or prediction unit (PU), a region covering a plurality of CTU / CTB / CU / PU.

[0202] FIG. 28 is a block diagram of a video processing apparatus 2800. The apparatus 2800 may be used to implement one or more of the methods described herein. The apparatus 2800 may be implemented in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver for a single device, etc. The apparatus 2800 may include one or more processing devices 2802, one or more memories 2804, and video processing hardware 2806. The one or more processing devices 2802 may be configured to implement one or more of the methods described herein. The memory or memories 2804 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 2806 may be used to implement the techniques described herein in a hardware circuit.

[0203] FIG. 29 is a flowchart of an example of a video decoding method 2900. This method 2900 includes maintaining a number of tables (e.g., look-up tables: LUTs) (2902), each table including a set of motion candidates, each motion candidate being associated with corresponding motion information derived from a previously coded video block, and performing a conversion between the current video block and the bitstream representation of the current video block in a video region (2904), and updating one or more tables based on an update rule (2906).

[0204] FIG. 30 is a flowchart of an example of a video decoding method 3000. This method 3000 includes checking a set of tables (e.g., look-up tables; LUTs) (3002), each table including one or more motion candidates, each motion candidate being associated with the motion information of the motion candidate, and processing the motion information of this video block based on one or more tables (3004), and updating one or more tables based on the video block generated by the processing (3006).

[0205] For the purposes of the foregoing description, specific embodiments of the technology of the present disclosure have been described, but it will be understood that various modifications can be made without departing from the scope of the invention. Accordingly, the technology of the present disclosure is not limited except as by the appended claims.

[0206] The disclosed and other embodiments, modules, and implementations of functional operations described herein may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a computer-readable medium for being executed by, or to control the operation of, a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that provides a machine-readable propagated signal, or a combination of one or more of them. The term “data processing apparatus” includes, for example, all apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus may include, in addition to hardware, code that creates an execution environment for the computer program, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., an electrical, optical, or electromagnetic signal generated by a machine, that is generated to encode information for transmission to an appropriate receiver device.

[0207] A computer program (also referred to as a program, software, software application, script, or code) can be described in any form of programming language, including compiled languages or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program may be recorded as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), stored in a single file dedicated to the program, or stored in multiple coordinating files (e.g., files that hold one or more modules, subprograms, or portions of code). One computer program can be deployed to be executed on one computer located at one site or on multiple computers distributed across multiple sites and interconnected by a communication network.

[0208] The processes and logic flows described herein can be performed by one or more programmable processing devices that execute one or more computer programs to function by operating on input data and generating output. The processes and logic flows can also be performed by special-purpose logic circuits, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), and the apparatus can also be implemented as special-purpose logic circuits.

[0209] A processing device suitable for the execution of a computer program includes, for example, both general-purpose and special-purpose microprocessors, as well as any one or more processing devices of any kind of digital computer. Generally, the processing device receives instructions and data from read-only memory or random access memory or both. The essential elements of a computer are a processing device for executing instructions and one or more storage devices for storing instructions and data. Generally, a computer may include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or may be operatively coupled to receive data from or transfer data to these mass storage devices. However, a computer need not have such devices. A computer-readable medium suitable for storing computer program instructions and data includes any form of non-volatile memory, medium, and storage device, including, for example, EPROM, EEPROM, flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and semiconductor storage devices such as CD-ROM and DVD-ROM disks. The processing device and the memory may be complemented by, or incorporated in, dedicated logic circuitry.

[0210] This patent specification includes many details, but these should not be construed as limiting the scope of any invention or the scope of the claims. Rather, they should be construed as descriptions of features that may be specific to particular embodiments of a particular invention. Specific features described in the context of another embodiment in this patent specification may be implemented in combination in one example. Conversely, various features described in the context of a single example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, although features may be described and initially claimed above as acting in a particular combination, one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of the sub-combination.

[0211] Similarly, although operations are shown in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order to achieve the desired result, or that all of the operations shown be performed. Also, the separation of various system modules in the examples described in this patent specification should not be understood as requiring such separation in all embodiments.

[0212] Only some implementations and examples are described, and based on what is described and illustrated in this patent specification, other embodiments, extensions, and variations are possible.

Claims

1. A method for processing video data, comprising: maintaining one or more tables that are history-based motion vector prediction (HMVP) tables, each of the one or more tables including one or more motion candidates derived from one or more coded video blocks, wherein the arrangement of the motion candidates in the table is based on the sequence of adding the motion candidates to the table; constructing a motion candidate list for a current video block, wherein during the construction, at least one motion candidate of the at least one table among the one or more tables is selectively checked in sequence to determine whether to add the at least one motion candidate to the motion candidate list; using the motion candidate list to determine motion information of the current video block; performing conversion between the current video block and a bitstream based on the determined motion information; resetting the number of at least one motion candidate of the one or more tables to zero at the start of coding a new region including a plurality of video blocks; including; the one or more tables include N tables, N is equal to K*P, K is an integer representing the number of tables per coding thread corresponding to a CTU row or tile of one slice of the video data, P is an integer representing the number of CTU rows or tiles in the slice; a method.

2. The table among the one or more tables has a size indicating the number of permitted motion candidates in the table. The method according to claim 1.

3. If the number of motion candidates in the table has reached the size of the table before adding the new motion candidate to the table, one candidate in the table is deleted by adding the new motion candidate to the table. The method according to claim 2.

4. The value of the size of the table is one of a predefined value or a value signaled in a syntax element. The method according to claim 2.

5. Further comprising maintaining a counter for the table. The counter represents the number of motion candidates in the table. The counter is less than or equal to the size of the table. The method according to claim 2.

6. The new region is a new coding tree unit (CTU) row, a new tile, or a new slice. The method according to claim 1.

7. The number of the tables is predefined. The method according to claim 1.

8. When one CTU row or one tile of a slice of the video data uses one table and N is equal to P, where P is an integer representing the number of CTU rows or tiles in the slice, the one or more tables include N tables. The method according to claim 1.

9. The number of the one or more tables is based on at least one of slice type, picture temporal layer index, and picture order count (POC) distance between one slice and the nearest intra-slice. The method according to claim 1.

10. The motion candidates in the table are associated with motion information including at least one of a prediction direction, a reference picture index, a motion vector value, an intensity compensation flag, an affine flag, a motion vector difference accuracy, or a motion vector difference value. The method according to claim 1.

11. The motion candidate list is one of a motion vector prediction list or a merge candidate list. The method according to claim 1.

12. Based on the result of the check, add one or more checked candidates to the motion candidate list. The method according to claim 1.

13. The conversion includes encoding the current video block into the bitstream. The method according to claim 1.

14. The conversion includes decoding the current video block from the bitstream. The method according to any one of claims 1.

15. An apparatus for coding video data, comprising a processing device and a non-transitory memory storing instructions. When executed by the processing device, the instructions cause the processing device to maintain one or more tables that are history-based motion vector prediction (HMVP) tables, each of the one or more tables including one or more motion candidates derived from one or more coded video blocks, and the arrangement of the motion candidates in the table being based on the sequence in which the motion candidates are added to the table; and construct a motion candidate list for a current video block, and during the construction, selectively check in order at least one motion candidate of the table among the one or more tables to determine whether to add at least one of the motion candidates to the motion candidate list. Using the motion candidate list to determine motion information of the current video block; Performing conversion between the current video block and a bitstream based on the determined motion information; Resetting the number of at least one motion candidate of the one or more tables to zero at the start of coding of a new region including a plurality of video blocks; causing to perform, The one or more tables include N tables, N is equal to K*P, K is an integer representing the number of tables per coding thread corresponding to a CTU row or tile of one slice of the video data, P is an integer representing the number of CTU rows or tiles in the slice, apparatus.

16. A non-transitory computer-readable storage medium storing instructions, The instructions cause a processing device to Maintain one or more tables that are history-based motion vector prediction (HMV P) tables, wherein each of the one or more tables includes one or more motion candidates derived from one or more coded video blocks, and an arrangement of the motion candidates in the table is based on a sequence of adding the motion candidates to the table; Construct a motion candidate list for a current video block, and selectively check, in order, at least one motion candidate of the one or more tables to determine whether to add at least one of the motion candidates to the motion candidate list during the construction; Using the motion candidate list to determine motion information of the current video block; Performing conversion between the current video block and a bitstream based on the determined motion information; Reset the number of at least one motion candidate of the one or more tables to zero at the start of coding of a new region including a plurality of video blocks, and cause the one or more tables to include N tables, where N is equal to K*P, K is an integer representing the number of tables per coding thread corresponding to a CTU row or tile of one slice of video data, and P is an integer representing the number of CTU rows or tiles in the slice, a non-transitory computer-readable storage medium. **Claim 17**: A method of storing a bitstream, comprising: maintaining one or more tables that are history-based motion vector prediction (HMV) tables, each of the one or more tables including one or more motion candidates derived from one or more coded video blocks, wherein an arrangement of the motion candidates within the table is based on a sequence of adding the motion candidates to the table; constructing a motion candidate list for a current video block, wherein, during the constructing, selectively checking, in order, at least one motion candidate of the at least one of the one or more tables to determine whether to add the at least one motion candidate to the motion candidate list; determining motion information of the current video block using the motion candidate list; generating the bitstream based on the determined motion information; storing the bitstream in a non-transitory computer-readable recording medium; and resetting the number of at least one motion candidate of the one or more tables to zero at the start of coding of a new region including a plurality of video blocks. and the one or more tables include N tables, N is equal to K * P, K is an integer representing the number of tables per coding thread corresponding to a CTU row or tile of one slice of video data, P is an integer representing the number of CTU rows or tiles in the slice, Method.

Citation Information

Patent Citations

  • Motion vector generation device, prediction image generation device, moving image decoding device, and moving image encoding device

    WO2018061522A1