Video encoding / decoding methods, apparatus, and media relating to conditions for updating lookup tables.
By maintaining and updating the lookup table of motion candidates, the conversion process between video blocks and bitstream representations is optimized, solving the problems of coding efficiency and bandwidth requirements in high-resolution video coding, and achieving more efficient video compression.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2019-07-01
- Publication Date
- 2026-05-26
Smart Images

Figure CN115134599B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] In accordance with applicable patent law and / or the rules of the Paris Convention, this application is to promptly claim priority and benefits to International Patent Application No. PCT / CN2018 / 093663, filed on 29 June 2018; International Patent Application No. PCT / CN2018 / 094929, filed on 7 July 2018; International Patent Application No. PCT / CN2018 / 101220, filed on 18 August 2018; and International Patent Application No. PCT / CN2018 / 093987, filed on 2 July 2018. For all purposes under U.S. law, the entire disclosures of International Patent Application No. PCT / CN2018 / 093663, International Patent Application No. PCT / CN2018 / 094929, International Patent Application No. PCT / CN2018 / 101220, and International Patent Application No. PCT / CN2018 / 093987 are incorporated by reference as a part of this application. This application is a divisional application of application No. 201910586572.5, filed July 1, 2019, entitled “Conditions for Updating Lookup Tables (LUTs)”. Technical Field
[0003] This patent document relates to video encoding and decoding technologies, devices, and systems. Background Technology
[0004] Despite advancements in video compression, digital video still accounts for the largest share of bandwidth usage on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This paper describes devices, systems, and methods related to encoding and decoding digital video using a set of tables containing coding candidates. The described methods can be applied to both existing video coding standards (e.g., High Efficiency Video Coding (HEVC)) and future video coding standards or video codecs.
[0006] In one representative aspect, the disclosed technology can be used to provide a video processing method, the method comprising: maintaining tables, wherein each table includes a set of motion candidates and each motion candidate is associated with corresponding motion information; performing a conversion between a first video block and a bitstream representation of a video including the first video block based on the tables; and updating zero or more tables based on update rules after performing the conversion.
[0007] In another representative aspect, a table is maintained, wherein each table includes a set of motion candidates and each motion candidate is associated with corresponding motion information; a conversion between a first video block and a bitstream representation of the video including the first video block is performed based on the table; and after performing the conversion, one or more tables are updated based on one or more video regions in the video until an update termination criterion is met.
[0008] In another representative aspect, the disclosed technology can be used to provide another video processing method, which includes: maintaining one or more tables including motion candidates, each motion candidate being associated with corresponding motion information; reordering the motion candidates in at least one of the one or more tables; and performing a conversion between a first video block and a bitstream representation of the video including the first video block based on the reordered motion candidates in the at least one table.
[0009] In another representative aspect, the disclosed technology can be used to provide another video processing method, which includes: maintaining one or more tables including motion candidates, each motion candidate being associated with corresponding motion information; performing a conversion between a first video block and a bitstream representation of a video including the first video block using the one or more tables; and updating the one or more tables by adding additional motion candidates to the tables and reordering the motion candidates in the tables based on the conversion of the first video block.
[0010] In another representative aspect, the above method is implemented in the form of processor-executable code and stored in a computer-readable program medium.
[0011] In another representative aspect, a device configured or operable to perform the above-described methods is disclosed. The device may include a processor programmed to implement the methods.
[0012] In another representative aspect, video decoder devices can implement the methods described in this paper.
[0013] The above and other aspects and features of the disclosed technology are described in more detail in the accompanying drawings, description and claims. Attached Figure Description
[0014] Figure 1 A typical block diagram of a High Efficiency Video Coding (HEVC) video encoder and decoder is shown.
[0015] Figure 2 An example of macroblock (MB) segmentation in H.264 / AVC is shown.
[0016] Figure 3An example of dividing a coding block (CB) into prediction blocks (PB) is shown.
[0017] Figure 4A and Figure 4B Examples are shown for subdividing the Coding Tree Block (CTB) into CBs and Transform Blocks (TBs), and the corresponding quadtrees.
[0018] Figure 5A and Figure 5B The breakdown of the Largest Coding Unit (LCU) and the corresponding QTBT (QuadTree plus Binary Tree) are shown as an example.
[0019] Figures 6A-6E An example of segmented coded blocks is shown.
[0020] Figure 7 An example segmentation of CB based on QTBT is shown.
[0021] Figures 8A-8I An example of CB segmentation supporting a generalized multi-tree type (MTT) as QTBT is shown.
[0022] Figure 9A An example of tree-structured signaling is shown.
[0023] Figure 9B An example of constructing a Merge candidate list is shown.
[0024] Figure 10 An example of a spatial candidate location is shown.
[0025] Figure 11 An example of candidate pairs for which a spatial Merge candidate is performed is shown.
[0026] Figure 12A and Figure 12B An example of the position of the second prediction unit (PU) based on the size and shape of the current block is shown.
[0027] Figure 13 An example of motion vector scaling for time merge candidates is shown.
[0028] Figure 14 An example of candidate positions for the time-merge candidate is shown.
[0029] Figure 15 An example of generating bidirectional prediction Merge candidates using a combination is shown.
[0030] Figure 16A and Figure 16B An example of the derivation process for motion vector prediction candidates is shown.
[0031] Figure 17 An example of motion vector scaling for spatial motion vector candidates is shown.
[0032] Figure 18A and Figure 18B An example of motion prediction using the Alternative Temporal Motion Vector Prediction (ATMVP) algorithm for the Coding Unit (CU) is shown.
[0033] Figure 19 An example of identifying source blocks and source images is shown.
[0034] Figure 20 An example of a coding unit (CU) with sub-blocks and neighboring blocks is shown, which is used by the Spatial-Temporal Motion Vector Prediction (STMVP) algorithm.
[0035] Figure 21 An example of bidirectional matching in the Pattern Matched Motion Vector Derivation (PMMVD) mode is shown as a special Merge mode based on the Frame-Rate Up Conversion (FRUC) algorithm.
[0036] Figure 22 An example of template matching in the FRUC algorithm is shown.
[0037] Figure 23 An example of unidirectional motion estimation in the FRUC algorithm is shown.
[0038] Figure 24 An example of a decoder-side motion vector refinement (DMVR) algorithm based on bidirectional template matching is shown.
[0039] Figure 25 An example of adjacent samples used to derive the illumination compensation (IC) parameters is shown.
[0040] Figure 26 An example of neighboring blocks used to derive spatial Merge candidates is shown.
[0041] Figure 27 Examples of the proposed 67 intra-frame prediction modes are shown.
[0042] Figure 28 An example of adjacent blocks used for the most likely pattern derivation is shown.
[0043] Figure 29A and Figure 29B The corresponding luminance and chromaticity sub-blocks in the I-strip with the QTBT structure are shown.
[0044] Figure 30 An example is shown illustrating how to select a representative location for lookup table updates.
[0045] Figure 31A and Figure 31B An example of updating a lookup table with a new set of motion information is shown.
[0046] Figure 32 An example of the decoding flowchart using the proposed HMVP method is shown.
[0047] Figure 33 An example of updating a table using the proposed HMVP method is shown.
[0048] Figures 34A-34B An example of a LUT update method based on redundancy removal is shown (where one redundant motion candidate is removed).
[0049] Figures 35A-35B An example of a LUT update method based on redundancy removal is shown (where multiple redundant motion candidates are removed).
[0050] Figure 36 An example encoding flow for LUT-based MVP / intra-frame mode prediction / IC parameters updated after a block is shown.
[0051] Figure 37 An example encoding flow for LUT-based MVP / intra-frame mode prediction / IC parameters updated after a region is shown.
[0052] Figures 38A-38D A flowchart of an example method for video processing according to the present disclosure is shown.
[0053] Figure 39 This is a block diagram of an example hardware platform used to implement the visual media decoding or visual media encoding technologies described in this document. Detailed Implementation
[0054] Due to the ever-increasing demand for higher resolution video, video coding methods and technologies are ubiquitous in modern technology. Video codecs typically consist of electronic circuitry or software that compresses or decompresses digital video and are constantly being improved to provide higher coding efficiency. Video codecs convert uncompressed video into a compressed format and vice versa. There is a complex relationship between video quality, the amount of data used to represent the video (determined by the bit rate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, and end-to-end latency (delay). Compression formats typically conform to standard video compression specifications, such as the High Efficiency Video Coding (HEVC) standard (also known as H.265 or MPEG-H Part 2), upcoming general video coding standards, or other current and / or future video coding standards.
[0055] Embodiments of the disclosed techniques can be applied to existing video coding standards (e.g., HEVC, H.265) and future standards to improve compression performance. Section headings are used in this document to improve readability and do not in any way limit the discussion or embodiments (and / or implementations) to the relevant section only.
[0056] 1. Example implementation of video encoding
[0057] Figure 1 A typical block diagram of a HEVC video encoder and decoder is shown (reference document [1]). The encoding algorithm that produces a bitstream conforming to HEVC is generally performed as follows. Each picture is divided into block regions, where the precise block segmentation is passed to the decoder. The first picture of the video sequence (and the first picture at each cleanrandom access point in the video sequence) is encoded using only intra-frame prediction (using some data prediction spatially from region to region within the same picture, but not depending on other pictures). For all the remaining pictures in the sequence or between the random access points, the inter-frame temporal prediction coding mode is generally used for most blocks. The encoding process for inter-frame prediction involves selecting motion data including the selected reference picture and the motion vector (MV) of the samples to be applied to predict each block. The encoder and decoder generate the same inter-frame prediction signaling by applying motion compensation (MC) using MV and mode decision data, which is sent as side information.
[0058] The residual signal of intra-frame or inter-frame prediction is transformed using a linear spatial transform, which represents the difference between the original block and its prediction. The transform coefficients are then scaled, quantized, entropy-coded, and sent along with the prediction information.
[0059] Encoder replicates decoder processing loop (see...) Figure 1 The gray shaded box in the image (indicated by the fact that both will generate the same prediction for subsequent data) is used. Therefore, the quantized transform coefficients are constructed by inverse scaling and then inversely transformed to replicate the decoded approximation of the residual signal. The residual is then added to the prediction, and the result can then be fed into one or two loop filters to smooth out artifacts caused by block processing and quantization. The final picture representation (i.e., a copy of the decoder's output) is stored in the decoded picture buffer for prediction of subsequent pictures. Typically, the order in which pictures are encoded or decoded is different from the order in which they arrive from the source; a distinction needs to be made between the decoder's decoding order (i.e., bitstream order) and the output order (i.e., display order).
[0060] Video material encoded in HEVC is typically expected to be input as progressive scan imagery (since the source video originates from this format or is generated from deinterlacing prior to encoding). There are no explicit encoding features in the HEVC design to support the use of interlaced scanning because interlaced scanning is no longer used for displays and has become largely uncommon for distribution. However, a metadata syntax has been provided in HEVC to allow the encoder to indicate whether interlaced scan video has been transmitted by encoding each field of the interlaced video (i.e., the even or odd lines of each video frame) as a separate picture or by encoding each interlaced frame as an HEVC-encoded picture. This provides an efficient way to encode interlaced video without requiring the decoder to support a special decoding process for it.
[0061] 1.1. Example of a split tree structure in H.264 / AVC
[0062] The core of the coding layer in the previous standard was a macroblock, which contained a 16×16 luma sample block and, in the typical case of 4:2:0 color sampling, two corresponding 8×8 chroma sample blocks.
[0063] Intra-coded blocks use spatial prediction to take advantage of spatial correlations among pixels. Two partitions are defined: 16×16 and 4×4.
[0064] Inter-frame coded blocks use temporal prediction rather than spatial prediction to estimate motion within the image. Motion can be estimated independently for each 16×16 macroblock or any of its sub-macroblocks (16×8, 8×16, 8×8, 8×4, 4×8, 4×4), such as... Figure 2 As shown in reference document [2], only one motion vector (MV) is allowed per sub-macroblock.
[0065] 1.2. Example of a segmented tree structure in HEVC
[0066] In HEVC, Coding Tree Units (CTUs) are divided into Coding Units (CUs) using a quadtree structure represented as a coding tree to accommodate various local characteristics. The decision of whether to use inter-frame (temporal) or intra-frame (spatial) prediction to encode a picture region is made at the CU level. Each CU can be further divided into one, two, or four PUs based on the prediction unit (PU) partitioning type. Within a PU, the same prediction process is applied, and relevant information is sent to the decoder based on the PU. After obtaining residual blocks by applying a prediction process based on the PU partitioning type, the CUs can be partitioned into Transform Units (TUs) according to another quadtree structure similar to the CU. One of the key features of the HEVC structure is that it has multiple partitioning concepts, including CUs, PUs, and TUs.
[0067] Some of the features involved in hybrid video coding using HEVC include:
[0068] (1) Coding Tree Unit (CTU) and Coding Tree Block (CTB) Structure The analog structure in HEVC is the Code Tree Unit (CTU), which has a size selected by the encoder and can be larger than a traditional macroblock. A CTU consists of a luma CTB, a corresponding chroma CTB, and syntax elements. The size L×L of the luma CTB can be selected as L = 16, 32, or 64 samples, with larger sizes generally achieving better compression. HEVC then supports using a tree structure and quadtree-like signaling to divide the CTB into smaller blocks.
[0069] (2) Coding Unit (CU) and Coding Block (CB) The quadtree syntax of the CTU specifies the size and location of its luma CB and chroma CB. The root of the quadtree is associated with the CTU. Therefore, the size of the luma CTB is the maximum supported size of the luma CB. Dividing the CTU into luma CB and chroma CB is notified by joint signaling. One luma CB and usually two chroma CBs together with the associated syntax form a coding unit (CU). The CTB may contain only one CU or may be divided to form multiple CUs, and each CU has a tree of associated partitioning and transforming units (TUs) partitioned into prediction units (PUs).
[0070] (3) Prediction cells and prediction blocks (PB) The determination of whether to use inter-frame or intra-frame prediction to encode image regions is performed at the CU level. The root of the PU segmentation structure is at the CU level. Based on the basic prediction type determination, the luma CB and chroma CB can then be further subdivided in size and predicted according to luma and chroma prediction blocks (PB). HEVC supports variable PB sizes from 64×64 to 4×4 samples. Figure 3 An example of a allowed PB for an M×M CU is shown.
[0071] (4) Transformer Unit (TU) and Transformer Block The prediction residuals are encoded using block transforms. The root of the TU tree structure is at the CU level. The luma CB residuals can be the same as the luma transform block (TB), or they can be further divided into smaller luma TBs. The same applies to the chroma TB. Integer basis functions similar to the Discrete Cosine Transform (DCT) are defined for square TB sizes of 4×4, 8×8, 16×16, and 32×32. For the 4×4 transform of the luma intra-frame prediction residuals, integer transforms derived from the form of the Discrete Sine Transform (DST) are alternately specified.
[0072] 1.2.1. Example of tree structure partitioning into TB and TU
[0073] For residual coding, the CB can be recursively partitioned into transform blocks (TBs). The partitioning is notified by residual quadtree signaling. Only square CB and TB partitions are specified, where blocks can be recursively divided into quadrants, such as... Figure 4A and Figure 4B As shown. For a given luma CB of size M×M, a flag signaling indicates whether it is divided into four blocks of size M / 2×M / 2. If further division is possible, as indicated by the maximum depth signaling of the residual quadtree in the Sequence Parameter Set (SPS), a flag indicating whether it is divided into four quadrants is assigned to each quadrant. The leaf node blocks generated by the residual quadtree are transform blocks that are further processed by transform coding. The encoder indicates the maximum and minimum luma TB sizes it will use. Division is implicit when the CB size is greater than the maximum TB size. No division is implicit when division would result in a luma TB size less than the indicated minimum. Except when the luma TB size is 4×4 (in which case a single 4×4 chroma TB is used for the area covered by four 4×4 luma TBs), the chroma TB size is half the luma TB size in each dimension. In the case of intra-predictive CUs, the decoded samples of the nearest neighboring TBs (inside or outside the CB) are used as reference data for intra-predictive.
[0074] In contrast to previous standards, the HEVC design allows the TB to span multiple PBs of the inter-frame predicted CUs to maximize the potential coding efficiency benefits of TB segmentation in a quadtree structure.
[0075] 1.2.2. Parent Node and Child Node
[0076] The CTB is partitioned according to a quadtree structure, where nodes are coding units. The quadtree structure comprises multiple nodes, including leaf nodes and non-leaf nodes. Leaf nodes have no child nodes in the tree structure (i.e., leaf nodes are not further partitioned). Non-leaf nodes include the root node of the tree structure. The root node corresponds to the initial video block of the video data (e.g., the CTB). For each corresponding non-root node among the multiple nodes, the corresponding non-root node corresponds to a video block that is a child block of the video block corresponding to the parent node of the corresponding non-root node in the tree structure. Each corresponding non-leaf node among the multiple non-leaf nodes has one or more child nodes in the tree structure.
[0077] 1.3. Example of a quadtree plus binary tree block structure with a larger CTU in JEM
[0078] In some embodiments, reference software known as the Joint Exploration Model (JEM) is used to explore future video coding techniques. In addition to binary tree structures, the JEM also describes quadtree-plus-binary-tree (QTBT) and ternary tree (TT) structures.
[0079] 1.3.1. Example of QTBT block partitioning structure
[0080] Unlike HEVC, the QTBT structure removes the concept of multiple segmentation types; that is, it removes the separation of CU, PU, and TU concepts and supports greater flexibility in the shape of CU segmentation. In the QTBT block structure, CU can have a square or rectangular shape. Figure 5A As shown, coding unit (CTU) is first segmented using a quadtree structure. The leaf nodes of the quadtree are then further segmented using a binary tree structure. There are two types of binary tree partitioning: symmetrical horizontal partitioning and symmetrical vertical partitioning. The leaf nodes of the binary tree are called coding units (CUs), and this segmentation is used for prediction and transform processing without any further segmentation. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. In JEM, CUs sometimes consist of coding blocks (CBs) with different color components; for example, a CU in the case of P-stripes and B-stripes in a 4:2:0 chroma format contains one luma CB and two chroma CBs. Sometimes, CUs consist of CBs with a single component; for example, a CU in the case of I-stripes contains only one luma CB or only two chroma CBs.
[0081] Define the following parameters for the QTBT segmentation scheme:
[0082] --CTU size: The size of the root node of the quadtree, the same concept as in HEVC.
[0083] --MinQTSize: Minimum allowed size of quadtree leaf nodes
[0084] --MaxBTSize: Maximum allowed size of the binary tree root node
[0085] --MaxBTDepth: Maximum allowed binary tree depth
[0086] --MinBTSize: Minimum allowed size of binary leaf nodes
[0087] In one example of a QTBT segmentation structure, the CTU size is set to 128×128 luminance samples with two corresponding 64×64 chroma sample blocks, MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize (width and height) is set to 4×4, and MaxBTDepth is set to 4. Quadtree segmentation is first applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If a quadtree leaf node is 128×128, it will not be further segmented by the binary tree because its size exceeds MaxBTSize (i.e., 64×64). Otherwise, the quadtree leaf node can be further segmented by the binary tree. Therefore, the quadtree leaf node is also the root node of the binary tree, and its binary tree depth is 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), further segmentation is not considered. When a binary tree node has a width equal to MinBTSize (i.e., 4), further horizontal partitioning is not considered. Similarly, when a binary tree node has a height equal to MinBTSize, further vertical partitioning is not considered. Leaf nodes of the binary tree are further processed through prediction and transformation without any further segmentation. In JEM, the maximum CTU size is 256×256 luminance samples.
[0088] Figure 5A An example of block partitioning using QTBT is shown, and Figure 5B The corresponding tree representation is shown. Solid lines indicate quadtree partitions, and dashed lines indicate binary tree partitions. In each partition (i.e., non-leaf) node of a binary tree, a signaling flag indicates which partition type (i.e., horizontal or vertical) was used, where 0 indicates a horizontal partition and 1 indicates a vertical partition. For quadtree partitions, it is not necessary to indicate the partition type because quadtree partitions always divide blocks horizontally and vertically to produce 4 sub-blocks of the same size.
[0089] Furthermore, the QTBT scheme supports the ability for luma and chroma to have separate QTBT structures. Currently, for P-strip and B-strip, the luma CTB and chroma CTB within a CTU share the same QTBT structure. However, for I-strip, the luma CTB is divided into CUs using a QTBT structure, and the chroma CTB is divided into chroma CUs using a separate QTBT structure. This means that a CU in I-strip consists of a coded block of the luma component or coded blocks of the two chroma components, while a CU in P-strip or B-strip consists of coded blocks of all three color components.
[0090] In HEVC, inter-frame prediction for small blocks is restricted to reduce memory accesses for motion compensation, resulting in no bidirectional prediction for 4×8 and 8×4 blocks, and no inter-frame prediction for 4×4 blocks. These restrictions are removed in JEM's QTBT.
[0091] 1.4. Tritree (TT) of Multifunctional Video Coding (VVC)
[0092] Figure 6A An example of quadtree (QT) partitioning is shown, and Figure 6B and Figure 6C Examples of vertical and horizontal binary tree (BT) partitioning are shown respectively. In some embodiments, in addition to quadtrees and binary trees, ternary tree (TT) partitioning is also supported, such as horizontal and vertical center-side ternary trees (e.g., Figure 6D and Figure 6E (As shown).
[0093] In some implementations, a two-level tree is supported: a region tree (quadtree) and a prediction tree (binary or ternary). The CTU is first segmented using a region tree (RT). RT leaves can be further subdivided using a prediction tree (PT). PT leaves can also be further subdivided using PT until the maximum PT depth is reached. PT leaves are the basic coding units. For convenience, they are still referred to as CUs. CUs cannot be further subdivided. Prediction and transformation are applied to the CUs in the same way as JEM. The entire segmentation structure is named a "multi-type tree".
[0094] 1.5. Examples of Segmentation Structures in Optional Video Coding Techniques
[0095] In some embodiments, a tree structure called a Multi-Tree Type (MTT) is supported as a generalized form of QTBT. In QTBT, such as... Figure 7 As shown, the coding tree unit (CTU) is first segmented using a quadtree structure. The leaf nodes of the quadtree are then further segmented using a binary tree structure.
[0096] MTT's structure consists of two types of tree nodes: Region Tree (RT) and Prediction Tree (PT), supporting nine types of splits, such as... Figures 8A-8I As shown. The region tree can recursively divide the CTU into squares until the leaf nodes of the region tree are 4×4 in size. At each node in the region tree, a prediction tree can be formed from one of three tree types (binary tree, ternary tree, and asymmetric binary tree). In PT partitioning, quadtree splits are prohibited in the branches of the prediction tree. As in JEM, the luminance tree and chrominance tree are separated in the I stripe. Figure 9A The signaling methods for RT and PT are shown in the figure.
[0097] 2. Examples of inter-frame prediction in HEVC / H.265
[0098] Over the years, video coding standards have improved significantly and now offer, in part, high coding efficiency and support for higher resolutions. Latest standards such as HEVC and H.265 are based on a hybrid video coding structure that utilizes time prediction plus transform coding.
[0099] 2.1 Examples of Predictive Patterns
[0100] Each inter-frame prediction PU (prediction unit) has motion parameters for one or two lists of reference images. In some embodiments, the motion parameters include motion vectors and reference image indices. In other embodiments, the use of one of the two reference image lists can also be notified using inter_pred_idc signaling. In other embodiments, the motion vectors can be explicitly encoded as increments relative to the predictor.
[0101] When a CU is encoded using the skip mode, a PU is associated with a CU, and there are no significant residual coefficients, no encoded motion vector increments, or reference picture indices. A Merge mode is specified, thereby obtaining the motion parameters of the current PU from neighboring PUs, including spatial and temporal candidates. The Merge mode can be applied to any PU for inter-frame prediction, not just skip mode. An alternative to the Merge mode is explicit transmission of motion parameters, where each PU explicitly signals the motion vectors, the corresponding reference picture index for each reference picture list, and the reference picture list usage.
[0102] When signaling indicates that one of two reference image lists should be used, a PU is generated from a sample block. This is called "one-way prediction". One-way prediction can be used for both P-strips and B-strips.
[0103] When signaling indicates that both from the reference image list should be used, a PU is generated from the two sample blocks. This is called "bidirectional prediction". Bidirectional prediction can only be used for B-strips.
[0104] 2.1.1 Implementation Examples for Constructing Candidate Merge Patterns
[0105] When predicting a PU using the Merge pattern, the indices pointing to entries in the Merge candidate list are parsed from the bitstream and used to retrieve motion information. The construction of this list can be outlined according to the following sequence of steps:
[0106] Step 1: Initial Candidate Derivation
[0107] Step 1.1: Spatial Candidate Derivation
[0108] Step 1.2: Redundancy check of spatial candidates
[0109] Step 1.3: Derivation of Time Candidates
[0110] Step 2: Add candidate insertions
[0111] Step 2.1: Create bidirectional prediction candidates
[0112] Step 2.2: Insert zero-motion candidates
[0113] Figure 9B An example of constructing a merge candidate list based on the steps outlined above is shown. For spatial merge candidate derivation, up to four merge candidates are selected from candidates located at five distinct positions. For temporal merge candidate derivation, up to one merge candidate is selected from two candidates. Since a constant number of candidates is assumed per PU at the decoder, additional candidates are generated when the number of candidates does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the stripe header. Because the number of candidates is constant, unary binarization (TU) is used to encode the index of the best merge candidate. If the size of the CU is equal to 8, all PUs of the current CU share a single merge candidate list, which is the same as the merge candidate list of a 2N×2N prediction unit.
[0114] 2.1.2 Constructing Space Merge Candidates
[0115] In the derivation of the spatial Merge candidates, located at Figure 10 Up to four merge candidates are selected from the candidates at the positions depicted. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, or A0 is unavailable (e.g., because it belongs to another stripe or slice) or if it is intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving coding efficiency.
[0116] To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only pairs with... Figure 11 The arrows in the list link pairs, and if the corresponding candidate used for redundancy checking does not have the same motion information, the candidate is only added to the list. Another source of duplicate motion information is a "second PU" associated with a segmentation different from 2N×2N. As an example, Figure 12A and Figure 12B The second PU is depicted for N×2N and 2N×N scenarios, respectively. When the current PU is segmented into N×2N, the candidate at position A1 is not considered for list construction. In some embodiments, adding this candidate may result in two prediction units having the same motion information, which is redundant for a coding unit with only one PU. Similarly, when the current PU is segmented into 2N×N, position B1 is not considered.
[0117] 2.1.3 Constructing Temporal Merge Candidates
[0118] In this step, only one candidate is added to the list. Specifically, in the derivation of the time-merge candidate, the scaling motion vector is derived based on the co-localized PU belonging to the image with the smallest POC difference from the current image within the given list of reference images. The list of reference images to be used for deriving the co-localized PU is explicitly signaled in the strip header.
[0119] Figure 13 An example of the derivation of the scaled motion vector for the temporal merge candidate (e.g., the dashed line) is shown, which is scaled from the motion vector of the co-localized PU using POC distances tb and td, where tb is defined as the POC difference between the current image and the reference image of the current image, and td is defined as the POC difference between the reference image of the co-localized image and the co-localized image. The reference image index of the temporal merge candidate is set to zero. For the B-strip, two motion vectors (one for reference image list 0 and the other for reference image list 1) are obtained and combined to produce a bidirectional predicted merge candidate.
[0120] In the co-location PU(Y) belonging to the reference frame, the position of the time candidate is selected between candidate C0 and C1, such as... Figure 14 As described in the text. If the PU at position C0 is unavailable, intra-coded, or outside the current CTU, then position C1 is used. Otherwise, position C0 is used in the derivation of the time merge candidate.
[0121] 2.1.4 Constructing Merge Candidates for Additional Types
[0122] In addition to spatiotemporal merge candidates, there are two additional types of merge candidates: combined bidirectional predictive merge candidates and zero merge candidates. Combined bidirectional predictive merge candidates are generated by utilizing spatiotemporal merge candidates. These combined bidirectional predictive merge candidates are only used for B-strips. Combined bidirectional predictive candidates are generated by combining the motion parameters of the first reference image list of the initial candidate with the motion parameters of the second reference image list of another initial candidate. If these two tuples provide different motion hypotheses, they will form new bidirectional predictive candidates.
[0123] Figure 15 An example of the process is shown, in which two candidates from the original list (1510 on the left) with mvL0 and refIdxL0 or mvL1 and refIdxL1 are used to create a bidirectional prediction Merge candidate that is added to the final list (1520 on the right).
[0124] Zero-motion candidates are inserted to populate the remaining entries in the Merge candidate list, thus reaching the MaxNumMergeCand capacity. These candidates have a zero-space displacement and a reference image index, which starts from zero and increases whenever a new zero-motion candidate is added to the list. The number of reference frames used by these candidates is 1 for unidirectional prediction and 2 for bidirectional prediction. In some embodiments, redundancy checks are not performed on these candidates.
[0125] 2.1.5 Example of motion estimation region for parallel processing
[0126] To accelerate the encoding process, motion estimation can be performed in parallel, thereby simultaneously deriving the motion vectors of all prediction units within a given region. Deriving merge candidates from spatial neighborhoods can interfere with parallel processing because a prediction unit cannot derive motion parameters from its neighboring PUs until its associated motion estimation is completed. To mitigate the trade-off between encoding efficiency and processing latency, a Motion Estimation Region (MER) can be defined. The size of the MER can be signaled in the Picture Parameter Set (PPS) using the "log2_parallel_merge_level_minus2" syntax element. When defining the MER, merge candidates falling into the same region are marked as unavailable and therefore not considered in the list construction.
[0127] Table 1 shows the syntax of the Raw Byte Sequence Payload (RBSP) for the Picture Parameter Set (PPS). The value of the variable Log2ParMrgLevel, specified by log2_parallel_merge_level_minus2 plus 2, is used in the derivation of the luminance motion vector for the Merge mode and in the derivation of the spatial Merge candidates specified in existing video coding standards. The value of log2_parallel_merge_level_minus2 must be in the range of 0 to CtbLog2SizeY-2, inclusive.
[0128] The variable Log2ParMrgLevel is derived as follows:
[0129] Log2ParMrgLevel=log2_parallel_merge_level_minus2+2
[0130] Note that the value of Log2ParMrgLevel indicates the built-in capability for parallel derivation of the Merge candidate list. For example, when Log2ParMrgLevel equals 6, the Merge candidate list for all prediction units (PUs) and coding units (CUs) contained in a 64×64 block can be derived in parallel.
[0131] Table 1: General Image Parameter Settings RBSP Syntax
[0132]
[0133] 2.2 Examples of Motion Vector Prediction in AMVP Mode
[0134] Motion vector prediction utilizes the spatiotemporal correlation between motion vectors and adjacent physical units (PUs) for explicit transmission of motion parameters. It constructs a motion vector candidate list by first checking the availability of temporally adjacent PU locations to the left and above, removing redundant candidates, and adding zero vectors to maintain a constant candidate list length. The encoder can then select the best predictor from the candidate list and send the corresponding index indicating the selected candidate. Similar to Merge index signaling, a rounding unary is used to encode the index of the best motion vector candidate.
[0135] 2.2.1 Example of constructing motion vector prediction candidates
[0136] Figure 16A and Figure 16B The derivation process for motion vector prediction candidates is outlined, and it can be implemented for each list of reference images with refidx as input.
[0137] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For the derivation of spatial motion vector candidates, the final result is based on the previously mentioned... Figure 10 The motion vector derivation for each PU at the five different locations shown in the figure derives two motion vector candidates.
[0138] For temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates derived based on two distinct co-localization locations. After generating the first spatiotemporal candidate list, duplicate motion vector candidates in the list are removed. If the number of potential candidates is greater than two, motion vector candidates with a reference image index greater than 1 in their associated reference image list are removed from the list. If the number of spatiotemporal motion vector candidates is less than two, additional zero motion vector candidates are added to the list.
[0139] 2.2.2 Constructing candidate spatial motion vectors
[0140] In the derivation of the spatial motion vector candidates, from the position as previously stated... Figure 10 At most two of the five potential candidates for the PU derivation at the positions shown are considered. These positions are the same as the positions of the motion merge. The derivation order to the left of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order to the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. For each side, there are therefore four cases that can be used as motion vector candidates, two of which do not require spatial scaling, and two of which do. The four different cases are summarized below:
[0141] --No space scaling
[0142] (1) The same list of reference images, and the same index of reference images (the same POC).
[0143] (2) Different lists of reference images, but the same reference image (same POC) -- spatial scaling
[0144] (3) Same list of reference images, but different reference images (different POCs)
[0145] (4) A list of different reference images, and different reference images (different POCs).
[0146] First, we examine the non-spatial scaling case, then the case where spatial scaling is allowed. Spatial scaling is considered when the POC differs between the reference image of the adjacent PU and the reference image of the current PU, regardless of the reference image list. If all candidate PUs on the left are unavailable or intra-coded, scaling for the above motion vectors is allowed to aid in the parallel derivation of the left and top MV candidates. Otherwise, spatial scaling is not allowed for the above motion vectors.
[0147] like Figure 17 As shown in the example, for spatial scaling, the motion vectors of adjacent PUs are scaled in a similar manner to temporal scaling. One difference is that a list of reference images and the index of the current PU are given as input; the actual scaling process is the same as that for temporal scaling.
[0148] 2.2.3 Constructing candidate time motion vectors
[0149] Except for the derivation of the reference image index, all the procedures for deriving the temporal merge candidate are the same as those for deriving the spatial motion vector candidate (e.g., ...). Figure 14 (As shown in the example). In some embodiments, reference image index signaling is notified to the decoder.
[0150] 2.2.4 Signaling for Merge / AMVP Messages
[0151] For AMVP mode, four parts can be signaled in the bitstream, such as prediction direction, reference index, MVD, and MV predictor candidate index, which are described in the context of the syntax shown in Table 2-4. For Merge mode, only the Merge index may need to be signaled.
[0152] Table 2: General Strip Segmentation Header Syntax
[0153]
[0154] Table 3: Prediction Unit Syntax
[0155]
[0156] Table 4: Syntax of Motion Vector Difference
[0157]
[0158] The corresponding semantics include:
[0159] `five_minus_max_num_merge_cand` specifies the maximum number of Merge MVP candidates supported from the stripes subtracted from 5. The maximum number of Merge MVP candidates, `MaxNumMergeCand`, is derived as follows:
[0160] MaxNumMergeCand=5-five_minus_max_num_merge_cand
[0161] The value of MaxNumMergeCand must be in the range of 1 to 5, inclusive.
[0162] `merge_flag[x0][y0]` specifies whether the inter-frame prediction parameters of the current prediction unit are inferred from adjacent inter-frame prediction segments. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the considered prediction block relative to the top-left luminance sample of the image.
[0163] When merge_flag[x0][y0] does not exist, the following inference is made:
[0164] --If CuPredMode[x0][y0] equals MODE_SKIP, then merge_flag[x0][y0] is inferred to be equal to 1.
[0165] Otherwise, merge_flag[x0][y0] is inferred to be equal to 0.
[0166] merge_idx[x0][y0] specifies the Merge candidate index in the Merge candidate list, where x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the prediction block under consideration relative to the top-left luminance sample of the image.
[0167] 3. Examples of inter-frame prediction methods in the Joint Exploration Model (JEM)
[0168] In some embodiments, reference software known as the Joint Exploration Model (JEM) is used to explore future video coding techniques. In the JEM, sub-block-based predictions are employed across several coding tools, such as affine prediction, optional temporal motion vector prediction (ATMVP), spatiotemporal motion vector prediction (STMVP), bidirectional optical flow (BIO), frame rate upconversion (FRUC), locally adaptive motion vector resolution (LAMVR), overlapped block motion compensation (OBMC), local illumination compensation (LIC), and decoder-side motion vector refinement (DMVR).
[0169] 3.1 Example of motion vector prediction based on sub-CU
[0170] In a JEM with a quadtree plus binary tree (QTBT), each CU can have at most one set of motion parameters for each prediction direction. In some embodiments, two sub-CU level motion vector prediction methods are considered in the encoder by dividing the large CU into sub-CUs and deriving the motion information of all sub-CUs of the large CU. An optional temporal motion vector prediction (ATMVP) method allows each CU to extract multiple sets of motion information from multiple blocks smaller than the current CU in the juxtaposed reference image. In the spatiotemporal motion vector prediction (STMVP) method, the motion vectors of sub-CUs are recursively derived using a temporal motion vector predictor and spatially adjacent motion vectors. In some embodiments, and to preserve a more accurate motion field for sub-CU motion prediction, motion compression of the reference frame can be disabled.
[0171] 3.1.1 Example of Optional Temporal Motion Vector Prediction (ATMVP)
[0172] In the ATMVP method, the temporal motion vector prediction (TMVP) method is modified by extracting multiple sets of motion information (including motion vectors and reference indices) from blocks smaller than the current CU.
[0173] Figure 18A and Figure 18B An example of the ATMVP motion prediction process for CU 1800 is shown. The ATMVP method predicts the motion vectors of sub-CUs 1801 within CU 1800 in two steps. The first step is to identify the corresponding block 1851 in reference image 1850 using time vectors. Reference image 1850 is also called the motion source image. The second step is to divide the current CU 1800 into sub-CUs 1801 and obtain the motion vectors from the blocks corresponding to each sub-CU, as well as the reference index for each sub-CU.
[0174] In the first step, reference image 1850 and the corresponding block are determined by the motion information of spatially adjacent blocks of the current CU 1800. To avoid repeated scanning of adjacent blocks, the first Merge candidate in the Merge candidate list of the current CU 1800 is used. The first available motion vector and its associated reference index are set to the index of the time vector and the motion source image. In this way, the corresponding block can be identified more accurately than TMVP, where the corresponding block (sometimes called the juxtaposed block) is always located in the lower right or center position relative to the current CU.
[0175] In one example, if the first Merge candidate comes from the left adjacent block (i.e. Figure 19 In A1), the associated MV and reference image are used to identify the source block and source image.
[0176] In the second step, the corresponding block of sub-CU 1851 is identified by adding a time vector to the coordinates of the current CU, using the time vector in the motion source image 1850. For each sub-CU, the motion information of its corresponding block (e.g., the smallest motion grid covering the center sample) is used to derive the sub-CU's motion information. After identifying the motion information of the corresponding N×N block, it is converted into the motion vector and reference index of the current sub-CU in the same manner as the TMVP of HEVC, where motion scaling and other processes apply. For example, the decoder checks whether a low-latency condition is met (e.g., the POC of all reference images of the current image is less than the POC of the current image), and may use the motion vector MVx (e.g., the motion vector corresponding to the reference image list X) to predict the motion vector MVy of each sub-CU (e.g., where X equals 0 or 1, and Y equals 1-X).
[0177] 3.1.2 Example of Spatiotemporal Motion Vector Prediction (STMVP)
[0178] In the STMVP method, the motion vector of the sub-CU is recursively derived according to the raster scan order. Figure 20 An example of a CU with four sub-blocks and adjacent blocks is shown. Consider an 8×8 CU 2000 comprising four 4×4 sub-CUs: A (2001), B (2002), CUC (2003), and D (2004). Adjacent 4×4 blocks in the current frame are labeled a (2011), b (2012), c (2013), and d (2014).
[0179] Motion derivation of sub-CU A begins by identifying its two spatial neighborhoods. The first neighborhood is the N×N block (block c 2013) above sub-CU A2001. If block c (2013) is unavailable or intra-coded, the other N×N blocks above sub-CU A (2001) are checked (starting from block c 2013, from left to right). The second neighborhood is the block to the left of sub-CU A2001 (block b2012). If block b (2012) is unavailable or intra-coded, the other blocks to the left of sub-CU A2001 are checked (starting from block b2012, from top to bottom). Motion information obtained from neighboring blocks in each list is scaled to the first reference frame for the given list. Next, the temporal motion vector predictor (TMVP) for sub-block A2001 is derived by following the same procedure as the TMVP derivation specified in HEVC. Motion information of the juxtaposed block at block D 2004 is extracted and scaled accordingly. Finally, after retrieving and scaling the motion information, all available motion vectors are averaged separately for each reference list. The averaged motion vector is specified as the motion vector of the current sub-CU.
[0180] 3.1.3 Example of Sub-CU Motion Prediction Mode Signaling
[0181] In some embodiments, the sub-CU mode is enabled as an additional Merge candidate, and no additional syntax element is required to signal the mode. Two additional Merge candidates are added to the Merge candidate list of each CU to represent the ATMVP mode and the STMVP mode. In other embodiments, up to seven Merge candidates can be used if the sequence parameter set indicates that ATMVP and STMVP are enabled. The encoding logic of the additional Merge candidates is the same as that of the Merge candidates in the HM, which means that for each CU in a P-strip or B-strip, two RD checks may be required for the two additional Merge candidates. In some embodiments, such as JEM, all bits of the Merge index are context-encoded using CABAC (Context-based Adaptive Binary Arithmetic Coding). In other embodiments, such as HEVC, only the first bit is context-encoded, and the remaining bits are context-bypass encoded.
[0182] 3.2 Example of Adaptive Motion Vector Difference Resolution
[0183] In some embodiments, when the use_integer_mv_flag in the stripe header is equal to 0, the motion vector difference (MVD) is signaled in units of quarter-luminance samples (QMS). Local Adaptive Motion Vector Resolution (LAMVR) is introduced in JEM. In JEM, MVD can be encoded in units of quarter-luminance samples, integer-luminance samples, or four-luminance samples. MVD resolution is controlled at the coding unit (CU) level, and for each CU having at least one non-zero MVD component, an MVD resolution flag is conditionally signaled.
[0184] For a CU with at least one non-zero MVD component, signaling informs a first flag to indicate whether quarter-luminance sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter-luminance sample MV precision is not used, signaling informs another flag to indicate whether integer luminance sample MV precision or four-luminance sample MV precision is used.
[0185] When the first MVD resolution flag of the CU is zero or not encoded for the CU (meaning all MVDs in the CU are zero), a quarter-luminance sample MV resolution is used for the CU. When the CU uses integer luminance sample MV precision or quadruple luminance sample MV precision, the MVPs in the CU's AMVP candidate list are rounded to the corresponding precision.
[0186] In the encoder, CU-level RD checks are used to determine which MVD resolution should be used for the CU. That is, for each MVD resolution, a CU-level RD check is performed three times. To speed up the encoder, the following encoding scheme is applied in JEM:
[0187] --During the RD check of a CU with a normal quarter-luminance sample MVD resolution, the motion information of the current CU (integer luminance sample accuracy) is stored. The stored motion information (rounded) is used as the starting point for further small-scale motion vector refinement during the RD check of the same CU with integer luminance sample and 4-luminance sample MVD resolutions, so that the time-consuming motion estimation process is not repeated three times.
[0188] --Conditionally invoke the RD check for a CU with 4-luminance sample MVD resolution. For a CU, skip the RD check for 4-luminance sample MVD resolution if the RD cost for integer luminance sample MVD resolution is much greater than the RD cost for quarter-luminance sample MVD resolution.
[0189] 3.2.1 Example of constructing an AMVP candidate list
[0190] In JEM, this process is similar to the HEVC design. However, when the current block selects a lower-precision MV (e.g., integer precision), a rounding operation can be applied. In the current implementation, after selecting two candidates from spatial location, if both are available, the two candidates are rounded and then pruned.
[0191] 3.3 Example of Motion Vector Derivation (PMMVD) for Pattern Matching
[0192] PMMVD mode is a special merge mode based on the Frame Rate Upconversion (FRUC) method. With this mode, motion information for blocks is not signaled but rather derived on the decoder side.
[0193] When the Merge flag of the CU is true, the FRUC flag can be signaled to the CU. When the FRUC flag is false, the Merge index can be signaled to use the regular Merge mode. When the FRUC flag is true, the additional FRUC mode flag can be signaled to indicate which method (e.g., bidirectional matching or template matching) to use to derive the block's motion information.
[0194] On the encoder side, the decision on whether to use the FRUC Merge mode for the CU is based on the RD cost selection made for normal Merge candidates. For example, multiple matching modes (e.g., bidirectional matching and template matching) are checked for the CU using RD cost selection. The matching mode that produces the minimum cost is further compared with other CU modes. If the FRUC matching mode is the most efficient mode, the FRUC flag is set to true for the CU, and the relevant matching mode is used.
[0195] Typically, the motion derivation process in the FRUC Merge model involves two steps: first, a CU-level motion search is performed, followed by sub-CU-level motion refinement. At the CU level, an initial motion vector is derived for the entire CU based on bidirectional matching or template matching. First, a candidate MV list is generated, and the candidate producing the minimum matching cost is selected as the starting point for further CU-level refinement. Then, a local search based on bidirectional matching or template matching is performed around the starting point. The MV result with the minimum matching cost is taken as the MV for the entire CU. Subsequently, using the derived CU motion vector as the starting point, the motion information is further refined at the sub-CU level.
[0196] For example, the following derivation process is performed for the motion information derivation of a W×HCU. In the first stage, the MV of the entire W×HCU is derived. In the second stage, the CU is further divided into M×M sub-CUs. The value of M is calculated according to formula (3). D is a predefined partitioning depth that is set to 3 by default in JEM. Then the MV of each sub-CU is derived.
[0197]
[0198] Figure 21 An example of bidirectional matching used in the Frame Rate Upconversion (FRUC) method is shown. Bidirectional matching is used to derive motion information of the current CU (2100) by finding the closest match between two blocks along the motion trajectory of the current CU (2100) in two different reference images (2110, 2111). Under the assumption of continuous motion trajectories, the motion vectors MV0 (2101) and MV1 (2102) pointing to the two reference blocks are proportional to the temporal distances between the current image and the two reference images (e.g., TD0 (2103) and TD1 (2104)). In some embodiments, when the current image 2100 is temporally between the two reference images (2110, 2111) and the temporal distances from the current image to the two reference images are the same, the bidirectional matching becomes a mirror-based bidirectional MV.
[0199] Figure 22An example of template matching used in the Frame Rate Upconversion (FRUC) method is shown. Template matching can be used to derive motion information of the current CU 2200 by finding the closest match between a template in the current image (e.g., the top and / or left adjacent block of the current CU) and a block in the reference image 2210 (e.g., of the same size as the template). In addition to the FRUC Merge mode described above, template matching can also be applied to the AMVP mode. In JEM and HEVC, AMVP has two candidates. Using the template matching method, a new candidate can be derived. If the newly derived candidate through template matching is different from the first existing AMVP candidate, it is inserted at the very beginning of the AMVP candidate list, and then (e.g., by removing the second existing AMVP candidate) the list size is set to two. When applied to the AMVP mode, only the CU-level search is applied.
[0200] The MV candidate set at the CU level may include the following: (1) the original AMVP candidate if the current CU is in AMVP mode, (2) all Merge candidates, (3) several MVs in the interpolated MV field (described later), and the upper and left adjacent motion vectors.
[0201] When using bidirectional matching, each valid MV of the Merge candidate can be used as input to generate MV pairs under the assumption of bidirectional matching. For example, a valid MV of the Merge candidate is (MVa, ref) a At reference list A. Then, find the reference image ref for its paired bidirectional MV in other reference list B. b , making ref a and ref b It is located on a different side of the current image in time. If such a ref b If it is not available in reference list B, then ref b Determined to be related to ref a Different references, and its time distance to the current image is the smallest among list B. In determining the reference... b Then, based on the current image and Refa and Ref... b The time distance between them is derived from MVb by scaling MVa.
[0202] In some implementations, four MVs from the interpolated MV field can also be added to the CU-level candidate list. More specifically, the interpolated MVs at the current CU positions (0, 0), (W / 2, 0), (0, H / 2), and (W / 2, H / 2) are added. When FRUC is applied to the AMVP pattern, the original AMVP candidates are also added to the CU-level MV candidate set. In some implementations, at the CU level, 15 MVs from the AMVP CU and 13 MVs from the Merge CU can be added to the candidate list.
[0203] The candidate set of MVs at the sub-CU level includes MVs determined from the CU level search, (2) top, left, top-left, and top-right adjacent MVs, (3) scaled versions of juxtaposed MVs from the reference image, (4) one or more ATMVP candidates (e.g., up to four), and (5) one or more STMVP candidates (e.g., up to four). The scaled MVs from the reference image are derived as follows. The reference images in both lists are traversed. The MVs at the juxtaposed positions of the sub-CUs in the reference image are scaled to the reference of the MVs at the starting CU level. The ATMVP and STMVP candidates can be the first four. At the sub-CU level, one or more MVs (e.g., up to 17) are added to the candidate list.
[0204] Generation of interpolated MV fields Before encoding the frames, an interpolated motion field is generated for the entire image based on a one-way ME. The motion field can then be used later as a CU-level or sub-CU-level MV candidate.
[0205] In some embodiments, the motion field of each reference image in the two reference lists is traversed at a 4×4 block level. Figure 23 An example of unidirectional motion estimation (ME)2300 in the FRUC method is shown. For each 4×4 block, if the motion associated with the block traverses the 4×4 blocks in the current image and the block is not assigned any interpolated motion, the motion of the reference block is scaled to the current image according to temporal distances TD0 and TD1 (in the same way as the MV scaling of TMVP in HEVC), and the scaled motion is assigned to the block in the current frame. If no scaled MV is assigned to the 4×4 block, the block's motion is marked as unavailable in the interpolated motion field.
[0206] Interpolation and matching costs When the motion vector points to the fractional sample point location, motion-compensated interpolation is required. To reduce complexity, bilinear interpolation can be used instead of conventional 8-tap HEVC interpolation for both bidirectional matching and template matching.
[0207] The calculation of matching cost differs slightly at different steps. When selecting candidates from the candidate set at the CU level, the matching cost can be the absolute sum difference (SAD) of bidirectional matching or template matching. After determining the starting MV, the matching cost C of bidirectional matching in the sub-CU level search is calculated as follows:
[0208]
[0209] Here, w is a weighting factor. In some embodiments, w can be empirically set to 4. MV and MV s These indicate the current MV and the starting MV, respectively. SAD can still be used as the matching cost for template matching in sub-CU level searches.
[0210] In FRUC mode, the motion signature (MV) is derived using only luma samples. The derived motion is used for both luma and chroma predictions in the inter-frame prediction (MC). After determining the MV, the final MC is performed using an 8-tap interpolation filter for luma and a 4-tap interpolation filter for chroma.
[0211] MV refinement is a pattern-based MV search with a criterion of bidirectional matching cost or template matching cost. JEM supports two search modes—Unrestricted Center-Biased Diamond Search (UCBDS) and adaptive cross-search for MV refinement at the CU and sub-CU levels, respectively. For CU and sub-CU level MV refinement, MVs are directly searched with quarter-luminance sample MV accuracy, followed by eighth-luminance sample MV refinement. The search range for MV refinement at the CU and sub-CU levels is set to equal 8 luminance samples.
[0212] In the bidirectional matching Merge mode, bidirectional prediction is applied because the motion information of the CU is derived based on the closest match between two blocks along the motion trajectory of the current CU in two different reference images. In the template matching Merge mode, the encoder can choose from unidirectional prediction from list 0, unidirectional prediction from list 1, or bidirectional prediction of the CU. The selection can be based on the template matching cost, as follows:
[0213] If costBi <= factor * min(cost0, cost1)
[0214] Use two-way forecasting;
[0215] Otherwise, if cost0 <= cost1
[0216] Use one-way prediction from list0;
[0217] otherwise,
[0218] Use one-way prediction from list1;
[0219] Here, cost0 is the SAD of template matching for list0, cost1 is the SAD of template matching for list1, and costBi is the SAD of bidirectional prediction template matching. For example, when the factor value is equal to 1.25, this means that the selection process is biased towards bidirectional prediction. Inter-frame prediction direction selection can be applied to the CU-level template matching process.
[0220] 3.4 Example of Decoder-Side Motion Vector Refinement (DMVR)
[0221] In bidirectional prediction, to predict a block region, two prediction blocks formed using motion vectors (MVs) from list0 and list1, respectively, are combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two bidirectional prediction motion vectors are further refined through a bidirectional template matching process. Bidirectional template matching is applied in the decoder to perform a distortion-based search between the bidirectional template and reconstructed samples in the reference image to obtain a refined MV without sending additional motion information.
[0222] In DMVR, bidirectional templates are generated from the initial MV0 of list0 and the MV1 of list1, respectively, as a weighted combination (i.e., average) of the two prediction blocks, such as... Figure 24 As shown. The template matching operation involves calculating a cost metric between the generated template and the sample region (around the initial prediction block) in the reference image. For each of the two reference images, the MV that produces the minimum template cost is considered the updated MV of that list to replace the original MV. In JEM, nine MV candidates are searched for each list. The nine MV candidates include the original MV and eight surrounding MVs, one of which has a brightness sample offset from the original MV in the horizontal or vertical direction or both directions. Finally, two new MVs (i.e., as shown) are selected. Figure 24 The MV0' and MV1' shown are used to generate the final bidirectional prediction results. The sum of absolute errors (SAD) is used as a cost metric.
[0223] DMVR is applied to bidirectional prediction merge patterns, where one MV comes from a past reference image and the other from a future reference image, without sending additional syntax elements. In JEM, DMVR is not applied when LIC, affine motion, FRUC, or subCU merge candidates are enabled for a CU.
[0224] 3.5 Local Illumination Compensation
[0225] Local illumination compensation (IC) is based on a linear model for illumination changes, using a scaling factor a and an offset b. Furthermore, local illumination compensation is adaptively enabled or disabled for each coding unit (CU) coded for each inter-frame mode.
[0226] When the IC is applied to the CU, the least squares error method is used to derive parameters a and b by using the neighboring samples of the current CU and their corresponding reference samples. More specifically, as... Figure 25 As shown, neighboring samples (2:1 subsampling) of the CU and corresponding samples in the reference image (identified by motion information from the current CU or subCU) are used. IC parameters are derived and applied to each prediction direction.
[0227] When the CU is encoded in Merge mode, the IC flag is copied from the adjacent block in a manner similar to motion information copying in Merge mode; otherwise, the IC flag is signaled to the CU to indicate whether the LIC applies.
[0228] When IC is enabled for an image, an additional CU-level RD check is required to determine whether LIC should be applied to the CU. When IC is enabled for a CU, mean-removed sum of absolute difference (MR-SAD) and mean-removed sum of absolute Hadamard-transformed difference (MR-SATD) are used for integer pixel motion search and fractional pixel motion search, respectively, instead of SAD and SATD.
[0229] To reduce encoding complexity, the following encoding scheme is applied in JEM. When there is no significant change in illumination between the current image and its reference images, IC (Integrated Illumination) is disabled for all images. To identify this situation, the histogram of the current image and each reference image of the current image are calculated at the encoder. If the histogram difference between the current image and each reference image of the current image is less than a given threshold, IC is disabled for the current image; otherwise, IC is enabled for the current image.
[0230] 3.6 Example of Merge / Skip Pattern with Bidirectional Matching Refinement
[0231] First, a Merge candidate list is constructed by inserting motion vectors and reference indices of spatially and temporally adjacent blocks into the candidate list using redundancy checks until the number of available candidates reaches the maximum candidate size of 19. Then, a Merge / Skip mode Merge candidate list is constructed by inserting spatial, temporal, affine, Advanced Temporal MVP (ATMVP), Spatial Temporal MVP (STMVP), and additional candidates (combined and zero candidates) used in HEVC according to a predefined insertion order. Figure 25 In the context of the numbered blocks shown:
[0232] (1) Spatial candidates of block 1-4
[0233] (2) Extrapolated affine candidates of blocks 1-4
[0234] (3) ATMVP
[0235] (4) STMVP
[0236] (5) Virtual Affine Candidates
[0237] (6) Spatial Candidates (Block 5) (Used only when the number of available candidates is less than 6)
[0238] (7) Extrapolation affine candidate (block 5)
[0239] (8) Time candidates (e.g., derived in HEVC)
[0240] (9) Non-nearby space candidates, followed by extrapolation affine candidates (blocks 6 to 49).
[0241] (10) Combined Candidates
[0242] (11) Zero candidate
[0243] It can be noted that, in addition to STMVP and affine, the IC flag is also inherited from the Merge candidate. Moreover, for the first four spatial candidates, bidirectional prediction candidates are inserted before candidates with unidirectional prediction.
[0244] 4. Examples of binarization methods and Merge index encoding
[0245] In some embodiments, several binarization methods can be selected. For a grammar element, the relevant values should first be binarized into a binary string based on the distribution. For each binary bit, it can be encoded using a context-based or side-channel encoding method.
[0246] 4.1 Exemplary Unary and Rounding Unary (TU) Binarization Process
[0247] For each unsigned integer value symbol x ≥ 0, the unary codeword in CABAC consists of x "1" bits followed by a terminating "0" bit. The Truncated Unary (TU) code is defined only for x where 0 ≤ x ≤ S, where for x < S, the code is given by the unary code, and for x = S, the terminating "0" bit is ignored, such that the TU code for x = S is given by a codeword consisting of only x "1" bits.
[0248] Table 5: Binary strings for unary binarization
[0249]
[0250] Table 6: Binary strings for truncated unary binarization
[0251]
[0252] 4.2 Exemplary K-th order Exp-Golomb (EGk) binarization process
[0253] For EGk binarization, the number of symbols with the same code length k + 2·l(x) + 1 grows geometrically. By inverting the Shannon relationship between the ideal code length and the symbol probability, we can, for example, easily infer that EG0 is the optimal code for the probability distribution function (pdf) p(x) = 1 / 2·(x + 1)-2 (where x ≥ 0). This means that for a suitably chosen parameter k, the EGk code represents a reasonably good first-order approximation of the ideal prefix-free code for the tails of the probability distribution functions typically observed.
[0254] Table 7: Binary strings for EG0 binarization
[0255]
[0256] 4.3 Exemplary Truncated Rice (TR) binarization process
[0257] The input to this process is a request for TR binarization, cMax, and cRiceParam.
[0258] The output of this process is the TR binarization that associates each value symbolVal with the corresponding binary string.
[0259] The TR binary string is the concatenation of a prefix binary string and a suffix binary string (when present).
[0260] For deriving the prefix binary string, the following applies:
[0261] -- The prefix value prefixVal of symbolVal is derived as follows:
[0262] prefixVal = symbolVal >> cRiceParam
[0263] -- The prefix of the TR binary string is specified as follows:
[0264] If prefixVal is less than cMax >> cRiceParam, the prefix binary string is a bit string of length prefixVal + 1 indexed by binIdx. The binary bits of binIdx less than prefixVal are equal to 1. The binary bit of binIdx equal to prefixVal is equal to 0.
[0265] When cMax is greater than symbolVal and cRiceParam is greater than 0, there is a suffix of the TR binary string, and it is derived as follows:
[0266] -- The suffix value suffixVal is derived as follows:
[0267] suffixVal = symbolVal - ((prefixVal) << cRiceParam)
[0268] -- The suffix of the TR binary string is specified by calling the fixed-length (FL) binarization process of suffixVal with a cMax value equal to (1 << cRiceParam) - 1.
[0269] Note that for the input parameter cRiceParam = 0, the TR binarization is exactly the truncating unary binarization, and it is always called with a cMax value equal to the maximum possible value of the decoded syntax element.
[0270] 4.4 Exemplary Fixed-Length (FL) Binarization Process
[0271] The input of this process is a request for FL binarization and cMax.
[0272] The output of this process is the FL binarization that associates each value symbolVal with the corresponding binary string.
[0273] The FL binarization is constructed by using the fixedLength-bit unsigned integer binary string of the symbol value symbolVal, where fixedLength = Ceil(Log2(cMax + 1)). The indices of the binary bits of the FL binarization are such that binIdx = 0 is associated with the most significant bit, and the value of binIdx increases towards the least significant bit.
[0274] Table 8: Binary String of FL Binarization (cMax = 7)
[0275]
[0276] 4.5Example encoding of merge_idx
[0277] As specified in the HEVC specification, if the total number of allowed Merge candidates is greater than 1, the Merge index is first binarized into a binary string.
[0278] Table 9: Binarization and Encoding Methods for merge_idx
[0279]
[0280] Use a TR with cRiceParam equal to 0, i.e., TU. The first bit of merge_idx is encoded with a context, and the remaining bits (if present) are encoded with a bypass.
[0281] 5. Example Implementation of Intra-Frame Prediction in JEM
[0282] 5.1 Example of intra-mode coding with 67 intra-prediction modes
[0283] To capture arbitrary edge directions presented in natural video, the number of directional intra-frame modes has been expanded from 33, as used in HEVC, to 65. Additional directional modes are available in... Figure 27 The arrows are depicted as light gray dashed lines, and the planar and DC modes remain unchanged. These denser directional intra-prediction modes are applicable to all block sizes as well as luma and chroma intra-prediction.
[0284] 5.2 Example of Luminosity Intra-Frame Mode Coding
[0285] In JEM, the total number of intra-prediction modes has increased from 35 in HEVC to 67. Figure 27 Examples of 67 intra-frame prediction modes are depicted.
[0286] To accommodate the increased number of directional intra-frame modes, an intra-frame mode coding method with six most probable modes (MPMs) is used. This involves two main technical aspects: 1) the derivation of the six MPMs, and 2) entropy coding of the six MPMs and non-MPM modes.
[0287] In JEM, the patterns included in the MPM list are categorized into three groups:
[0288] --Neighborhood Intra-Frame Mode
[0289] --Derived intra-frame mode
[0290] --Default intra-frame mode
[0291] The MPM list is formed using five adjacent intra-prediction modes. The positions in the five adjacent blocks are the same as those used in the Merge mode, i.e., ... Figure 28 The diagram shows the left (L), top (A), bottom left (BL), top right (AR), and top left (AL). The initial MPM list is formed by inserting five neighboring intra-frame modes, along with the planar and DC modes, into the MPM list. A pruning process is used to remove duplicate modes so that only unique modes can be included in the MPM list. The initial mode inclusion order is: left, top, planar, DC, bottom left, top right, and then top left.
[0292] If the MPM list is not full (i.e., fewer than 6 MPM candidates), a derivation mode is added; these intra-frame modes are obtained by adding -1 or +1 to the angular mode already included in the MPM list. Such additional derivation modes are not generated from non-angular modes (DC or plane).
[0293] Finally, if the MPM list is still not complete, add the default patterns in the following order: vertical, horizontal, pattern 2, and diagonal pattern. As a result of this process, a unique list of 6 MPM patterns is generated.
[0294] For entropy encoding of the selected pattern using 6 MPMs, rounding unary binarization is used. The first three bits are encoded with context, which depends on the MPM pattern associated with the current signaling notification bit. MPM patterns are classified into one of three categories: (a) predominantly horizontal patterns (i.e., the number of MPM patterns is less than or equal to the number of patterns in the diagonal direction), (b) predominantly vertical patterns (i.e., the number of MPM patterns is greater than the number of patterns in the diagonal direction), and (c) non-angular (DC and planar) classes. Therefore, based on this classification, three contexts are used for the signaling notification MPM index.
[0295] The encoding for selecting the remaining 61 non-MPMs is performed as follows. The 61 non-MPMs are first divided into two sets: the selected mode set and the unselected mode set. The selected mode set contains 16 modes, and the remaining modes (45 modes) are assigned to the unselected mode set. The mode set to which the current mode belongs is indicated in a bitstream with a flag. If the mode to be indicated is within the selected mode set, the selected mode is notified using a 4-bit fixed-length code signaling; and if the mode to be indicated comes from the unselected mode set, the selected mode is notified using a rounded binary code signaling. The selected mode set is generated by subsampling the 61 non-MPM modes as follows:
[0296] --Selected pattern set = {0,4,8,12,16,20…60}
[0297] --Unselected pattern set = {1,2,3,5,6,7,9,10…59}
[0298] On the encoder side, a similar two-stage intra-mode decision process to HM is used. In the first stage (i.e., the intra-mode pre-selection stage), N intra-prediction modes are pre-selected from all available intra-modes using the lower-complexity Sum of Absolute Transform Difference (SATD) cost. In the second stage, a higher-complexity RD cost selection is further applied to select one intra-prediction mode from the N candidates. However, when applying 67 intra-prediction modes, the complexity of the intra-mode pre-selection stage increases because the total number of available modes roughly doubles, so directly using the same encoder mode decision process as HM would also increase the complexity. To minimize the increase in encoder complexity, a two-step intra-mode pre-selection process is performed. In the first step, based on the Sum of Absolute Transform Difference (SATD) metric, N intra-prediction modes are pre-selected from (… Figure 27 In the first step, N (N depends on the intra-prediction block size) modes are selected from the original 35 intra-prediction modes (indicated by the black solid arrows); in the second step, the direct neighbors of the selected N modes are further examined using SATD (e.g., ...). Figure 27 The additional intra-prediction directions (indicated by the light gray dashed arrows) are then used to update the list of the selected N modes. Finally, if MPMs are not yet included, the first M MPMs are added to the N modes, and a final list of candidate intra-prediction modes is generated for the second-stage RD cost check, which is done in the same way as HM. Based on the original settings in HM, the value of M is increased by 1, and N is slightly decreased, as shown in Table 10 below.
[0299] Table 10: Number of mode candidates in the intra-frame mode pre-selection step
[0300] Intra-prediction block size 4×4 8×8 16×16 32×32 64×64 >64×64 HM 8 8 3 3 3 3 JEM with 67 intra-frame prediction modes 7 7 2 2 2 2
[0301] 5.3 Example of chroma intra-frame mode coding
[0302] In JEM, a total of 11 intra-frame modes are allowed for chroma CB coding. These modes include 5 traditional intra-frame modes and 6 cross-component linear model modes. The list of chroma mode candidates consists of the following three parts:
[0303] ●CCLM mode
[0304] ●DM mode, an intra-prediction mode derived from the luminance CBs of five juxtaposed positions covering the current chroma block.
[0305] The five positions checked in sequence are: the center (CR), top left (TL), top right (TR), bottom left (BL), and bottom right (BR) 4×4 blocks within the corresponding luma block of the current chroma block of the I stripe. For the P and B stripes, only one of these five sub-blocks is checked because they have the same pattern index. Figure 29A and Figure 29B The image shows an example of five juxtaposed brightness positions.
[0306] ● Based on the chromaticity prediction mode of spatially adjacent blocks:
[0307] ○ Five chroma prediction modes: based on spatially adjacent blocks on the left, top, bottom left, top right, and top left.
[0308] ○ Plane and DC modes
[0309] ○ Add derived patterns by adding -1 or +1 to angle patterns already included in the list to obtain these intra-frame patterns.
[0310] ○ Vertical, Horizontal, Pattern 2
[0311] A pruning process is applied whenever a new chroma intra mode is added to the candidate list. The size of the non-CCLM chroma intra mode candidate list is then trimmed to 5. For mode signaling, a signaling flag is first used to indicate whether to use one of the CCLM modes or one of the traditional chroma intra prediction modes. Several more flags can then be followed to specify the exact chroma prediction mode used for the current chroma CB.
[0312] 6. Examples of existing implementation methods
[0313] Current HEVC designs can better encode motion information by using the correlation between the current block and its neighboring blocks (the blocks immediately adjacent to the current block). However, neighboring blocks may correspond to different objects with different motion trajectories. In this case, prediction based on its neighboring blocks is inefficient.
[0314] Predictions based on motion information from non-neighboring blocks can bring additional coding gains, but there is a cost to storing all motion information (typically at the 4×4 level) in a cache, which significantly increases the complexity of the hardware implementation.
[0315] Univariate binarization is suitable for a small number of allowed merge candidates. However, it may become suboptimal when the total number of allowed candidates becomes larger.
[0316] The HEVC design for constructing the AMVP candidate list only invokes pruning among the two spatial AMVP candidates. Full pruning (of each available candidate compared to all others) is not utilized because the encoding loss due to finite pruning is negligible. However, pruning becomes important if more AMVP candidates are available. Furthermore, when LAMVR is enabled, how to construct the AVMP candidate list should be investigated.
[0317] 7. Example methods for motion vector prediction based on LUTs
[0318] Embodiments of this disclosure overcome the shortcomings of existing implementations, thereby providing higher coding efficiency for video coding. To overcome the shortcomings of existing implementations, various embodiments can implement a LUT-based motion vector prediction technique using one or more tables (e.g., lookup tables) with at least one motion candidate stored to predict block motion information, to provide higher coding efficiency for video coding. A lookup table is an example of a table that can be used to include motion candidates for predicting block motion information, and other implementations are also possible. Each LUT may include one or more motion candidates, each motion candidate associated with corresponding motion information. The motion information of a motion candidate may include prediction direction, reference index / picture image, motion vector, LIC flag, affine flag, motion vector derivation (MVD) precision and / or some or all of the MVD value. The motion information may further include block location information to indicate where the motion information originates.
[0319] LUT-based motion vector prediction, which can enhance existing and future video coding standards based on the disclosed techniques, is illustrated in the following examples describing various implementations. Because LUTs allow encoding / decoding processes to be performed based on historical data (e.g., blocks that have already been processed), LUT-based motion vector prediction can also be referred to as a History-based Motion Vector Prediction (HMVP) method. In LUT-based motion vector prediction methods, one or more tables with motion information from previously encoded blocks are maintained during the encoding / decoding process. These motion candidates stored in the LUT are named HMVP candidates. During the encoding / decoding of a block, the associated motion information in the LUT can be added to the motion candidate list (e.g., a Merge / AMVP candidate list), and after encoding / decoding a block, the LUT can be updated. Subsequent blocks are then encoded using the updated LUT. Therefore, the updating of motion candidates in the LUT is based on the encoding / decoding order of the blocks.
[0320] LUT-based motion vector prediction, which can enhance existing and future video coding standards, is illustrated in the following examples describing various implementations. The examples of the disclosed techniques provided below illustrate general concepts and are not intended to be construed as limiting. In one example, the various features described in these examples may be combined unless explicitly indicated to the contrary.
[0321] Regarding terminology, the following example of an entry for a LUT is a motion candidate. The term motion candidate is used to indicate a set of motion information stored in a lookup table. For conventional AMVP or Merge modes, AMVP or Merge candidates are used to store motion information. As will be described below, and in non-limiting examples, the concept of a LUT with motion candidates for motion vector prediction is extended to a LUT with intra-prediction modes for intra-frame mode coding, or to a LUT with illumination compensation parameters for IC parameter coding, or to a LUT with filter parameters. LUT-based methods for motion candidates can be extended to other types of coding information, such as existing and future video coding standards as described in this patent document.
[0322] Example of a lookup table
[0323] Example A: Each lookup table can contain one or more motion candidates, each of which is associated with its motion information.
[0324] In one example, the table size (e.g., the maximum allowed number of motion candidates) and / or the number of tables may depend on the sequence resolution, the maximum coding unit size, and the size of the Merge candidate list.
[0325] Update the lookup table
[0326] Example B1: After a block is encoded with motion information (i.e., IntraBC mode, inter-frame coding mode), one or more lookup tables can be updated.
[0327] a. In one example, whether to update the lookup table can reuse the rules used to select the lookup table. For example, when a lookup table can be selected to encode the current block, the selected lookup table can be further updated after the block is encoded / decoded.
[0328] b. The lookup table to be updated can be selected based on the encoding information and / or the location of the block / LCU.
[0329] c. If the block is encoded with motion information notified by direct signaling (such as AMVP mode, MMVD mode for normal / affine inter-frame mode, AMVR mode for normal / affine inter-frame mode), the block's motion information can be added to the lookup table.
[0330] i. Alternatively, if a block is encoded with motion information directly inherited from spatially adjacent blocks without any refinement (e.g., no refined spatial merge candidates), then the block's motion information should not be added to the lookup table.
[0331] ii. Alternatively, if the block is encoded with refined motion information (such as DMVR, FRUC) directly inherited from spatially adjacent blocks, the block's motion information should not be added to any lookup table.
[0332] iii. Alternatively, if the block is encoded with motion information inherited directly from motion candidates stored in a lookup table, the block's motion information should not be added to any lookup table.
[0333] iv. In one example, such motion information can be added directly to a lookup table, such as as the last entry in the table or as an entry to store the next available motion candidate.
[0334] v. Alternatively, such motion information can be added directly to the lookup table without pruning, for example, without any pruning.
[0335] vi. Alternatively, such motion information can be used to reorder the lookup table.
[0336] vii. Alternatively, such motion information can be used to update the lookup table with limited pruning (e.g., compared to the latest one in the lookup table).
[0337] d. Select M (M>=1) representative positions within the block, and the motion information associated with the representatives is used to update the lookup table.
[0338] i. In one example, representative positions are defined as the four corner positions within the block (e.g., Figure 30 One of C0-C3 in the series.
[0339] ii. In one example, the representative location is defined as the center location within the block (e.g., Figure 30 (Ca-Cd in the middle).
[0340] iii. When sub-block prediction is not allowed for a block, M is set to 1.
[0341] iv. When sub-block prediction is allowed for a block, M can be set to 1, or the total number of sub-blocks, or any other value between [1, number of sub-blocks] (excluding 1 and the number of sub-blocks).
[0342] v. Alternatively, when sub-block prediction is allowed for a block, M can be set to 1 and the selection of representative sub-blocks can be based on the following:
[0343] 1. The frequency of motion information used.
[0344] 2. Is it a bidirectional prediction block?
[0345] 3. Based on reference image index / reference image
[0346] 4. Motion vector difference compared to other motion vectors (e.g., selecting the maximum MV difference).
[0347] 5. Other encoding information.
[0348] e. When selecting M (M>=1) representative locations to update the lookup table, other conditions can be checked before adding them as additional motion candidates to the lookup table.
[0349] i. Pruning can be applied to a new set of motion information for existing motion candidates in a lookup table.
[0350] ii. In one example, the new set of motion information should not be the same as any or part of the existing motion candidates in the lookup table.
[0351] iii. Alternatively, for a new set of motion information and the same reference image of an existing motion candidate, the MV difference should be no less than one or more thresholds. For example, the horizontal and / or vertical components of the MV difference should be greater than 1 pixel distance.
[0352] iv. Alternatively, when K>L, the new set of motion information is pruned only with the last K candidates or the first K%L existing motion candidates to allow the reactivation of old motion candidates.
[0353] v. Alternatively, trimming should not be applied.
[0354] f. If M sets of motion information are used to update the lookup table, the corresponding counter should be incremented by M.
[0355] g. Assuming that before encoding the current block, K represents the counter of the lookup table to be updated, then after encoding the block, for a selected set of motion information (using the method described above), it is added as an additional motion candidate with an index equal to K%L (where L is the lookup table size). An example is shown in... Figure 31A and Figure 31B As shown in the image.
[0356] i. Alternatively, it can be added as an additional motion candidate with an index equal to min(K+1, L-1). Alternatively, if K>=L, the first motion candidate (index equal to 0) is removed from the lookup table, and the indices of the subsequent K candidates are decreased by 1.
[0357] ii. For the two methods above (adding new motion candidates to entries with an index equal to K%L or adding them to entries with an index equal to min(K+1, L-1)), they attempt to preserve the latest set of motion information from previously encoded blocks, regardless of whether the same / similar motion candidates exist.
[0358] iii. Alternatively, when a new set of motion information is added to the LUT as a motion candidate, a redundancy check is applied first. In this case, the LUT will retain the most recent few sets of motion information from previously encoded blocks; however, redundant sets of motion information can be removed from the LUT. Such a method is called a redundancy removal-based LUT update method.
[0359] 1. If there are redundant motion candidates in the LUT, the counter associated with the LUT may not be incremented or decremented.
[0360] 2. Redundancy checks can be defined as a pruning process during the construction of the Merge candidate list, for example, checking whether the reference image / reference image index is the same and whether the motion vector difference is within range or the same.
[0361] 3. If a redundant motion candidate is found in the LUT, the redundant motion candidate is moved from its current position to the last position of the LUT.
[0362] a. Similarly, if a redundant motion candidate is found in the LUT, it is removed from the LUT. Additionally, all motion candidates inserted into the LUT after the redundant motion candidate are shifted forward to refill the entry for the removed redundant motion candidate. After the shift, the new motion candidate is added to the LUT.
[0363] b. In this case, the counter remains unchanged.
[0364] c. Once a candidate for redundant motion is identified in the LUT, the redundancy check process is terminated.
[0365] 4. It can identify multiple redundant motion candidates. In this case, all redundant motion candidates are removed from the LUT. Additionally, all remaining motion candidates can be moved forward sequentially.
[0366] a. In this case, the counter is reduced (the number of redundant motion candidates is reduced by 1).
[0367] b. Terminate the redundancy check process after identifying maxR redundant motion candidates (maxR is a positive integer variable).
[0368] 5. The redundancy check process can start from the first motion candidate and go to the last motion candidate (i.e., in the order they are added to the LUT and in the order of the decoding process of the block from which the motion information comes).
[0369] 6. Alternatively, when redundant motion candidates exist in the LUT, instead of removing one or more of the redundant motion candidates from the LUT, virtual motion candidates can be derived from the redundant motion candidates and used to replace the redundant motion candidates.
[0370] a. Virtual motion candidates can be derived from redundant motion candidates by adding offsets(multiple) to the horizontal and / or vertical components of one or more motion vectors; or by averaging two motion vectors if they point to the same reference image. Alternatively, virtual motion candidates can be derived from any function that takes motion vectors from a lookup table as input. Exemplary functions are: adding two or more motion vectors together; averaging two or more motion vectors. The motion vectors can be scaled before being input to the function.
[0371] b. Virtual motion candidates can be added to the same positions as redundant motion candidates.
[0372] c. Virtual motion candidates can be added before all other motion candidates (e.g., starting from the smallest entry index, such as zero).
[0373] d. In one example, it is applied only under certain conditions (such as when the current LUT is not full).
[0374] 7. Under certain conditions, the LUT update method based on redundancy removal can be invoked, such as:
[0375] a. The current block is encoded using Merge mode.
[0376] b. The current block is encoded using the AMVP pattern, but at least one component of the MV difference is non-zero;
[0377] c. Whether the current block is encoded using a sub-block-based motion prediction / motion compensation method (e.g., not using affine patterns for encoding).
[0378] d. The current block is encoded using the Merge pattern, and motion information is associated with some type (e.g., from a spatially adjacent block, from the left adjacent block, from a temporal block).
[0379] h. After encoding / decoding a block, one or more lookup tables can be updated by inserting only the set of M motion information to the end of the table (i.e., after all existing candidates).
[0380] i. Alternatively, some existing motion candidates in the table can also be removed.
[0381] 1. In one example, if the table is full after inserting M sets of motion information, the first few entries of the motion candidates can be removed from the table.
[0382] 2. In one example, if the table is full before inserting M sets of motion information, the first few entries of the motion candidates can be removed from the table.
[0383] ii. Alternatively, if the block is encoded using motion candidates from a table, the motion candidates in the table can be reordered so that the selected motion candidate is placed in the last entry of the table.
[0384] i. In one example, before encoding / decoding the block, the lookup table may include HMVP0, HMVP1, HMVP2, ..., HMVP K-1 HMVP K HMVP K+1 ...HMVP L-1 The motion candidates are represented, among which HMVP i This represents the i-th entry in the lookup table. If from HMVP... K If the block is predicted using K in the range [0, L-1, inclusive], then after encoding / decoding the block, the lookup table is reordered as: HMVP0, HMVP1, HMVP2, ..., HMVP K-1 HMVP K HMVP K+1 ...HMVP L-1 HMVP K .
[0385] j. After encoding an intra-frame constrained block, the lookup table can be cleared.
[0386] k. If entries for motion information are added to the lookup table, more entries for motion information can be added to the table through derivation from the motion information. In this case, the counter associated with the lookup table can increase by more than 1.
[0387] i. In one example, the MV of the motion information entries is scaled and placed in a table;
[0388] ii. In one example, the MV of the motion information entry is added to (dx, dy) and placed in the table;
[0389] iii. In one example, calculate the average MV of two or more entries of motion information and put it in a table.
[0390] Example B2: If a block is located at the boundary of an image / strip / piece, then updating the lookup table should never be allowed.
[0391] Example B3: Motion information for the LCU line above can be disabled to encode the current LCU line.
[0392] a. In this case, the number of available motion candidates can be reset to 0 at the beginning of the new strip / slice / LCU line.
[0393] Example B4: At the beginning of encoding a strip / slice with a new time layer index, the number of available motion candidates can be reset to 0.
[0394] Example B5: A lookup table can be continuously updated using a stripe / slice / LCU row / multiple stripes with the same time-level index.
[0395] a. Alternatively, the lookup table can be updated only after encoding / decoding every S (S>=1) CTU / CTB / CU / CB or after encoding / decoding a region (e.g., with a size equal to 8×8 or 16×16).
[0396] b. Alternatively, the lookup table may be updated only after encoding / decoding (e.g., S inter-coded blocks) of every S (S>=1) blocks (e.g., CU / CB) using certain modes. Alternatively, the lookup table may be updated only after encoding / decoding of every S (S>=1) inter-coded blocks (e.g., CU / CB) that are encoded without using sub-block-based motion prediction / motion compensation methods (e.g., without affine and / or ATMVP modes).
[0397] c. Alternatively, the lookup table can be updated only when the top-left coordinate of the encoded / decoded block satisfies some conditions. For example, the lookup table can be updated only when (x & M == 0) && (y & M == 0), where (x, y) are the top-left coordinates of the encoded / decoded block. M is an integer, such as 2, 4, 8, 16, 32, or 64.
[0398] d. Alternatively, a lookup expression can stop updating once it reaches the maximum allowed counter.
[0399] e. In one example, a predefined counter can be used. Alternatively, it can be signaled in areas covering multiple CTUs / CTBs / CUs / PUs, including the Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), strip header, title sequence, Code Tree Unit (CTU), Code Tree Block (CTB), Code Unit (CU), or Prediction Unit (PU).
[0400] Example B6: When encoding blocks using Merge or AMVP mode, the LUT can be updated using motion information associated with the block.
[0401] Example B7: Pruning can be applied before updating the LUT by adding motion candidates obtained from the coded block.
[0402] Example B8: The LUT can be updated periodically.
[0403] Example C: A reordering of motion candidates in the LUT can be applied.
[0404] (a) In one example, after a block has been encoded, new motion candidates can be obtained from that block. These can first be added to the LUT, and then a reordering can be applied. In this case, for subsequent blocks, it will utilize the reordered LUT. For example, a reordering occurs after encoding of a cell (e.g., an LCU, a row of LCUs, multiple LCUs, etc.) has been completed.
[0405] (b) In one example, motion candidates in the LUT are not reordered. However, when encoding blocks, the motion candidates can be reordered first, and then checked and inserted into Merge / AMVP / or other kinds of motion information candidate lists.
[0406] Example D: Similar to the use of LUTs with motion candidates for motion vector prediction, it is proposed that one or more LUTs can be constructed and / or updated to store intra-prediction patterns from previously encoded blocks, and the LUTs can be used to encode / decode intra-coded blocks.
[0407] 8. Additional Examples of LUT-Based Motion Vector Prediction
[0408] A history-based MVP (HMVP) method is proposed, where HMVP candidates are defined as motion information of previously encoded blocks. A table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is cleared when a new stripe is encountered. Whenever an inter-coded block exists, the associated motion information is added to the last entry of the table as a new HMVP candidate. The entire encoding process... Figure 32 Described in the text.
[0409] In one example, setting the table size to L (e.g., L = 16, 6, or 44) indicates that up to L HMVP candidates can be added to the table.
[0410] In one embodiment (corresponding to example B1.gi), if there are more than L HMVP candidates from previously encoded blocks, a first-in-first-out (FIFO) rule is applied so that the table always contains the latest L previously encoded motion candidates. Figure 33 An example is depicted in the proposed method where FIFO rules are applied to remove HMVP candidates and add new HMVP candidates to the table.
[0411] In another embodiment (corresponding to invention B1.g.iii), whenever a new motion candidate is added (such as the current block being inter-coded and non-affine mode), a redundancy check process is first applied to identify whether the same or similar motion candidate exists in the LUT.
[0412] Some examples are described below:
[0413] Figure 34A An example is shown when the LUT is full before adding a new motion candidate.
[0414] Figure 34B An example is shown when the LUT is not full before adding a new motion candidate.
[0415] Figure 34A and Figure 34B Together, an example of a LUT update method based on redundancy removal is shown (where one redundant motion candidate is removed).
[0416] Figure 35A and 35B Example implementations of two cases of the LUT update method based on redundancy removal are shown (where multiple redundant motion candidates (the two candidates in the figure) are removed).
[0417] Figure 35A This shows an example case where the LUT is full before adding a new motion candidate.
[0418] Figure 35BThis shows an example case where the LUT is not full before adding a new motion candidate.
[0419] HMVP candidates can be used during the Merge candidate list construction process. All HMVP candidates, from the last entry to the first entry (or the last K0 HMVPs, e.g., K0 equals 16 or 6), are inserted into the table after the TMVP candidates. Pruning is applied to the HMVP candidates. The Merge candidate list construction process terminates once the total number of available Merge candidates reaches the maximum allowed Merge candidates for signaling notification. Alternatively, motion candidate extraction from the LUT terminates once the total number of added motion candidates reaches a given value.
[0420] Similarly, HMVP candidates can also be used in the AMVP candidate list construction process. The motion vectors of the last K1 HMVP candidates are inserted into the table after the TMVP candidates. The AMVP candidate list is constructed using only HMVP candidates with the same reference image as the AMVP target reference image. Pruning is applied to the HMVP candidates. In one example, K1 is set to 4. Figure 36 An example of the encoding flow for a LUT-based prediction method is depicted. In the context of periodic updates, the update process is completed after the region is decoded. An example of avoiding frequent LUT updates is provided in [the document / section]. Figure 37 Described in the text.
[0421] The above examples can be incorporated into the context of the methods described below (e.g., methods 3810, 3820, 3830, 3840, which can be implemented at the video decoder and / or video encoder).
[0422] Figure 38A A flowchart of an exemplary method for video processing is shown. Method 3810 includes: at step 3812: maintaining tables, wherein each table includes a set of motion candidates, and each motion candidate is associated with corresponding motion information. Method 3810 further includes: at step 3814: performing a conversion between a first video block and a bitstream representation of the video including the first video block based on the tables. Method 3810 further includes: at step 3816: after performing the conversion, updating zero or more tables based on update rules.
[0423] Figure 38BAnother flowchart of an exemplary method for video processing is shown. Method 3820 includes: at step 3822, maintaining tables, wherein each table includes a set of motion candidates, and each motion candidate is associated with corresponding motion information. Method 3820 further includes: at step 3824, performing a conversion between a first video block and a bitstream representation of the video including the first video block based on the tables. Method 3820 further includes: at step 3826, after performing the conversion, updating one or more tables based on one or more video regions in the video until an update termination criterion is met.
[0424] Figure 38C Another flowchart of an exemplary method for video processing is shown. Method 3830 includes: at step 3832, maintaining one or more tables including motion candidates, each motion candidate associated with corresponding motion information. Method 3830 further includes: at step 3834, reordering the motion candidates in at least one of the one or more tables. Method 3830 further includes: at step 3836, performing a conversion between a first video block and a bitstream representation of the video including the first video block using the one or more tables based on the reordered tables.
[0425] Figure 38D Another flowchart of an exemplary method for video processing is shown. Method 3840 includes: at step 3842, maintaining one or more tables including motion candidates, each motion candidate associated with corresponding motion information. Method 3840 further includes: at step 3844, performing a conversion between a first video block and a bitstream representation of the video including the first video block using one or more of the reordered tables. Method 3840 further includes: at step 3846, updating one or more tables by adding additional motion candidates to the tables based on the conversion of the first video block and reordering the motion candidates in the tables.
[0426] 9. Example implementations of the disclosed technology
[0427] Figure 39This is a block diagram of a video processing apparatus 3900. Apparatus 3900 can be used to implement one or more methods described herein. Apparatus 3900 can be implemented in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3900 may include one or more processors 3902, one or more memories 3904, and video processing hardware 3906. The processors (multiple) 3902 can be configured to implement one or more methods described in this document (including, but not limited to, methods 3810-3840). Memory 3904 can be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 3906 can be used to implement some of the techniques described in this document in hardware circuitry.
[0428] In some embodiments, video encoding and decoding methods can be used as described in [the relevant context]. Figures 38A-38D It is implemented by a device on the described hardware platform.
[0429] The additional features and embodiments of the above methods / techniques are described below using a clause-based description format.
[0430] 1. A video processing method, comprising: maintaining tables, wherein each table includes a set of motion candidates and each motion candidate is associated with corresponding motion information; performing a conversion between a first video block and a bitstream representation of a video including the first video block based on the tables; and updating zero or more tables based on update rules after performing the conversion.
[0431] 2. As described in Clause 1, wherein the update rule does not allow updating the table of video blocks located at the boundaries of images, stripes, or segments of a video.
[0432] 3. A video processing method, comprising: maintaining tables, wherein each table includes a set of motion candidates and each motion candidate is associated with corresponding motion information; performing a conversion between a first video block and a bitstream representation of a video including the first video block based on the tables; and updating one or more tables based on one or more video regions in the video after performing the conversion, until an update termination criterion is met.
[0433] 4. The method described in Clause 1 or 3, wherein the table is updated only within a stripe, slice, maximum coding unit (LCU) row, or a stripe with the same time layer index.
[0434] 5. The method according to Clause 1 or 3, wherein the table is updated after a conversion is performed on S video regions or after a conversion is performed on a video region having a certain size, wherein S is an integer.
[0435] 6. The method described in Clause 3, wherein the update termination criterion is satisfied when the counter associated with the table being updated reaches the maximum allowed number.
[0436] 7. The method according to Clause 3, wherein the update termination criterion is satisfied when the counter associated with the table being updated reaches a predetermined value.
[0437] 8. The method according to Clause 7, wherein the predetermined value is signaled in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), strip header, title sequence, coding tree unit (CTU), coding tree block (CTB), coding unit (CU) or prediction unit (PU), or in a video area covering multiple CTUs, multiple CTBs, multiple CUs or multiple PUs.
[0438] 9. The method described in Clause 1 or 3, wherein the update rule updates the table when the top-left coordinate (x, y) of the first video block satisfies a certain condition defined in the update rule.
[0439] 10. The method described in Clause 9, wherein the table is updated when (x & M == 0) && (y & M == 0), where M is 2, 4, 8, 16, 32 or 64.
[0440] 11. The method according to Clause 1 or 3, wherein after performing the transformation on S video blocks, the rule update table is updated, where S is an integer not less than 1.
[0441] 12. The method according to Clause 11, wherein the S video blocks are inter-frame coded blocks.
[0442] 13. The method according to Clause 11, wherein the S video blocks are not encoded using a sub-block-based motion prediction or sub-block-based motion compensation method.
[0443] 14. The method according to Clause 11, wherein the S video blocks are not encoded using affine mode or optional temporal motion vector prediction (ATMVP) mode.
[0444] 15. A video processing method, comprising: maintaining one or more tables including motion candidates, each motion candidate being associated with corresponding motion information; reordering the motion candidates in at least one of the one or more tables; and performing a conversion between a first video block and a bitstream representation of a video including the first video block based on the reordered motion candidates in the at least one table.
[0445] 16. The method described in Clause 15 further includes: updating one or more tables based on the transformation.
[0446] 17. A video processing method, comprising: maintaining one or more tables including motion candidates, each motion candidate being associated with corresponding motion information; performing a conversion between a first video block and a bitstream representation of a video including the first video block using the one or more tables; and updating the one or more tables by adding additional motion candidates to the tables and reordering the motion candidates in the tables based on the conversion of the first video block.
[0447] 18. The method according to Clause 15 or 17 further comprises: performing a conversion between subsequent video blocks of the video and a bitstream representation of the video based on a reordered table.
[0448] 19. The method according to Clause 17, wherein, after performing a conversion on a video unit including a maximum coding unit (LCU), an LCU row, and at least one of a plurality of LCUs, a reordering is performed.
[0449] 20. The method according to any one of clauses 1-19, wherein performing the conversion includes generating a bitstream representation from the first video block.
[0450] 21. The method according to any one of clauses 1-19, wherein performing the conversion includes generating a first video block from a bitstream representation.
[0451] 22. The method according to any one of clauses 1-21, wherein the motion candidate is associated with motion information, wherein the motion information includes at least one of: predicted direction, reference image index, motion vector value, intensity compensation flag, affine flag, motion vector difference accuracy, or motion vector difference value.
[0452] 23. The method according to any one of clauses 1-14 and 16-22, wherein updating one or more tables includes updating one or more tables based on motion information of the first video block after performing the transformation.
[0453] 24. The method according to Clause 23 further includes: performing a conversion between subsequent video blocks of the video and the bitstream representation of the video based on an updated table.
[0454] 25. An apparatus comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method according to any one of claims 1 to 24.
[0455] 26. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for performing the method according to any one of clauses 1 to 24.
[0456] Based on the foregoing, it is understood that specific embodiments of the disclosed technology have been described for illustrative purposes, but various modifications can be made without departing from the scope of the invention. Therefore, the disclosed technology is not limited to any aspect other than the appended claims.
[0457] The subject matter and functional operations described in this patent document can be implemented in various systems, digital electronic circuits, or computer software, firmware, or hardware (including the structures disclosed herein and their equivalents, or combinations thereof). The subject matter described herein can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for operation by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of material affecting machine-readable propagation signals, or combinations thereof. The terms "data processing unit" or "data processing apparatus" encompass all means, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an operating environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof.
[0458] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file storing one or more modules, subroutines, or code sections). Computer programs can be deployed to run on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected through a communications network.
[0459] The processes and logic flows described in this specification can be executed by one or more programmable processors running one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be executed by dedicated logic circuits, and the device can be implemented as dedicated logic circuits, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits).
[0460] Processors suitable for running computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from, transfer data to, or receive data from and transfer data to such mass storage devices. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0461] Instruction manual and accessories Figure 1 The words are intended to be illustrative, where illustrative means example. As used herein, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” are intended to also include the plural forms. Additionally, unless the context clearly indicates otherwise, the use of “or” is intended to include “and / or.”
[0462] While this patent document contains numerous details, these details should not be construed as limiting any invention or potentially claimed scope, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in this patent document in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be excluded from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[0463] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in a sequential manner, or performing all shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0464] Only some implementation methods and examples are described, and other implementations, enhancements and variations can be made based on what is described and shown in this patent document.
Claims
1. A method for encoding and decoding video data, comprising: Maintain one or more tables, wherein each table includes one or more motion candidates derived from one or more video blocks that have been encoded or decoded, and the arrangement of motion candidates in each table is based on the order in which the motion candidates were added to the table; Construct a motion candidate list for the current video block; The motion information of the current video block is determined using the motion candidate list; and Based on the determined motion information, the current video block is encoded and decoded; The determination of motion information for updating one or more tables is based on the encoding / decoding information of the current video block. After the update, the table is used to construct a motion candidate list for subsequent video blocks, and the encoding / decoding information includes the encoding / decoding mode of the current video block.
2. The method of claim 1, wherein, The encoding / decoding information also includes at least one of the position of the current video block or the size of the current video block.
3. The method of claim 1, wherein, If the conditions are met, the determined motion information is used to update one or more tables, wherein the conditions are obtained based on the encoding / decoding information.
4. The method according to claim 1, wherein, If the conditions are not met, the determined motion information will not be used to update the tables in the one or more tables, wherein the conditions are obtained based on the encoding / decoding information.
5. The method according to claim 3 or 4, wherein, The conditions include: The encoding / decoding mode of the current video block belongs to at least one preset mode.
6. The method according to claim 5, wherein, The at least one preset mode includes at least one intra-block copy (IBC) mode or inter-frame encoding / decoding mode.
7. The method according to claim 5, wherein, The at least one preset mode does not include a sub-block-based inter-frame encoding / decoding mode.
8. The method according to claim 7, wherein, The sub-block-based inter-frame coding and decoding modes include affine mode or optional temporal motion vector prediction (ATMVP) mode.
9. The method according to claim 2, wherein, The position of the current video block includes the top-left coordinates (x, y) of the current video block.
10. The method according to claim 1, wherein, If the determined motion information is used to update a table in one or more tables, the table to be updated is selected based on at least one of the following: the codec information, the position of the current video block, and the position of the maximum codec unit (LCU).
11. The method according to any one of claims 1-4, wherein, If the determined motion information is used to update one or more tables, then the table to be updated is the same as the table used to construct the motion candidate list for the current video block.
12. The method according to any one of claims 1-4, wherein, If the determined motion information is used to update one or more tables, the table is updated by adding motion information derived from the current video block to the table.
13. The method according to any one of claims 1-4, wherein, The motion candidates in the table are associated with motion information, which includes at least one of the following: block location information indicating the source of the motion information, prediction direction, reference image index, motion vector value, intensity compensation flag, affine flag, motion vector difference accuracy or motion vector difference, and filter parameters used in the filtering process.
14. The method according to any one of claims 1-4, wherein, The encoding and decoding includes generating a bitstream from the current video block.
15. The method according to any one of claims 1-4, wherein, The encoding and decoding includes generating the current video block from the bitstream.
16. An apparatus for encoding and decoding video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: Maintain one or more tables, wherein each table includes one or more motion candidates derived from one or more video blocks that have been encoded or decoded, and the arrangement of motion candidates in each table is based on the order in which the motion candidates were added to the table; Construct a motion candidate list for the current video block; The motion information of the current video block is determined using the motion candidate list; and Based on the determined motion information, the current video block is encoded and decoded; The determination of motion information for updating one or more tables is based on the encoding / decoding information of the current video block. After the update, the table is used to construct a motion candidate list for subsequent video blocks, and the encoding / decoding information includes the encoding / decoding mode of the current video block.
17. The apparatus according to claim 16, wherein, The encoding / decoding information also includes at least one of the position of the current video block or the size of the current video block.
18. The apparatus according to claim 16 or 17, wherein, If the conditions are met, the determined motion information is used to update one or more tables, wherein the conditions are obtained based on the encoding / decoding information.
19. A non-transitory computer-readable medium storing instructions, said instructions causing a processor to: Maintain one or more tables, where Each table includes one or more motion candidates derived from one or more video blocks that have been encoded or decoded, and the arrangement of motion candidates in each table is based on the order in which the motion candidates were added to the table; Construct a motion candidate list for the current video block; The motion information of the current video block is determined using the motion candidate list; as well as Based on the determined motion information, the current video block is encoded and decoded; The determination of motion information for updating one or more tables is based on the encoding / decoding information of the current video block. After the update, the table is used to construct a motion candidate list for subsequent video blocks, and the encoding / decoding information includes the encoding / decoding mode of the current video block.
20. A method for storing a bitstream of video, comprising: Maintain one or more tables, wherein each table includes one or more motion candidates derived from one or more video blocks that have been encoded or decoded, and the arrangement of motion candidates in each table is based on the order in which the motion candidates were added to the table; Construct a motion candidate list for the current video block; The motion information of the current video block is determined using the motion candidate list; and Based on the determined motion information, a bitstream of the current video block is generated; and The bitstream is stored in a non-transitory computer-readable medium. The determination of motion information for updating one or more tables is based on the encoding / decoding information of the current video block. After the update, the table is used to construct a motion candidate list for subsequent video blocks, and the encoding / decoding information includes the encoding / decoding mode of the current video block.