Motion vector (MV) candidate reordering

TWI935170BActive Publication Date: 2026-08-11QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
TW111131315
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-08-18
Filing Date
2022-08-19
Publication Date
2026-08-11
Estimated Expiration
2042-08-18

AI Technical Summary

Technical Problem

Existing video decoding technologies face challenges in efficiently compressing video data while maintaining high quality, leading to increased communication network and equipment burdens due to the large amount of data required for high-resolution video consumption.

Method used

Implementing a multi-stage self-adjusting reordering of merge candidates (ARMC) technique for motion vector candidate reordering in video decoding, which includes grouping and reordering prediction candidates using various methods to optimize the merge candidate list construction.

Benefits of technology

Enhances video decoding efficiency by reducing bit rate requirements without compromising video quality, thus alleviating the burden on communication networks and processing equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001905105_001
    Figure TWG2TB001905105_001
  • Figure TWG2TB001905105_002
    Figure TWG2TB001905105_002
  • Figure TWG2TB001905105_003
    Figure TWG2TB001905105_003
Patent Text Reader

Abstract

Systems and techniques for decoding video data are provided. In some instances, the process may include: obtaining a first plurality of prediction candidates associated with the video data. The process may also include: determining a first group of prediction candidates, at least in part, by applying a first grouping method to the first plurality of prediction candidates. The process may include: reordering the first group of prediction candidates and selecting a first merged candidate from the reordered first group of prediction candidates. The process may also include: adding the first merged candidate to a candidate list.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to video decoding (e.g., encoding and / or decoding of video data). For example, the various forms of this application relate to motion vector (MV) candidate reordering (e.g., for merged patterns). [Previous Technology]

[0002] Many devices and systems allow video data to be processed and output for consumption. Digital video data typically includes a large amount of data to meet the needs of video consumers and providers. For example, video data consumers expect high-quality, high-fidelity, high-resolution, high-view rate, and other high-speed video. Therefore, the large amount of video data required to meet these needs places a burden on the communication networks and equipment that process and store the video data.

[0003] Various video decoding technologies can be used to compress video data. Video decoding technologies can be performed by an encoder-decoder (called a transcoder) according to one or more video decoding standards and / or formats. For example, video decoding standards and formats include High Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), Moving Picture Experts Group (MPEG) Part 2 Decoding, VP9, ​​Open Media Consortium (AOMedia) Video 1 (AV1), Basic Video Decoding (EVC), etc. Video decoding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), which take advantage of redundancy present in video images or sequences. A key goal of video decoding technology is to compress video data to a form using a lower bit rate while avoiding or minimizing video quality degradation. As evolving video services become available, there is a need for encoding technologies with improved decoding accuracy or efficiency. [Summary of the Invention]

[0004] This document describes systems and techniques for improved video processing, such as video encoding and / or decoding. For example, the system may perform motion vector (MV) candidate reordering (e.g., for merged modes), such as multi-level (e.g., two-level) reordering using merged candidate self-adjusting reordering (ARMC) techniques.

[0005] In one illustrative example, an apparatus for processing video data includes: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain a first plurality of prediction candidates associated with the video data; determine a first group of prediction candidates by at least partially applying a first grouping method to the first plurality of prediction candidates; reorder the first group of prediction candidates; select a first merged candidate from the reordered first group of prediction candidates; and add the first merged candidate to a candidate list.

[0006] In another example, a method for decoding video data includes: obtaining a first plurality of prediction candidates associated with the video data; determining a first group of prediction candidates by at least partly applying a first grouping method to the first plurality of prediction candidates; reordering the first group of prediction candidates; selecting a first merged candidate from the reordered first group of prediction candidates; and adding the first merged candidate to a candidate list.

[0007] In another example, a non-transitory computer-readable medium having instructions stored thereon, which, when executed by at least one processor, cause at least one processor to perform the following operations: obtain a first plurality of prediction candidates associated with video data; determine a first group of prediction candidates by at least partly applying a first grouping method to the first plurality of prediction candidates; reorder the first group of prediction candidates; select a first merged candidate from the reordered first group of prediction candidates; and add the first merged candidate to a candidate list.

[0008] In another example, an apparatus for decoding video data includes: means for obtaining a first plurality of prediction candidates associated with the video data; means for determining a first group of prediction candidates at least in part by applying a first grouping method to the first plurality of prediction candidates; means for reordering the first group of prediction candidates; means for selecting a first merged candidate from the reordered first group of prediction candidates; and means for adding the first merged candidate to a candidate list.

[0009] In some embodiments, the system may be or may be a subset of the following: mobile devices (e.g., mobile phones or so-called "smartphones," tablets, or other types of mobile devices), network-connected wearable devices, extended reality devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), personal computers, laptops, server computers (e.g., video servers or other server devices), televisions, vehicles (or computing devices or systems of vehicles), cameras (e.g., digital cameras, Internet Protocol (IP) cameras, etc.), multi-camera systems, robotic devices or systems, aviation devices or systems, or other devices. In some embodiments, the system includes at least one camera for capturing one or more images or video frames. For example, the system may include a camera (e.g., an RGB camera) or multiple cameras for capturing one or more images and / or one or more videos including video frames. In some embodiments, the system includes a display for displaying one or more images, videos, notifications, or other displayable data. In some embodiments, the system includes a transmitter configured to transmit one or more video frames and / or syntax data to at least one device over a transmission medium. In some embodiments, the system described above may include one or more sensors. In some embodiments, the processor includes a neural processing unit (NPU), a central processing unit (CPU), a graphics processing unit (GPU), or other processing devices or components.

[0010] The foregoing and other features and examples will become more apparent after referring to the following description, requests and drawings.

Implementation Method

[0031] Certain forms and examples of this disclosure are provided below. As will be apparent to those skilled in the art, some of these forms and examples can be applied independently, and some can be applied in combination. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of the examples of this application. However, it will be apparent that the various examples can be practiced without such specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0032] The following description provides only illustrative examples and is not intended to limit the scope, applicability, or configuration of this disclosure. Specifically, the subsequent description of each example will provide those skilled in the art with feasible descriptions for implementing particular examples. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0033] Video decoding equipment implements video compression technology to efficiently encode and decode video data. For example, Universal Video Decoding (VVC) is the latest video decoding standard developed by the Joint Video Experts Team (JVET) of ITU-T and ISO / IEC to achieve substantial compression capabilities exceeding HEVC for a wider range of applications. The VVC specification was finalized in July 2020 and published by both ITU-T and ISO / IEC. The VVC specification defines standard bitstream and picture formats, High-Level Syntax (HLS) and decoding unit-level syntax, as well as parsing and decoding processes. VVC also specifies profile / layer / level (PTL) limits, byte stream formats, hypothetical reference decoders, and supplementary enhancement information (SEI) in its appendices. Recently, JVET has been developing Enhanced Compression Model (ECM) software to enhance compression capabilities beyond VVC. The decoding toolset in ECM software includes all functional blocks in the hybrid video decoding framework (including intra-frame prediction, inter-frame prediction, transform and coefficient decoding, intra-loop filtering, and entropy decoding).

[0034] Video compression techniques may include applying different prediction modes (including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques) to reduce or remove redundancy inherent in a video sequence. A video encoder may divide each image of the original video sequence into rectangular regions, which are referred to as video blocks or decoding units (described in more detail below). Specific prediction modes may be used to encode these video blocks.

[0035] A video block may be divided into one or more smaller blocks in one or more ways. A block may include a decode tree block, a prediction block, a transform block, and / or other suitable blocks. Unless otherwise specified, the reference to "block" generally refers to such a video block (e.g., a decode tree block, a decode block, a prediction block, a transform block, or other suitable block or sub-block, as will be understood by those skilled in the art). Furthermore, each of these blocks may also be interchangeably referred to herein as a "unit" (e.g., a decode tree unit (CTU), a decode unit, a prediction unit (PU), a transform unit (TU), etc.). In some cases, a unit may refer to a decoded logic unit encoded in a bitstream, while a block may refer to a portion of the video frame buffer to which the process is targeted.

[0036] For inter-frame prediction modes, the video encoder can search for blocks similar to the block being encoded at another time location within a frame (or image), referred to as reference frames or reference images. The video encoder can restrict the search to a specific spatial displacement from the block to be encoded. A two-dimensional (2D) motion vector, including horizontal and vertical displacement components, can be used to locate the best match. For intra-frame prediction modes, the video encoder can use spatial prediction techniques to form prediction blocks based on data from previously encoded adjacent blocks within the same image.

[0037] The video encoder can determine the prediction error. For example, the prediction can be determined as the difference between the primitive values ​​in the block being encoded and the predicted block. The prediction error can also be called the residual. The video encoder can also apply a transform to the prediction error using transform decoding (e.g., using the form of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), or other suitable transforms) to produce transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and motion vectors can be represented using syntax elements and, together with control information, form a decoded representation of the video sequence. In some cases, the video encoder can perform entropy decoding on the syntax elements, thereby further reducing the number of bits required for its representation.

[0038] The video decoder can use the syntax elements and control information discussed above to construct prediction data (e.g., prediction blocks) for decoding the current frame. For example, the video decoder can add the prediction block to the compressed prediction error. The video decoder can determine the compressed prediction error by weighting the transform basis function using quantization coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.

[0039] In some cases, video decoding devices (e.g., video encoders, video decoders, or combined encoder-decoders or transcoders) can apply one or more merging patterns to the current block of video data being decoded (e.g., encoded and / or decoded) to inherit information from another block. For example, when applying a merging pattern to the current block, the video decoding device can obtain the same one or more motion vectors, prediction directions, and / or one or more reference image indices from another block of the current block (e.g., another inter-frame prediction PU or other block). Efficient systems and techniques are required to implement the merging patterns.

[0040] As described in more detail below, this document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively, "systems and techniques") for performing motion vector (MV) candidate reordering (such as for merging patterns). In some forms, multi-level (e.g., two-level) merging candidate self-adjusting reordering (ARMC) is provided. For example, in a first ARMC level, a decoding device (e.g., an encoding device, a decoding device, or a combined encoder-decoder device or transcoder) can use a first grouping method to group the predicted candidates. The decoding device can apply reordering within each group (e.g., individually within each group). The decoding device can then construct a first merging candidate list in the order of the groups processed by the first ARMC level. The decoding device can use the first merging candidate list as input to a second ARMC level. In the second ARMC level, the decoding device can use a second grouping method to group the candidates. The decoding device can apply reordering within each group of the candidates generated by the second grouping method (e.g., individually within each group). The decoding device can then construct a second merge candidate list in the order of the groups processed by the second ARMC stage. Further details are provided below.

[0041] The systems and techniques described herein can be applied to one or more of various block-based video decoding technologies, where video is reconstructed on a block-by-block basis. For example, the systems and techniques described herein can be applied to any existing video transcoder (e.g., VVC, HEVC, AVC, or other suitable existing video transcoders), and / or can be efficient decoding tools for any video decoding standard under development and / or future video decoding standards. For example, the instances described herein can be performed using video transcoders such as VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein can also be applied to other decoding standards, transcoders, or formats such as MPEG, JPEG (or other decoding standards for still images), VP9, ​​AV1, their extensions, or other suitable decoding standards that are already available or not yet available or under development. For example, in some instances, the systems and techniques can operate according to proprietary video transcoders / formats (such as AV1, extensions of AVI, and / or subsequent versions of AV1 (e.g., AV2)) or other proprietary formats or industry standards. Therefore, although the techniques and systems described herein may be described with reference to a specific video decoding standard, those skilled in the art will understand that the description should not be construed as applicable only to that specific standard.

[0042] Various aspects of the systems and technologies described herein will be discussed herein with reference to the accompanying drawings.

[0043] FIG1 is a block diagram illustrating an example of a system 100 including an encoding device 104 and a decoding device 112 capable of performing one or more of the techniques described herein. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device. The source device and / or receiving device may include electronic devices such as mobile or landline phones (e.g., smartphones, cellular phones, etc.), desktop computers, laptops or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices. In some instances, the source device and receiving device may include one or more wireless transceivers for wireless communication. The decoding techniques described herein are applicable to video decoding in a variety of multimedia applications, including streaming video transmission (e.g., via the Internet), television broadcasting or transmission, encoding of digital video for storage on data storage media, decoding of digital video stored on data storage media, or other applications. As used herein, the term decoding may refer to encoding and / or decoding. In some instances, System 100 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video replay, video broadcasting, gaming, and / or video telephony.

[0044] Encoding device 104 (or encoder) can be used to encode video data using video decoding standards, formats, transcoders, or protocols to produce an encoded video bitstream. Examples of video decoding standards and formats / transcoders include ITU-T H.261, ISO / IEC MPEG-1 Vision, ITU-T H.262 or ISO / IEC MPEG-2 Vision, ITU-T H.263, ISO / IEC MPEG-4 Vision, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its Scalable Video Decoding (SVC) and Multi-View Video Decoding (MVC) extensions), High Efficiency Video Decoding (HEVC) or ITU-T H.265, and Universal Video Decoding (VVC) or ITU-T H.266. Various extensions of HEVC exist for handling multi-layer video decoding, including range and screen content decoding extensions, 3D video decoding (3D-HEVC), multi-view extensions (MV-HEVC), and scalable extensions (SHVC). The ITU-T Video Decoding Experts Group (VCEG) and the Joint Coordinating Group on Video Decoding of the ISO / IEC Moving Picture Experts Group (MPEG) (JCT-VC), as well as the Joint Coordinating Group on the Development of 3D Video Decoding Extensions (JCT-3V), have developed HEVC and its extensions. VP9, ​​AOMedia Video 1 (AV1) developed by the Open Media Consortium (AOMedia) of Open Media, and Basic Video Decoding (EVC) are the technologies described in this paper that can be applied to other video decoding standards.

[0045] Referring to FIG1, video source 102 can provide video data to encoding device 104. Video source 102 may be part of a source device, or it may be part of a device other than a source device. Video source 102 may include video capturing devices (e.g., cameras, camera phones, video phones, etc.), video archives containing stored video, video servers or content providers providing video data, video feed interfaces receiving video from video servers or content providers, computer graphics systems for generating computer graphics video data, combinations of such sources, or any other suitable video source.

[0046] Video data from video source 102 may include one or more input pictures or frames. A picture or frame is a still image that is part of the video in some cases. In some instances, the data from video source 102 may be a still image that is not part of the video. In HEVC, VVC, and other video decoding specifications, a video sequence may include a series of pictures. A picture may include three sampling arrays (denoted as SL, SCb, and SCr). SL is a two-dimensional array of luminance samples, SCb is a two-dimensional array of Cb chrominance samples, and SCr is a two-dimensional array of Cr chrominance samples. Chrominance sampling may also be referred to herein as "chroma" sampling. A primitive may refer to all three components (luminance and chrominance samples) at a given location in the array of pictures. In other cases, a picture may be monochrome and may consist only of an array of luminance samples; in such cases, the terms primitive and sample may be used interchangeably. The example techniques described herein for illustrative purposes, which refer to various sampling methods, can be applied to primitives (e.g., all three sampled components at a given location in an array of images).

[0047] The encoder engine 106 (or encoder) of the encoding device 104 encodes the video data to produce an encoded video bitstream. In some instances, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more decoded video sequences. The decoded video sequence (CVS) includes a series of access units (AUs) starting from an AU that has a random access point picture in the base layer and has certain attributes, and continuing to the next AU that has a random access point picture in the base layer and has certain attributes, but not including that next AU. For example, certain attributes of the random access point picture that starts the CVS may include a RASL flag equal to 1 (e.g., NoRaslOutputFlag). Otherwise, a random access point picture (with a RASL flag equal to 0) does not start the CVS. The access unit (AU) includes one or more decoded pictures and control information corresponding to the decoded pictures that share the same output time. Decoded slices of an image are encapsulated as data units (called Network Abstraction Layer (NAL) units) at the bitstream level. For example, an HEVC video bitstream may include one or more CVSs, each of which includes NAL units. Each NAL unit has a NAL unit header. In one instance, the header is a byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. The syntax elements in the NAL unit header use specified bits and are therefore visible to all kinds of systems and transport layers, such as transport streams, Real-Time Transport (RTP) protocols, file formats, etc.

[0048] In the HEVC standard, there are two types of NAL units: Video Decoding Layer (VCL) NAL units and non-VCL NAL units. A VCL NAL unit includes a slice or fragment of decoded picture data (described below), and a non-VCL NAL unit includes control information related to one or more decoded pictures. In some cases, a NAL unit may be referred to as a packet. A HEVC AU includes: a VCL NAL unit containing decoded picture data, and a non-VCL NAL unit (if any) corresponding to the decoded picture data.

[0049] The NAL unit may contain a sequence of bits forming a decoded representation of video data (e.g., an encoded video bitstream, a bitstream CVS, etc.), such as a decoded representation of an image in the video. The encoder engine 106 generates a decoded representation of an image by dividing each image into multiple slices. Each slice is independent of other slices, allowing information in that slice to be decoded without relying on data from other slices within the same image. A slice includes one or more segments, which include independent segments and (if present) one or more dependent segments that depend on previous segments. The slice is divided into decoded tree blocks (CTBs) for luma sampling and chroma sampling. The luma-sampled CTB and one or more chroma-sampled CTBs, together with the syntax used for sampling, are called decoded tree units (CTUs). A CTU may also be called a "tree block" or a "maximum decoded unit" (LCU). A CTU is the basic processing unit used for HEVC encoding. A CTU can be separated into multiple decoded units (CUs) of different sizes. The CU contains luminance and chrominance sampling arrays called decoding blocks (CBs).

[0050] Luminance and chrominance CBs can be further separated into prediction blocks (PBs). A PB is a sampled block of the luminance or chrominance component that uses the same motion parameters for inter-frame prediction or intra-block copy prediction (when available or enabled for use). The luminance PB and one or more chrominance PBs, together with their associated syntax, form a prediction unit (PU). For inter-frame prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU and used for inter-frame prediction of the luminance PB and one or more chrominance PBs. Motion parameters may also be referred to as motion information. CBs can also be divided into one or more transform blocks (TBs). A TB represents a square block of sampled color components, to which a residual transform (e.g., in some cases, the same two-dimensional transform) is applied to decode the prediction residual signal. A transform unit (TU) represents the TB of luminance and chrominance sampling and the corresponding syntax elements.

[0051] The size of the CU corresponds to the size of the decoding mode and can be square. For example, the size of the CU can be 8x8 samples, 16x16 samples, 32x32 samples, 64x64 samples, or any other suitable size up to the corresponding CTU size. The phrase "NxN" is used herein to refer to the primitive size of the video block in both the vertical and horizontal dimensions (e.g., 8 primitives x 8 primitives). Primitives in a block can be arranged in rows and columns. In some instances, the block may not have the same number of primitives in the horizontal direction as it does in the vertical direction. The syntax data associated with the CU can describe, for example, dividing the CU into one or more PUs. The segmentation mode can differ between whether the CU is encoded using an in-frame prediction mode or an inter-frame prediction mode. PUs can be segmented into non-square shapes. The syntax data associated with the CU can also describe, for example, dividing the CU into one or more TUs according to the CTU. TUs can be square or non-square shapes.

[0052] According to the HEVC standard, transform units (TUs) can be used to perform transforms. The TU can be different for different CUs. The size of the TU can be set based on the size of the PU within a given CU. The TU can be the same size as the PU or smaller than the PU. In some instances, a quadtree structure called a residual quadtree (RQT) can be used to subdivide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can correspond to the TU. The primitive differences associated with the TU can be transformed to produce transform coefficients. The transform coefficients can be quantized by the encoder engine 106.

[0053] Once the images of the video data are segmented into CUs, the encoder engine 106 uses a prediction mode to predict each PU. The prediction unit or prediction block is subtracted from the original video data to obtain a residual (described below). For each CU, the prediction mode can be signaled within the bitstream using syntax data. The prediction mode can include intra-frame prediction (or intra-image prediction) or inter-frame prediction (or inter-image prediction). Intra-frame prediction utilizes the correlation between samples that are spatially adjacent within the image. For example, using intra-frame prediction, each PU is predicted from adjacent image data in the same image using, for example, DC prediction to find an average value for the PU, planar prediction to adapt a planar surface to the PU, orientation prediction to infer from adjacent data, or any other suitable type of prediction. Inter-frame prediction uses temporal correlation between images to derive motion-compensated predictions for blocks of image samples. For example, using inter-frame prediction, each PU is predicted from image data in one or more reference images (in the output order before or after the current image). For example, a decision can be made at the CU level as to whether to use inter-image prediction or intra-image prediction to decode image regions.

[0054] Encoder engine 106 and decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, the video decoder (such as encoder engine 106 and / or decoder engine 116) segments the image into a plurality of decoder tree units (CTUs) (where the CTB for luminance sampling and one or more CTBs for chrominance sampling, along with the syntax used for sampling, are collectively referred to as CTUs). The video decoder can segment CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure removes the concept of multiple segmentation types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level segmented according to quadtree segmentation and a second level segmented according to binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to decoder units (CUs).

[0055] In the MTT partitioning structure, blocks can be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. A ternary tree partition is a partition where a block is divided into three sub-blocks. In some instances, ternary tree partitioning divides a block into three sub-blocks without partitioning the original block via a center. The partitioning types in MTT (e.g., quadtree, binary tree, and ternary tree) can be symmetric or asymmetric.

[0056] When operating according to the AV1 transcoder, the encoding device 104 and the decoding device 112 can be configured to decode video data in blocks. In AV1, the largest decoded block that can be processed is called a superblock. In AV1, a superblock can be a 128x128 luminance sample or a 64x64 luminance sample. However, in subsequent video decoding formats (e.g., AV2), the superblock can be defined by different (e.g., larger) luminance sample sizes. In some instances, the superblock is the highest level of the block quadtree. The encoding device 104 can further divide the superblock into smaller decoded blocks. The encoding device 104 can use square or non-square partitioning to divide the superblock and other decoded blocks into smaller blocks. Non-square blocks can include N / 2xN, NxN / 2, N / 4xN, and NxN / 4 blocks. The encoding device 104 and the decoding device 112 can perform separate prediction and transformation processes for each decoded block in the decoded block.

[0057] AV1 also defines tiles for video data. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, the encoding device 104 and the decoding device 112 can encode and decode the decoding blocks within a tile separately, without using video data from other tiles. However, the encoding device 104 and the decoding device 112 can perform filtering across tile boundaries. The size of the tiles can be uniform or non-uniform. Tile-based decoding enables parallel processing and / or multithreading for encoder and decoder implementations.

[0058] In some instances, the encoding device 104 and the decoding device 112 may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other instances, the video decoder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luminance component and another QTBT or MTT structure for the two chrominance components (or two QTBT and / or MTT structures for the respective chrominance components).

[0059] The encoding device 104 and the decoding device 112 can be configured to use a quadtree segmentation, QTBT segmentation, MTT segmentation, or other segmentation structure according to HEVC.

[0060] In some instances, one or more slices of an image are assigned slice types. Slice types include I-slices, P-slices, and B-slices. An I-slice (intra-frame, independently decodable) is a slice of an image that is decoded solely by intra-frame prediction and is therefore independently decodable because an I-slice only requires data within the frame to predict any prediction unit or prediction block of the slice. A P-slice (one-way prediction frame) is a slice of an image that can be decoded using both intra-frame prediction and one-way inter-frame prediction. Each prediction unit or prediction block within a P-slice is decoded using either intra-frame prediction or inter-frame prediction. When inter-frame prediction is applied, the prediction unit or prediction block is predicted using only one reference image, and therefore the reference sample comes from only one reference region of a frame. A B-slice (two-way prediction frame) is a slice of an image that can be decoded using both intra-frame prediction and inter-frame prediction (e.g., double prediction or single prediction). Bidirectional prediction of prediction units or blocks in a B-slice can be performed from two reference images, where each image contributes a reference region, and the sample sets of the two reference regions are weighted (e.g., using equal weights or different weights) to generate the prediction signal for the bidirectional prediction block. As explained above, a slice of an image is decoded independently. In some cases, an image may be decoded as a single slice.

[0061] As mentioned above, intra-frame prediction utilizes the correlation between spatially adjacent samples within the image. There are multiple intra-frame prediction modes (also referred to as "intra-frame patterns"). In some instances, intra-frame prediction for a luminance block includes 35 modes, including a planar mode, a DC mode, and 33 angular modes (e.g., a diagonal intra-frame prediction mode and angular modes adjacent to the diagonal intra-frame prediction mode). Table 1 below indexes the 35 intra-frame prediction modes. In other instances, more intra-frame patterns can be defined, including prediction angles that may not yet be represented by the 33 angular modes. In other instances, the prediction angle associated with an angular mode may differ from the prediction angles used in HEVC. Table 3 Specification of Intra-Frame Prediction Modes and Associated Names In-frame prediction mode Associated names 0 INTRA_PLANAR 1 INTRA_DC 2..34 INTRA_ANGULAR2..INTRA_ANGULAR34

[0062] Inter-image prediction utilizes the temporal correlation between images to derive motion compensation predictions for blocks sampled from images. Using a translational motion model, the position of a block in a previously decoded image (reference image) is indicated by a motion vector (), which specifies the horizontal displacement of the reference block relative to the current block's position, and the vertical displacement of the reference block relative to the current block's position. In some cases, the motion vector () can be integer sampling precision (also known as integer precision), in which case the motion vector points to the integer primitive grid (or integer primitive sampling grid) of the reference frame. In other cases, the motion vector () can have fractional sampling precision (also known as fractional primitive precision or non-integer precision) to capture the movement of the underlying object more accurately, without being limited to the integer primitive grid of the reference frame. The precision of the motion vector can be expressed by the quantization level of the motion vector. For example, the quantization level can be integer precision (e.g., 1 primitive) or fractional primitive precision (e.g., ¼ primitive, ½ primitive, or other sub-primitive values). When the corresponding motion vector has fractional sampling accuracy, interpolation is applied to the reference image to derive the predicted signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate values ​​at fractional positions. The previously decoded reference image is indicated by a reference index (refIdx) in a list of reference images. The motion vector and reference index can be referred to as motion parameters. Two types of inter-image prediction (including single prediction and double prediction) can be performed.

[0063] In the case of using dual prediction for inter-frame prediction (also known as bidirectional inter-frame prediction), two sets of motion parameters (x, y) are used to generate two motion-compensated predictions (from the same reference image or possibly from different reference images). For example, in the case of dual prediction, each prediction block uses two motion-compensated prediction signals and generates B prediction units. The two motion-compensated predictions are combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another instance, weighted prediction can be used, in which different weights can be applied to each motion-compensated prediction. The reference images that can be used in dual prediction are stored in two separate lists (represented as list 0 and list 1). The motion parameters can be derived at the encoder using a motion estimation process.

[0064] In the case of using single prediction for inter-frame prediction (also known as one-way inter-frame prediction), a set of motion parameters is used to generate motion-compensated predictions from a reference image. For example, in the case of single prediction, at most one motion-compensated prediction signal is used per prediction block, and P prediction units are generated.

[0065] The PU may include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when the PU is encoded using intra-frame prediction, the PU may include data describing the intra-frame prediction mode used for the PU. As another example, when the PU is encoded using inter-frame prediction, the PU may include data defining the motion vectors used for the PU. The data defining the motion vectors used for the PU may describe, for example, the horizontal component of the motion vector, the vertical component of the motion vector, the resolution used for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference picture to which the motion vector points, the reference index, a list of reference pictures used for the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.

[0066] AV1 includes two common techniques for encoding and decoding decoding blocks of video data. These two common techniques are intra-frame prediction (e.g., intra-frame prediction or spatial prediction) and inter-frame prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when using an intra-frame prediction mode to predict blocks of the current frame of video data, the encoding device 104 and the decoding device 112 do not use video data from other frames of the video data. For most intra-frame prediction modes, the video encoding device 104 encodes the blocks of the current frame based on the difference between the sampled value in the current block and the predicted value generated from a reference sample in the same frame. The video encoding device 104 determines the predicted value generated from the reference sample based on the intra-frame prediction mode.

[0067] After performing prediction using intra-frame prediction and / or inter-frame prediction, the encoding device 104 can perform transformation and quantization. For example, following prediction, the encoder engine 106 can calculate a residual value corresponding to the PU. The residual value can include the primitive difference between the current block (PU) of the primitive being decoded and the prediction block used to predict the current block (e.g., a prediction version of the current block). For example, after generating a prediction block (e.g., issuing inter-frame prediction or intra-frame prediction), the encoder engine 106 can generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of primitive differences that quantize the difference between the primitive values ​​of the current block and the primitive values ​​of the prediction block. In some instances, the residual block can be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of primitive values). In such instances, the residual block is a two-dimensional representation of the primitive values.

[0068] Block transforms are used to transform any residual data that may remain after prediction is performed. These block transforms can be based on discrete cosine transform, discrete sine transform, integer transform, wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., sizes of 32x32, 16x16, 8x8, 4x4, or other suitable sizes) can be applied to the residual data in each CU. In some instances, TUs can be used for transform and quantization processes implemented by encoder engine 106. A given CU with one or more PUs can also include one or more TUs. As described further in detail below, residual values ​​can be transformed into transform coefficients using block transforms, and TUs can be used for quantization and scanning to produce serialized transform coefficients for entropy decoding.

[0069] In some instances, following intra-frame prediction decoding or inter-frame prediction decoding using the PU of the CU, the encoder engine 106 can calculate residual data for the TU of the CU. The PU may include primitive data in the spatial domain (or primitive domain). The TU may include coefficients in the transform domain following the application of block transform. As mentioned above, the residual data may correspond to the primitive difference between the primitives of the uncoded image and the predicted value corresponding to the PU. The encoder engine 106 can form a TU including the residual data for the CU, and can transform the TU to produce transform coefficients for the CU.

[0070] The encoder engine 106 can perform quantization of the transform coefficients. Quantization provides further compression by reducing the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one instance, a coefficient with an n-bit value can be rounded down to an m-bit value during quantization, where n is greater than m.

[0071] Once quantization is performed, the decoded video bitstream includes quantized transform coefficients, prediction information (e.g., prediction patterns, motion vectors, block vectors, etc.), segmentation information, and any other suitable data (such as other grammatical data). The different elements of the decoded video bitstream can be entropy-encoded by encoder engine 106. In some instances, encoder engine 106 can scan the quantized transform coefficients using a predefined scanning order to produce a serialized vector that can be entropy-encoded. In some instances, encoder engine 106 can perform a self-adjusting scan. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), encoder engine 106 can entropy-encode that vector. For example, encoder engine 106 can use context-adjusting variable-length decoding, context-adjusting binary arithmetic decoding, grammar-based context-adjusting binary arithmetic decoding, probabilistic interval segmentation entropy decoding, or another suitable entropy coding technique.

[0072] The output 110 of the encoding device 104 can transmit NAL units constituting the encoded video bitstream data to the decoding device 112 of the receiving device over the communication link 120. The input 114 of the decoding device 112 can receive the NAL units. The communication link 120 may include a channel provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network may include any wireless interface or combination of wireless interfaces and may include any suitable wireless network (e.g., the Internet or other wide area networks, packet-based networks, WiFi™, radio frequency (RF), UWB, WiFi Direct, cellular, Long Term Evolution (LTE), WiMax™, etc.). The wired network may include any wired interface (e.g., fiber optic, Ethernet, powerline Ethernet, coaxial cable Ethernet, digital signal line (DSL), etc.). Various devices (such as base stations, routers, access points, bridges, gateways, switches, etc.) can be used to implement wired and / or wireless networks. Encoded video bitstream data can be modulated according to communication standards such as wireless communication protocols, and the encoded video bitstream data can be sent to the receiving device.

[0073] In some instances, encoding device 104 may store encoded video bitstream data in storage device 108. Output 110 may retrieve encoded video bitstream data from encoder engine 106 or from storage device 108. Storage device 108 may include any of a variety of distributed or local access data storage media. For example, storage device 108 may include hard disks, optical disks, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. Storage device 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. In other instances, storage device 108 may correspond to a file server or another intermediate storage device that may store encoded video generated by a source device. In this case, receiving device including decoding device 112 may access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing and transmitting encoded video data to a receiving device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device can access the encoded video data via any standard data connection (including an internet connection) and can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. Transmission of the encoded video data from storage 108 can be streaming, downloading, or a combination thereof.

[0074] The input 114 of the decoding device 112 receives encoded video bitstream data and can provide the video bitstream data to the decoder engine 116 or to the storage 118 for later use by the decoder engine 116. For example, the storage 118 may include a DPB for storing reference pictures used in inter-frame prediction. A receiving device including the decoding device 112 can receive the encoded video data to be decoded via the storage 108. The encoded video data can be modulated according to a communication standard such as a wireless communication protocol and transmitted to the receiving device. The communication medium used to transmit the encoded video data can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, a switch, a base station, or any other means that can be used to facilitate communication from the source device to the receiving device.

[0075] Decoder engine 116 can decode encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extraction of elements constituting one or more decoded video sequences of encoded video data. Decoder engine 116 can rescale and perform an inverse transform on the encoded video bitstream data. Residual data is passed to the prediction stage of decoder engine 116. Decoder engine 116 predicts blocks of primitives (e.g., PUs). In some instances, the prediction is added to the output of the inverse transform (residual data).

[0076] Decoding device 112 can output the decoded video to video destination device 122, which may include a display or other output device for displaying the decoded video data to a content consumer. In some embodiments, video destination device 122 may be part of a receiving device that includes decoding device 112. In some embodiments, video destination device 122 may be part of a separate device different from the receiving device.

[0077] In some instances, the video encoding device 104 and / or the video decoding device 112 may be integrated with the audio encoding device and the audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 may also include other hardware or software necessary for implementing the decoding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), individual logic, software, hardware, firmware, or any combination thereof. The video encoding device 104 and the video decoding device 112 may be integrated as part of a combined encoder / decoder (transcoder) in the respective device. An example of specific details of the encoding device 104 is described below with reference to Figure 17. An example of specific details of the decoding device 112 is described below with reference to Figure 18.

[0078] The example system illustrated in Figure 1 is an illustrative example that can be used herein. The techniques used to process video data using the techniques described herein can be implemented by any digital video encoding and / or decoding device. Although the techniques disclosed herein are typically implemented by video encoding or video decoding devices, such techniques can also be implemented by a combined video encoder-decoder, commonly referred to as a "CODEC". Furthermore, the techniques disclosed herein can also be implemented by a video preprocessor. The source device and receiving device are merely examples of such decoding devices, where the source device generates decoded video data for transmission to the receiving device. In some instances, the source device and receiving device can operate in a substantially symmetrical manner, such that each of these devices includes video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video replay, video broadcasting, or video telephony.

[0079] Extensions to the HEVC standard include the Multi-View Video Decoding Extension (MV-HEVC) and the Scalable Video Decoding Extension (SHVC). MV-HEVC and SHVC extensions share the concept of layered decoding, where different layers are included in the encoded video bitstream. Each layer in the decoded video sequence is addressed by a unique layer identifier (ID). The layer ID can be present in the header of the NAL unit to identify the layer associated with the NAL unit. In MV-HEVC, different layers can represent different views of the same scene in the video bitstream. In SHVC, different scalable layers are provided to represent the video bitstream at different spatial resolutions (or image resolutions) or different reconstruction fidelities. A scalable layer can include a base layer (where layer ID = 0) and one or more enhancement layers (where layer IDs = 1, 2 to n). The base layer can conform to the first version of the HEVC profile and represents the lowest available layer in the bitstream. Compared to the base layer, enhancement layers offer increased spatial resolution, temporal resolution, or playback rate and / or reconstruction fidelity (or quality). Enhancement layers are organized hierarchically and may or may not depend on lower layers. In some instances, a single-standard transcoder can be used to decode different layers (e.g., using HEVC, SHVC, or other decoding standards to encode all layers). In other instances, a multi-standard transcoder can be used to decode different layers. For example, AVC can be used to decode the base layer, while the MV-HEVC extension of the SHVC and / or HEVC standards can be used to decode one or more enhancement layers.

[0080] Typically, a layer includes a set of VCL NAL units and a corresponding set of non-VCL NAL units. Specific layer ID values ​​are assigned to NAL units. Layers can be hierarchical in the sense that they can depend on lower layers. A layer set refers to a self-contained set of layers represented within a bitstream, meaning that layers within a layer set can depend on other layers in that layer set during decoding, but not on any other layers for decoding. Therefore, each layer in a layer set can form an independent bitstream that can represent video content. The layer set within a layer set can be obtained from another bitstream through sub-bitstream extraction operations. When the decoder wants to operate based on certain parameters, the layer set can correspond to the layer set to be decoded.

[0081] As previously described, the HEVC bitstream includes a set of NAL units (including VCL NAL units and non-VCL NAL units). VCL NAL units include decoded picture data that forms the decoded video bitstream. For example, a bit sequence forming the decoded video bitstream exists within a VCL NAL unit. Non-VCL NAL units may contain parameter sets with high-level information relating to the encoded video bitstream, among other information. For example, parameter sets may include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). Examples of objectives for parameter sets include bit rate efficiency, error resilience, and providing a system-level interface. Each slice references a single active PPS, SPS, and VPS to access information that the decoding device 112 can use to decode the slice. Identifiers (IDs) (including VPS ID, SPS ID, and PPS ID) can be decoded for each parameter set. An SPS includes an SPS ID and a VPS ID. The PPS includes a PPS ID and an SPS ID. Each slice header includes a PPS ID. Using these IDs, the set of active parameters can be identified for a given slice.

[0082] The PPS includes information applicable to all slices in a given image. In some instances, all slices in an image refer to the same PPS. Slices in different images may also refer to the same PPS. The SPS includes information applicable to all images in the same decoded video sequence (CVS) or bitstream. As previously described, a decoded video sequence is a series of access units (AUs) that begin with a random access point image in the base layer (e.g., an Instant Decoding Reference (IDR) image or a Broken Link Access (BLA) image or other suitable random access point image) and have certain properties (described above), until the next AU (or the end of the bitstream) has a random access point image in the base layer and has certain properties, and does not include that next AU. The information in the SPS may not change between images within the decoded video sequence. Images in a decoded video sequence may use the same SPS. The VPS includes information applicable to all layers within the decoded video sequence or bitstream. A VPS includes a syntax structure with syntax elements applicable to the entire decoded video sequence. In some instances, a VPS, SPS, or PPS can be transmitted in-band along with the encoded bitstream. In some instances, a VPS, SPS, or PPS can be transmitted out-of-band in a separate transmission, distinct from the NAL unit containing the decoded video data.

[0083] In general, this disclosure may involve "signaling" certain information (such as syntax elements). The term "signaling" can generally refer to the transmission of values ​​for syntax elements and / or other data for decoding encoded video data. For example, video encoding device 104 may signal values ​​for syntax elements in a bitstream. Generally, signaling refers to generating values ​​in a bitstream. As described above, video source 102 may transmit the bitstream to video destination device 122 substantially immediately or not immediately (such as when syntax elements are stored in memory 108 for later retrieval by video destination device 122).

[0084] The video bitstream may also include supplemental enhancement information (SEI) messages. For example, an SEI NAL unit may be part of the video bitstream. In some cases, the SEI message may contain information that is not required for the decoding process. For example, the information in the SEI message may not be necessary for the decoder to decode the video images in the bitstream, but the decoder can use the information to improve the display or processing of the images (e.g., the decoded output). The information in the SEI message may be embedded relay data. In an illustrative example, the decoder-side entity may use the information in the SEI message to improve the visibility of the content. In some cases, certain application standards may mandate the presence of such SEI messages in the bitstream so that the quality improvement can be brought to all devices that comply with the application standard (e.g., among many other examples, frame-compatible planar stereoscopic 3DTV video formats carry frame-encapsulated SEI messages, where an SEI message is carried for each frame of the video, processing restore point SEI messages, and using generalized scanning to scan rectangular SEI messages in DVB).

[0085] As described above, for each block, a set of motion information (also referred to herein as motion parameters) may be available. The set of motion information may contain motion information for the forward prediction direction and the backward prediction direction. Here, the forward prediction direction and the backward prediction direction are two prediction directions in a bidirectional prediction mode, and the terms "forward" and "backward" do not necessarily have a geometric meaning. Instead, forward and backward may correspond to a reference image list 0 (RefPicList0) and a reference image list 1 (RefPicList1) for the current image, slice, or block. In some instances, when only one reference image list is available for an image, slice, or block, only RefPicList0 is available, and the motion information for each block of the slice is always forward. In some instances, RefPicList0 includes reference images that are temporally preceding the current image, and RefPicList1 includes reference images that are temporally following the current image. In some cases, motion vectors and associated reference indices may be used during decoding. Such motion vectors with associated reference indices are represented as a unidirectional prediction set of motion information.

[0086] For each prediction direction, motion information may include a reference index and a motion vector. In some cases, for simplicity, the motion vector may have associated information, based on which it can be assumed that the motion vector has an associated reference index. The reference index may be used to identify a reference image in the current list of reference images (RefPicList0 or RefPicList1). The motion vector may have horizontal and vertical components, which provide the offset from the coordinate position in the current image to the coordinates in the reference image identified by the reference index. For example, the reference index may indicate a specific reference image that should be used for a block in the current image, and the motion vector may indicate where the best-matching block in the reference image (the block that best matches the current block) is located in the reference image.

[0087] Picture order counts (POCs) can be used in video decoding standards to identify the display order of pictures. Although it is possible for two pictures within a decoded video sequence to have the same POC value, it is generally not the case that two pictures with the same POC value occur within a decoded video sequence. When there are multiple decoded video sequences in a bitstream, pictures with the same POC value may be closer to each other in terms of decoding order. The POC value of a picture can be used for reference picture list construction, such as reference picture set derivation in HEVC, and / or motion vector scaling, etc.

[0088] In some instances, encoding device 104 and / or decoding device 112 may utilize a merging mode that allows a block (e.g., an inter-frame prediction PU or other block) to inherit the same one or more motion vectors, prediction directions, and / or one or more reference picture indices from another block (e.g., another inter-frame prediction PU or other block). In ECM, a general merging candidate list is constructed by sequentially including the following six types of candidates:

[0089] Spatial MVP (SMVP) from Spatial Neighbor CUs: Up to four merge candidates can be selected from the candidates located at the positions depicted in Figure 2. The derivation order is B0, A0, B1, A1, and B2. In some cases, position B2 is only considered if one or more of the CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because it belongs to another slice or tile) or if it is decoded within a frame.

[0090] Temporary MVP from Co-located CU (TMVP): In some cases, only one TMVP candidate is added to the list. In the derivation of this TMVP candidate, a scaled motion vector is derived based on the co-located CU belonging to the co-located reference image. The list of reference images to be used for deriving the co-located CU is explicitly signaled in the slice header. The scaled motion vector for the TMVP candidate is obtained, as shown by the dashed line in Figure 3, which is scaled from the motion vector of the co-located CU using POC distances tb and td, where tb is defined as the POC difference between the reference image of the current image and the current image, and td is defined as the POC difference between the reference image of the co-located image and the co-located image. That is, the MV of currCU (referring to the current CU) is equal to the MV of col_CU multiplied by tb / td. The reference image index of the TMVP candidate is set to zero. The position of the temporal candidate is selected between candidates C0 and C1, as depicted in Figure 4. If the CU at position C0 is unavailable, decoded within the frame, or outside the current line of the CTU, then position C1 is used. Otherwise, position C0 is used when deriving the TMVP candidate.

[0091] Non-adjacent spatial MVPs (NA-SMVPs) from spatially non-adjacent neighboring points CU: Non-adjacent spatial merge candidates (e.g., those described in JVET-L0399, which are incorporated herein by reference in their entirety and for all purposes) are inserted after the TMVP in the general merge candidate list. Figure 5 illustrates an example of a spatial merge candidate pattern, where blocks 1 to 5 are used for SMVPs and blocks 6 to 23 are used for NA-SMVPs. The distance between a non-adjacent spatial candidate and the current decoded block is based on the width and height of the current decoded block.

[0092] History-based MVP (HMVP) from the FIFO table: Motion information of previously decoded blocks is stored in the table and used as the HMVP of the current CU. A table with multiple HMVP candidates is maintained during encoding and / or decoding. Whenever a non-sub-block inter-frame decoding CU exists, the associated motion information is added as the last entry of the table as a new HMVP candidate. The HMVP table size is set to 6, indicating that up to 6 HMVP candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is used, where redundancy checks are first applied to check if the same HMVP exists in the table.

[0093] Pairwise Average MVP (PA-MVP): Pairwise average candidates are generated by averaging predefined pairs of candidates in the existing merge candidate list. The predefined pairs are defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the merge index of the merge candidate list. The average motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged, even if they point to different reference images; if only one motion vector is available, that motion vector is used directly; if no motion vector is available, the list remains invalid.

[0094] Zero MV: When the merge list is not full after adding pairwise average merge candidates, insert zero MVP at the end until the maximum number of merge candidates is encountered.

[0095] In some cases, the encoding device 104 and / or the decoding device 112 can construct a template matching (TM) merge candidate list. For example, in ECM, the TM merge candidate list is constructed based on six types of candidates in the same order as those used in the general merge candidate list described above. TM is a decoder-side MV derivation method used to refine the MV information of each candidate in the TM merge candidate list by finding the closest match between the template in the current image and a block having the same size as the template in the reference image. TM can work with block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check. When both BM and TM are enabled for CU, the TM search process stops at half primitive MVD precision and the resulting MV is further refined by using the same model-based MVD derivation method as in DMVR.

[0096] In some cases, the encoding device 104 and / or the decoding device 112 can perform sub-block-based temporal motion vector prediction (SbTMVP). For example, SbTMVP can be used to predict the motion vectors of sub-CUs within the current CU in two steps, as shown in Figures 6A and 6B. For example, in the first step, spatial neighbor A1 in Figure 6A is examined. If A1 has a motion vector that uses a co-location image as its reference image, then that motion vector is selected as the motion shift to be applied. If no such motion is identified, the motion shift is set to (0, 0). In the second step, the motion shift identified in step 1 is added to the coordinates of the current block to obtain sub-CU-level motion information (motion vector and reference index) from the co-location image, as shown in Figure 6B. The example in Figure 6B assumes that the motion shift is set to the motion of block A1. After identifying the motion information of the co-located sub-CU, this motion information is converted into motion vectors and reference indices in the current sub-CU in a manner similar to the TMVP process in VVC, where temporal motion scaling is applied to align the reference image of the temporal motion vector with the reference image of the current CU. The SbTMVP predictor is added as the first entry in the list of sub-block-based merge candidates, followed by affine merge candidates. SbTMVP and TMVP differ in at least the following ways:

[0097] TMVP predicts motion at the CU level, but SbTMVP predicts motion at the subCU level;

[0098] Although TMVP extracts temporal motion vectors from co-location blocks in the co-location image (the co-location block is the lower right or center block relative to the current CU), SbTMVP applies motion shift before extracting temporal motion information from the co-location image, wherein the motion shift is obtained from the motion vector of one of the spatially adjacent blocks from the current CU.

[0099] In some instances, the encoding device 104 and / or the decoding device 112 can construct a sub-block merging candidate list. For example, the following four types of sub-block merging candidates can be used to construct the sub-block merging candidate list (e.g., the first entry in the sub-block merging candidate list is SbTMVP, and the other entries are affine merging candidates):

[0100] SbTMVP: Add an SbTMVP predictor as described above (e.g., add an SbTMVP predictor as the first entry in the list of sub-block-based merge candidates, followed by affine merge candidates).

[0101] Inherited Affine Merging Candidates Extrapolated from CPMV of Neighboring CUs (I-AffineMVP): There are at most two inherited affine candidates, derived from the affine motion models of adjacent blocks, one from the left-side neighboring CU and one from the upper-side neighboring CU. Candidate blocks are illustrated in Figure 2. For the left predictor, the scanning order is A0->A1, and for the upper predictor, the scanning order is B0->B1->B2. Only the first inherited candidate from each side is selected. No pruning check is performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMV of the affine merging candidate for the current CU.

[0102] Construction of Affine Merging Candidates (C-AffineMVP) using Translation MV Derivation of Neighboring Points CU: Constructing affine candidates means constructing candidates by combining the translation motion information of each control point's neighboring points. The motion information used for the control points is derived from the specified spatial and temporal neighbors shown in Figure 7. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, the B2->B3->A2 block is checked, and the MV of the first available block is used. For CPMV2, the B1->B0 block is checked, and for CPMV3, the A1->A0 block is checked. If the TMVP is available, it is used as CPMV4. The following combinations of control points MV are used to construct them sequentially: {CPMV1,CPMV2,CPMV3}, {CPMV1,CPMV2,CPMV4}, {CPMV1,CPMV3,CPMV4}, {CPMV2,CPMV3,CPMV4}, {CPMV1,CPMV2}, {CPMV1,CPMV3}. Combinations of three CPMVs construct a 6-parameter affine merge candidate, and combinations of two CPMVs construct a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control points MV are discarded if their reference indices differ.

[0103] ZeroMV: After checking the inheritance affine merge candidates and the construction affine merge candidates, if the list is still not full, insert zeroMV at the end of the list.

[0104] In some cases, the encoding device 104 and / or the decoding device 112 may perform merge candidate self-adjusting reordering (ARMC) (referred to as ECM ARMC). For example, in ECM, merge candidates are reordered self-adjustingly using TM. The reordering method can be applied to the general merge candidate list, the TM merge candidate list, and / or the affine merge candidate list (excluding the sub-block merge candidate list of SbTMVP candidates). For TM merge mode, the merge candidates are reordered before the TM refinement process.

[0105] After constructing the merge candidate list, the merge candidates are divided into several subgroups. For the general merge mode and the TM merge mode, the subgroup size is set to 5. For the affine merge mode, the subgroup size is set to 3. The merge candidates in each subgroup are reordered incrementally according to the TM-based cost value. In some instances, for simplicity, the merge candidates in the last subgroup (instead of the first subgroup) are not reordered.

[0106] The TM cost of a merge candidate can be measured by the sum of absolute differences (SAD) (or other measurement) between a sample of the template in the current block and its corresponding reference sample. The template includes a set of reconstructed samples adjacent to the current block. The reference sample of the template is located using motion information of the merge candidate.

[0107] When merging candidates utilize bidirectional prediction, the reference sample of the merging candidate template is also generated through bidirectional prediction, as shown in Figure 8. For a sub-block-based merging candidate with a sub-block size equal to Wsub × Hsub, the above template includes several sub-templates of size Wsub × 1, and the left template includes several sub-templates of size 1 × Hsub. As shown in Figure 9, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample of each sub-template.

[0108] In some instances, the encoding device 104 and / or the decoding device 112 can generate a merge candidate list for geometric segmentation mode (e.g., a single prediction merge candidate list). For example, in VVC, geometric segmentation mode (which may be referred to as GEO mode) is supported for inter-frame prediction. When using GEO mode, the CU or other blocks can be separated into two parts by geometrically positioned straight lines, as shown in Figure 10.

[0109] The position of the separation line can be mathematically derived from the angle and offset parameters of a specific partition. Each part of the geometric partition in the CU is used for inter-frame prediction using its own motion; in some cases, only a single prediction is allowed for each partition, in which case each part has a motion vector and a reference index.

[0110] The encoding device 104 and / or decoding device 112 can directly derive the single prediction candidate list for the GEO mode from the general merging candidate list, as shown in Figure 11. For example, denoted as n, which is the index of the single prediction motion in the geometric single prediction candidate list, the LX motion vector of the nth merging candidate (where X equals the isotope of n (even or odd)) is used as the nth single prediction motion vector for the geometric segmentation mode. These motion vectors are marked with "x" in Figure 11. In the absence of a corresponding LX motion vector for the nth extended merging candidate, the L(1-X) motion vector of the same candidate is used instead as the single prediction motion vector for the geometric segmentation mode.

[0111] As described above, the systems and techniques described herein can utilize multi-level ARMC (e.g., two-level ARMC) technology. Using two-level ARMC as an illustrative example, in the first ARMC level, the encoding device 104 and / or decoding device 112 can group available candidates using a first grouping method, and can apply a reordering within each group (e.g., reordering individually within each group). The encoding device 104 and / or decoding device 112 can then construct a first merged candidate list in the order of the groups processed by the first ARMC. The input to the second ARMC level may include the first merged candidate list. In the second ARMC level, the encoding device 104 and / or decoding device 112 can group candidates (e.g., from the first merged candidate list) using a second grouping method. The encoding device 104 and / or decoding device 112 can apply a reordering within each group of candidates generated by the second grouping method (e.g., reordering individually within each group). The encoding device 104 and / or decoding device 112 can then construct a second merged candidate list in the order of the groups processed by the second ARMC. The systems and techniques described in this article can be applied individually or in any combination.

[0112] In some instances, in the first ARMC, when adding candidates to the first merged candidate list, the last X candidates in each group after reordering (e.g., the last three candidates, the last two candidates, the last candidate, or other number of candidates) may be discarded to reduce the number of candidates. The first grouping method may be different from the second grouping method, or it may be the same grouping method. The reordering criteria in the first ARMC may also be different from the reordering criteria in the second ARMC, or it may be the same reordering.

[0113] In one illustrative example, the first grouping method in the first ARMC is based on candidate types (e.g., SMVP, TMVP, NA-TMVP, HMVP, SbTMVP, I-AffineMVP, and / or C-AffineMVP candidate types). In some cases, the candidates for each type are reordered individually. A first merged candidate list can then be constructed (e.g., by encoding device 104 and / or decoding device 112) in a predefined order of candidate types. After constructing the first merged candidate list, encoding device 104 and / or decoding device 112 can apply a second ARMC to the first merged candidate list to further group the candidates in the first merged candidate list. The candidates in each group can then be reordered. After completing the second ARMC processing, encoding device 104 and / or decoding device 112 can construct a second merged candidate list as described above. In some cases, the second merged candidate list is the final candidate list for a specific merging pattern.

[0114] In another alternative or additional illustrative instance, the second grouping method in the second ARMC is based on the candidate index. An illustrative example of the second ARMC is the ARMC currently used in ECM, such as the ECM ARMC described above. For example, the first ARMC and the second ARMC can reorder the merge candidates based on the TM cost value, as described above regarding the ECM ARMC.

[0115] In another alternative or additional illustrative example, the first ARMC and the second ARMC are implemented after the merge candidate list is constructed. For example, after constructing the merge candidate list, the first ARMC may include grouping candidates based on candidate type and reordering candidates for each candidate type. The merge candidate list reordered according to the first ARMC may be further grouped and reordered according to the second ARMC.

[0116] In some cases, there are multi-level ARMCs, where the grouping and reordering methods may differ at at least two levels. In some instances, there are P1 ARMCs before constructing the first merge candidate list and P2 ARMCs after constructing the first merge candidate list, where P1 and P2 are positive integers. In some cases, the proposed two-level or multi-level ARMCs can be applied to the construction of merge candidate lists in any merge mode, such as general merge lists, TM merge lists, MMVD merge lists, CIIP merge lists, GPM merge lists, sub-block merge lists, etc.

[0117] Figure 15 is a block diagram illustrating an example of a multi-level ARMC 1500 according to the present disclosure. As indicated above, in some cases, the encoding device 104 and / or the decoding device 112 may utilize a merging mode that allows a block (e.g., an inter-frame prediction PU or other block) to inherit the same one or more motion vectors, prediction directions, and / or one or more reference picture indices from another block (e.g., another inter-frame prediction PU or other block). This merging mode may be based on a merging candidate list. In some cases, a multi-level ARMC (such as multi-level ARMC 1500) may be used to select merging candidates for the merging candidate list. In this example, one or more grouping techniques may be applied to a set of prediction candidates (e.g., another inter-frame prediction PU or other block) to produce prediction candidate group 1502. For example, a TMVP (discussed in detail below) may be applied to a set of up to 30 TMVP prediction candidates to determine a prediction candidate group (e.g., for prediction candidate group 1502) comprising 9 prediction candidates. As another example, NA-SMVP (discussed in detail below) can be applied to a set of up to 80 NA-SMVP prediction candidates to determine a group of 18 prediction candidates (such as prediction candidate group 1502).

[0118] Prediction candidate group 1502 can be reordered 1504 into a reordered prediction candidate group 1506. In some cases, prediction candidate group 1502 can be reordered 1504 based on the ARMC TM cost value of the prediction candidates in prediction candidate group 1502. From the reordered prediction candidate group 1506, a merge candidate 1510 can be selected 1508. For example, the first prediction candidate from the reordered prediction candidate group 1506 can be selected as merge candidate 1510. This merge candidate 1510 can be added to the merge candidate list 1512 (e.g., a second candidate list). In some cases, the merge candidate list 1512 can be a TM merge candidate list. In some cases, the merge candidate list 1512 can include other merge candidates 1514. Other merge candidates 1514 can be added to the merge candidate list 1512 via other grouping methods (such as SMVP). In some cases, the merge candidate list 1512 may include l merge candidates, which are added to the merge candidate list 1512 via grouping methods (such as the various candidate types discussed below). In some cases, l can be 10. Better merge candidates can be obtained by obtaining merge candidates via first-level ARMC (such as via grouping methods, reordering, and selection of predictive candidates). For example, first-level ARMC allows for the selection of TMVP candidates from a wider range of predictive candidates, rather than obtaining TMVP candidates from two possible candidates (as shown in Figure 4 and discussed above). The selection process of first-level ARMC also allows for the selection of more predictive merge candidates compared to using any available predictive candidates.

[0119] In some cases, merge candidate 1510 can be checked against other merge candidates 1514 that are already in merge candidate list 1512, and if merge candidate 1510 is also not in merge candidate list 1512, then merge candidate 1510 is added. If merge candidate 1510 is already in merge candidate list 1512, then merge candidate 1510 can be discarded and merge candidate list 1512 can be zero-padded.

[0120] In some cases, the second reordering method 1516 can be applied to the merge candidate list 1512 to produce a reordered merge candidate list 1518. For example, the merge candidate list 1512 can be reordered 1516 based on the ARMC™ cost values ​​of the merge candidates in the merge candidate list 1512. Subsequently, a merge block 1522 can be selected from the reordered merge candidate list 1518. For example, a first merge candidate can be selected from the reordered merge candidate list 1518.

[0121] In some illustrative examples, the first ARMC of the two-level ARMC described herein groups the candidates into multiple groups and reorders the N candidates in each group based on a cost criterion. In such an example, M candidates are selected from the N candidates, where M and N are positive integers and M ≤ N. After the first ARMC processes each group, a first merged candidate list is constructed in a predefined order of the groups.

[0122] In some cases, the first ARMC will group candidates with the same candidate type into one group. For example, in the general merging mode, SMVP candidates are grouped into one group, and the three other candidate types (e.g., TMVP candidates, NA-TMVP candidates, and HMVP candidates) are grouped into three different groups respectively. In some cases, the values ​​of M and N may be different in different groups. In some cases, at least one of M and N may be different for different CU sizes. For example, for larger CU sizes, there may be a larger number of M or N. In some instances, N is the number of all candidates in the candidate type.

[0123] Illustrative examples of the first ARMC of various candidate types are described below:

[0124] TMVP Reordering: For example, M1 TMVP candidates can be selected from the reordered N1 TMVP candidates based on the ARMC TM cost value, where M1 ≤ N1. In some cases, one TMVP candidate can be selected from the reordered 9 TMVP candidates. Let the i-th TMVP candidate be denoted as TMVPi, and the N1 TMVP candidates consist of different positions in the co-located image. As an illustrative example, referring to FIG12, TMVPi is derived from position Ci, and TMVPj is derived from position Cj. Ci and Cj can be any position adjacent to the current CU. In another illustrative example, the N1 TMVP candidates consist of distinct position pairs in the co-located image; one example is that TMVPi is derived from the C2 and C3 pair and TMVPj is derived from the Ci and Cj pair, where for the prediction list LX (e.g., X equals 0 or 1), if C2 in LX is available, then TMVPi in LX is derived from C2; otherwise, if C3 in LX is available, then TMVPi in LX is derived from C3. If Ci in LX is available, then TMVPj in LX is derived from Ci; otherwise, if Cj in LX is available, then TMVPj in LX is derived from Cj.

[0125] In another illustrative example, the N1 TMVP candidates consist of the same positions in different co-located images. For example, TMVPi can be derived from position Ci in co-located image A, and TMVPj can be derived from the same position in co-located image B.

[0126] Alternatively or in some instances, the N1 TMVP candidates include the same positions using different prediction lists. In an illustrative example, TMVP0 is derived from the Ci and Cj pairs, where for prediction list LX (X equals 0 or 1), if Ci in LX is available, then TMVP0 in LX is derived from Ci; otherwise, if Cj in LX is available, then TMVP0 in LX is derived from Cj. In such an instance, TMVP1 uses prediction list 0 (L0), such as deriving TMVP1 equal to TMVP0 by using only L0 (e.g., TMVP0 without L1). In such an instance, TMVP2 uses prediction list 1 (L1), such as deriving TMVP2 equal to TMVP0 by using only L1 (e.g., TMVP0 without L0).

[0127] For example, TMVP candidates can be constructed using either double-prediction positions (where two motion vectors are used) or single-prediction positions (where a single motion vector is used). When constructing TMVP candidates, position Ci can be examined, and subsequently position Cj can be examined. Each of Ci and Cj can be either double-prediction or single-prediction. If the MV prediction at a position is double-prediction, then there are two motion vectors at that position, one from L0 and one from L1. Similarly, if the MV prediction is single-prediction, then there is one motion vector, either from L0 or from L1. In an illustrative example, 10 TMVP candidates can be constructed from 10 pairs of positions, and these 10 TMVP candidates can be double-prediction and / or single-prediction candidates. In some cases, TMVP candidates can be used to generate another set of TMVP candidates from L0 and another set of TMVP candidates from L1. For example, continuing with the example above, for a total of 30 possible prediction candidates, 10 TMVP candidates can be used to generate another 10 TMVP candidates from L0 and another 10 TMVP candidates from L1. In some cases, the first nine positions can be used to predict candidate groups (such as predicting candidate group 1502).

[0128] Alternatively or in some instances, the N1 TMVP candidates include the same positions using different scaling factors (e.g., a × tb / td), where a can be any non-zero value. In an illustrative example, TMVP0 uses a scaling factor tb / td (a is set to 1) to derive the TMVP from the Ci and Cj pairs as described above. In this instance, TMVP1 uses a scaling factor (9 / 8) × tb / td (a is set to 9 / 8) to derive the TMVP from the Ci and Cj pairs as described above. In this instance, TMVP2 uses a scaling factor (1 / 8) × tb / td (a is set to 1 / 8) to derive the TMVP from the Ci and Cj pairs as described above.

[0129] The above TMVP instances can be used individually or combined in any form. In one instance of a combination of TMVP instances, the N1 TMVP candidates include different position pairs and use different prediction lists. For example, in this instance, TMVP0 is derived from the C0 and C1 pair, where for prediction list LX (X equals 0 or 1), if C0 in LX is available, then TMVP0 in LX is derived from C0; otherwise, if C1 in LX is available, then TMVP0 in LX is derived from C1. In this instance, TMVP1 uses prediction list 0 (L0), such as deriving TMVP1 equal to TMVP0 by using only L0 (e.g., TMVP0 without L1). In this instance, TMVP2 uses prediction list 1 (L1), such as deriving TMVP2 equal to TMVP0 by using only L1 (e.g., TMVP0 without L0). In this example, TMVP3 is derived from the Ci and Cj pairs, where for a prediction list LX (X equals 0 or 1), if Ci in LX is available, then TMVP3 in LX is derived from Ci; otherwise, if Cj in LX is available, then TMVP3 in LX is derived from Cj. Furthermore, in this example, TMVP4 uses prediction list 0 (L0), such as deriving TMVP4 equal to TMVP3 using only L0 (e.g., TMVP3 without L1). TMVP5 uses prediction list 1 (L1), such as deriving TMVP5 equal to TMVP3 using only L1 (e.g., TMVP3 without L0). In another illustrative example of a combination of TMVP instances, N1 TMVP candidates comprise different position pairs, using different prediction lists and different scaling factors a×tb / td.

[0130] NA-SMVP Reordering (sometimes also referred to as Non-Adjacent MVP): In an illustrative example, M2 NA-SMVP candidates can be selected from the reordered N2 NA-SMVP candidates based on ARMC TM cost values, where M2 ≤ N2. The N2 NA-SMVP candidates consist of MVs from spatially distinct adjacent blocks, as shown in Figure 5. Alternatively, in some cases, adjacent blocks consist of different location types. In one instance, the type can be based on geometric orientation. For example, as an illustrative example, referring to Figure 13, blocks 53, 57, 61, 65, 69, and 73 are in one orientation, and blocks 51, 55, 59, 63, 67, and 71 are in another orientation. In another instance, the type can be based on geometric layers. As an illustrative example, referring again to Figure 13, blocks 47, 49, 42, 71, 45, 73, 46, 74, 44, 72, 43, 50, and 48 are in one layer, and blocks 38, 40, 33, 67, 36, 69, 37, 70, 35, 68, 34, 41, and 39 are in another layer. The type can be based on any combination of geometric orientations, geometric layers, and / or other factors associated with the blocks. In some patterns, positions are classified into groups g1 based on a first position type, where g1 is a positive integer greater than or equal to 2. In this pattern, each group can be further classified into subgroups g2 based on a second position type, where g2 is a positive integer greater than or equal to 2. Subsequently, N2 NA-SMVP candidates can be constructed in the order of group 1, group 2, to the final group g1, and within each group i, NA-SMVP candidates are constructed in the order of subgroup 1, subgroup 2, to the final subgroup g2. As an illustrative example, referring again to Figure 13, positions can be classified into two groups (g1 = 2) based on geometric orientation. Square blocks from five geometric orientations and diamond blocks from four geometric orientations are classified as Group 1, and circular blocks forming four geometric orientations are classified as Group 2. Each group is further classified into seven subgroups (g2 = 7) based on geometric layers. Taking Group 1 as an example, subgroup 1 includes 1, 4, 5, 3, and 2; subgroup 2 includes 11, 13, 6, 9, 10, 8, 7, 14, and 12; and subgroup 3 includes 20, 22, 15, 18, 19, 17, 16, 23, and 21, etc. In some cases, N2 NA-SMVP candidates are first constructed from the blocks in Group 1. If N2 is not reached, candidates from the blocks in Group 2 are further added to the NA-SMVP candidates until N2 is reached. In group i, NA-SMVP candidates are constructed in the order of subgroup 1, subgroup 2 to subgroup 7, until N2 candidates are added to the NA-SMVP list. In some cases, N2 can be 18 candidates.

[0131] SbTMVP Reordering: For example, M3 SbTMVP candidates can be selected from the reordered N3 SbTMVP candidates based on the ARMC TM cost value, where M3 ≤ N3. Let the i-th SbTMVP candidate be denoted as SbTMVPi, and the N3 SbTMVP candidates consist of different motion shifts applied to the co-location picture. In an illustrative example, SbTMVPi applies the MV information of the adjacent block A1 in Figure 6A to the coordinates of the current block to obtain sub-CU level motion information from the co-location picture, and SbTMVPj applies the MV information of the adjacent block B1 in Figure 6A to the coordinates of the current block to obtain sub-CU level motion information from the co-location picture. In another example of SbTMVP reordering, the N3 SbTMVP candidates can consist of different motion shifts derived from a general merge list, where the general merge list has been reordered by ARMC. In another example, SbTMVPi first performs an SMVP reordering, where M3 SMVP candidates are selected from the reordered N10 SMVP candidates based on the ARMC TM cost value, where M3 ≤ N10. One or more i-th SMVP candidates consist of MVs from spatially adjacent blocks, as shown in Figure 2. The selected M3 SMVPs are then used to shift co-located blocks in the current CU during SbTMVP execution. In an illustrative example, M3 = 1, which indicates that an SMVP is selected from the reordered N10 SMVP candidates. One SMVP (according to M3 = 1) is used to shift co-located blocks in the current CU during SbTMVP operation.

[0132] In some cases, the same reordering method can be applied to SMVP, PA-MVP, I-AffineMVP and C-AffineMVP to obtain a predefined number of candidates from the reordered candidates in the merged candidate type.

[0133] For example, assuming the size of the TM merge candidate list is M4, we can first derive N4 TM merge candidates, where M4 ≤ N4, and then the second ARMC reorders the N4 TM merge candidates. In this instance, M4 TM merge candidates are selected from the N4 reordered TM merge candidates.

[0134] In another instance, some candidate types (such as HMVP and SMVP) are not reordered by the first ARMC. In another instance, some candidate types (such as TMVP) are reordered by the first ARMC in one merge list (such as a general merge list), but not in another merge list (such as a TM merge list).

[0135] In another example, the first grouping method in the first ARMC is to group at least two candidate types into one group. For example, the encoding device 104 and / or the decoding device 112 can group HMVP and PA-MVP into one group. For example, assuming there are X1 and X2 candidates in HMVP and PA-MVP, the X1+X2 candidates in HMVP and PA-MVP are reordered by ARMC TM cost, and Y candidates are selected from the X1+X2 candidates, where Y≤X1+X2. Another example is to group SMVP and PA-MVP into one group. Assume there are X1 and X2 candidates in SMVP and PA-MVP. Subsequently, the X1+X2 candidates in SMVP and PA-MVP are reordered by ARMC TM cost, and Y candidates are selected from the X1+X2 candidates, where Y≤X1+X2. In another illustrative example, the best PA-MVP candidate is selected from X2 PA-MVP candidates reordered by the first ARMC, and this PA-MVP candidate and X1 SMVP candidates are further reordered by the first ARMC. In another illustrative example, the best SMVP candidate is selected from X1 SMVP candidates by ARMC TM cost value, and the best PA-MVP candidate is selected from X2 PA-MVP candidates by ARMC TM cost value. Then, the best SMVP candidate and the best PA-MVP candidate are further compared by ARMC TM cost value as follows: if the TM cost value of the best PA-MVP candidate is less than the TM cost value of the best SMVP candidate, the candidate type order is {best PA-MVP candidate, SMVP candidate}; otherwise, it is {SMVP candidate, best PA-MVP candidate}. In another illustrative example, if the candidate type order is {best PA-MVP candidate, SMVP candidate}, more PA-MVP candidates (e.g., N9 PA-MVP candidates) are constructed and added to the merged candidate list. In another illustrative example, N9 PA-MVP candidates can be reordered by a first ARMC, and M9 candidates are selected from the N9 PA-MVP candidates, where M9 ≤ N9. In another illustrative example, if the candidate type order is {best PA-MVP candidate, SMVP candidate}, then the i-th merged candidate mergeCand_i of at least one of the candidate types in the first ARMC and / or the unreordered group in the second ARMC is replaced with a PA-MVP by averaging the best PA-MVP candidate and that candidate mergeCand_i (e.g., pairwiseAverage(mergeCand_i, best PA-MVP candidate)), where pairwiseAverage is performed in the same way as the PA-MVP in the current ECM.In another illustrative example, the MV information of the i-th merged candidate mergeCand_i of at least one of the candidate types in the first ARMC and / or the unordered group in the second ARMC is replaced with PA-MVP by averaging the first SMVP candidate, the second SMVP candidate, and the candidate mergeCand_i (e.g., pairwiseAverage(mergeCand_i, first SMVP candidate, second SMVP candidate)), wherein the average motion vector is computed separately for each reference list, and averaging is performed as long as at least two MVs of mergeCand_i, the first SMVP candidate, and the second SMVP candidate are available in the reference list.

[0136] In some cases, pruning schemes or procedures can be applied to at least one candidate list in various candidate groups (e.g., TMVP, NA-SMVP, etc.) to remove redundant candidates. In an illustrative example, a pruning procedure can be applied when constructing an NA-SMVP candidate list with a list size equal to N². In this example, a comparison can be performed between the candidate to be added to the NA-SMVP candidate list and the candidates already added to the NA-SMVP list. Based on the comparison results, the candidate under consideration may not be added to the NA-SMVP list. For example, the i-th candidate is not added to the list if at least one of the following conditions is true in the j-th candidate of the candidate group (e.g., an NA-SMVP candidate list with a list size of N²): 1. The i-th and j-th candidates use the same list of reference images (e.g., L0 or L1). 2. The i-th and j-th candidates use the same index of the list of reference images. 3. If the absolute value of the horizontal MV difference between the i-th and j-th candidates is not greater than the predefined MV difference threshold Tx, and the absolute value of the vertical MV difference between the i-th and j-th candidates is not greater than the predefined MV difference threshold Ty (checking both L0 MV and L1 MV), where Tx and Ty can be any positive pre-assigned value, such as 1 / 4, 1 / 2, 1, and / or other values. The values ​​of Tx and Ty can be equal.

[0137] In some cases, Tx and Ty are pattern-dependent. In one instance, when constructing the NA-SMVP list, Tx1 and Ty1 for one merge pattern and Tx2 and Ty2 for another merge pattern are different, where Tx1 is not equal to Tx2 and Ty1 is not equal to Tx2.

[0138] In some configurations, a first merge candidate list is constructed after the candidates are reordered in the group by a first ARMC (e.g., based on candidate type), and no second ARMC is applied to the first merge candidate list. In this configuration, the first merge candidate list is the final merge candidate list of the merge pattern.

[0139] In some instances, the systems and techniques described herein include non-adjacent TMVPs. For example, encoding device 104 and / or decoding device 112 can derive TMVPs from non-adjacent co-located blocks and can add TMVPs to merge lists (such as general merge lists and / or TM merge lists). Figure 14 illustrates an example where blocks 1 to 5 are used for SMVPs, blocks 6 to 23 are used for NA-SMVPs, blocks C0 and C1 are used for TMVPs, and blocks C2 to C11 are used for the proposed non-adjacent TMVP (NA-TMVP). Note that the pattern is not limited to the pattern depicted in Figure 14. An NA-TMVP candidate can use blocks located anywhere that is not adjacent to the current CU.

[0140] In some cases, regarding the order of candidate types in the merge list, the encoding device 104 and / or the decoding device 112 may insert NA-TMVP candidate types into the merge list between NA-SMVP candidate types and HMVP candidate types. In another instance, the encoding device 104 and / or the decoding device 112 may insert NA-TMVP candidates into NA-SMVP candidates based on their distance to the current CU. For example, NA-SMVP and NA-TMVP candidates with similar distances to the current CU are grouped together, and groups with shorter distances are ordered with higher priority for insertion into the merge candidate list. As an illustrative example, referring to Figure 14, blocks 6, 7, 8, and C2 are in the first group; blocks 9, 10, 11, 12, 13, C3, C4, and C5 are in the second group; blocks 14, 15, 16, 17, 18, C6, C7, and C8 are in the third group; and blocks 19, 20, 21, 22, 23, C9, C10, and C11 are in the fourth group. In some cases, the order of insertion into the merge candidate list is the first, second, third, and fourth groups, and NA-SMVPs have a higher priority than NA-TMVPs within the same group.

[0141] In one instance, the i-th NA-TMVP candidate is represented as NA-TMVPi, and NA-TMVPi is derived from position Ci. In another instance, NA-TMVPi is derived from a pair of positions Cj and Ck. For example, for a prediction list LX, if Cj in LX is available, then NA-TMVPi in LX is derived from Cj; otherwise, if Ck in LX is available, then NA-TMVPi in LX is derived from Ck.

[0142] In some cases, the first ARMC level described above can be applied to NA-TMVP. For example, M5 NA-TMVP candidates can be selected from a reordered N5 NA-TMVP candidates based on the ARMC TM cost value, where M5 ≤ N5. In another instance, NA-SMVP and NA-TMVP candidates with similar distances to the current CU are grouped into a group. For example, assuming there are N6 candidates in this group, M6 candidates are selected from the reordered N6 candidates based on the ARMC TM cost value, where M6 ≤ N6.

[0143] In another instance, an NA-TMVP candidate can be inserted between the TMVP candidate and the NA-SMVP candidate. In this instance, the TMVP and NA-TMVP can be grouped together by the first ARMC level and reordered together.

[0144] In some cases, when candidates in a candidate group come from different distances, candidates can be added to the list from nearest to farthest until the list is full. For example, candidates in the NA-SMVP candidate group come from different distances (e.g., from any one or more blocks 6, 7 to 23 in Figure 14). As an example, referring to Figure 14, blocks 6, 7, and 8 are at the same distance level 1, blocks 9, 10, 11, 12, and 13 are at the same distance level 2, blocks 14, 15, 16, 17, and 18 are at the same distance level 3, and blocks 19, 20, 21, 22, and 23 are at the same distance level 4. In one illustrative example, N2 is set to 18. In other examples, N2 can be set to any other suitable value. As indicated above, N2 is the size of the NA-SMVP list. The N2 candidates in the NA-SMVP list will be reordered, and the M2 NA-SMVP candidates with the lowest TM cost will be selected from the N2 candidates and added to the merge list. Subsequently, the N2 NA-SMVP candidates (e.g., N2 = 18) are added to the NA-SMVP candidate list from farthest to nearest (e.g., in the order of blocks 6, 7, 8 to 23). In some cases, candidates at each distance level are further divided into Q1 subgroups, and candidates are added to the list from nearest to farthest and from subgroup 1 to subgroup Q1 until the list is full. For example, if Q1 is set to 2, then blocks 6 and 7 are set as subgroup 1 at distance level 1, block 8 is set as subgroup 2 at distance level 1, blocks 9, 10, 11 and 12 are set as subgroup 1 at distance level 2, block 13 is set as subgroup 2 at distance level 2, blocks 14, 15, 16 and 17 are set as subgroup 1 at distance level 3, block 18 is set as subgroup 2 at distance level 3, blocks 19, 20, 21 and 22 are set as subgroup 1 at distance level 4, and block 23 is set as subgroup 2 at distance level 4. Then, N2 NA-SMVP candidates (e.g., N2 = 18) can be added to the NA-SMVP candidate list in order from farthest to nearth and from subgroup 1 to subgroup Q1 (e.g., in the order of blocks 6, 7, 9, 10, 11, 12, 14, 15, 16, 17, 19, 20, 21, 22, and then 8, 13, 18, 23).

[0145] In some cases, the encoding device 104 and / or the decoding device 112 can reorder the candidates in the GEO merging candidate list. For example, the encoding device 104 and / or the decoding device 112 can construct the GEO merging candidate list independently of the general merging candidate list. In one instance, the encoding device 104 and / or the decoding device 112 can construct the GEO merging candidate list based on at least one of the following candidate types: SMVP, TMVP, NA-SMVP, NA-TMVP, HMVP, and PA-MVP. In some cases, if a candidate is a double-prediction candidate (generated based on double prediction), the double-prediction candidate is separated into two single-prediction candidates. Subsequently, for each candidate type, ARMC is applied to reorder the candidates in the candidate type based on the TM cost value. Using the SMVP candidate type in Figure 2 as an example, there are 5 blocks (including blocks B0, A0, B1, A1, and B2). If B0 is a dual-prediction candidate, it is split into two SMVP candidates: SMVP0, which uses the MV information from prediction list 0 of block B0, and SMVP1, which uses the MV information from prediction list 1 of block B0. If A0 is a single-prediction candidate, SMVP2 uses the MV information from block A0. In some instances, assuming N7 single-prediction SMVP candidates are collected, M7 candidates are selected from the reordered N7 candidates based on TM cost, where M7 ≤ N7. In some cases, the same method can be applied to TMVP, NA-SMVP, NA-TMVP, and HMVP.

[0146] In some instances, there are two independent GEO merge candidate lists, such as one for geometry partition 0 and another for geometry partition 1. In such instances, list construction and ARMC can be applied to both lists. In some cases, there are multiple independent GEO merge candidate lists, where the GEO merge candidate lists correspond to specific geometry partition angles and specific geometry partition indices. When ARMC is applied to the GEO merge list, the TM cost can be calculated based on a predefined template using the TM lookup table. Table 1 below illustrates an example of a TM lookup table, where only the top template of the current CU in Figure 8 is used to calculate the TM cost of the GEO merge list corresponding to partition angle index 0 and the first partition, and the top-left template of the current CU is used to calculate the TM cost of the GEO merge list corresponding to partition angle index 0 and the second partition. Table 1. TM Lookup Table for Calculating the TM Cost of GEO Merge Candidate Lists Partition angle 0 2 3 4 5 8 11 12 13 14 First Division A A A A L+A L+A L+A L+A A A Second Division L+A L+A L+A L L L L L+A L+A L+A Partition angle 16 18 19 20 twenty one twenty four 27 28 29 30 First Division A A A A L+A L+A L+A L+A A A Second Division L+A L+A L+A L L L L L+A L+A L+A

[0147] In some instances, the GEO merge candidate list is constructed based on a general merge candidate list, as described above regarding the construction of a single-prediction GEO merge candidate list. For example, if the candidate in the general merge list is a double-prediction candidate, the encoding device 104 and / or the decoding device 112 can separate a double-prediction candidate into two single-prediction candidates, and can apply ARMC to select one of these two single-prediction candidates based on the ARMC TM cost value.

[0148] Figure 16 is a flowchart illustrating a process 1600 for performing bit rate estimation 1600 according to various embodiments of the present disclosure. At operation 1602, process 1600 includes: obtaining a first plurality of prediction candidates associated with video data. In some cases, at least one camera is configured to capture one or more frames associated with the video data. At operation 1604, process 1600 includes: determining a first group of prediction candidates by at least partially applying a first grouping method to the first plurality of prediction candidates. In some cases, the first grouping method is based on a plurality of candidate types associated with the first plurality of prediction candidates. In some cases, the candidate list includes a first merged candidate in a predefined order based on the plurality of candidate types. In some cases, the plurality of candidate types includes at least one of the following: Spatial Motion Vector Predictor (SMVP) type, Temporal Motion Vector Predictor (TMVP) type, Non-Adjacent Temporal Motion Vector Predictor (NA-TMVP) candidate, History-Based Motion Vector Predictor (HMVP) candidate, Sub-Block-Based Temporal Motion Vector Predictor (SbTMVP) candidate, Inherited Affine Merging (I-AffineMVP) candidate, or Constructive Affine Merging (C-AffineMVP) candidate. In some cases, the first grouping method is one of Temporal Motion Vector Predictor (TMVP) or Non-Adjacent Temporal Motion Vector Predictor (NA-TMVP). In some cases, the first group of prediction candidates includes fewer prediction candidates than the first plurality of prediction candidates.

[0149] At operation 1606, process 1600 includes: reordering the first set of prediction candidates. In some cases, process 1600 includes: reordering the first set of prediction candidates based on cost values. In some cases, process 1600 includes: reordering the first set of prediction candidates by powers of the cost values. In some cases, the cost values ​​are based on template matching. In some cases, process 1600 includes: discarding at least one candidate from the reordered first set of prediction candidates before adding a first merged candidate to the candidate list.

[0150] At operation 1608, process 1600 includes: selecting a first merge candidate from the reordered first group of prediction candidates. At operation 1610, process 1600 includes: adding the first merge candidate to a candidate list. In some cases, the candidate list is a merge candidate list for a merge pattern. In some cases, process 1600 includes: determining a second group of prediction candidates at least in part by applying a second grouping method to the candidate list. In some cases, process 1600 includes: reordering the second group of prediction candidates; selecting a second merge candidate from the reordered second group of prediction candidates; and adding the second merge candidate to the candidate list. In some cases, process 1600 includes: determining that the first merge candidate is not in the candidate list; and adding the first merge candidate to the candidate list based on the determination that the first merge candidate is not in the candidate list.

[0151] In some cases, process 1600 includes: generating a prediction of the current block of the video data based on a candidate list. In some cases, process 1600 includes: decoding the current block of the video data based on the prediction. In some cases, process 1600 includes: encoding the current block of the video data based on the prediction. In some cases, process 1600 includes: displaying an image from the video data.

[0152] In some instances, the processes described herein may be performed by a computing device or apparatus (such as encoding device 104, decoding device 112, and / or any other computing device). In some cases, the computing device or apparatus may include a processor, microprocessor, microcomputer, or other components of a device configured to perform the steps of the processes described herein. In some instances, the computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) including video frames. For example, the computing device may include a camera device, which may or may not include a video transcoder. As another example, the computing device may include a mobile device with a camera (e.g., a camera device (such as a digital camera, IP camera, etc.), a mobile phone or tablet device including a camera, or other types of devices with a camera). In some cases, the computing device may include a display for displaying images. In some instances, the camera or other capturing device for capturing video data is separate from the computing device, in which case the computing device receives the captured video data. The computing device may also include a network interface, transceiver, and / or transmitter configured to transmit video data. Network interfaces, transceivers, and / or transmitters can be configured to transmit Internet Protocol (IP) based data or other network data.

[0153] The processes described herein can be implemented using hardware, computer instructions, or a combination thereof. In the context of computer instructions, these operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0154] Furthermore, the processes described herein can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented by hardware, or a combination thereof. As mentioned above, the code can be stored, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors on a computer-readable or machine-readable storage medium. The computer-readable or machine-readable storage medium can be non-transitory.

[0155] The decoding techniques discussed herein can be implemented in an example video encoding and decoding system (e.g., System 100). In some instances, the system includes a source device that provides encoded video data to be decoded later by a destination device. Specifically, the source device provides the video data to the destination device via computer-readable media. The source and destination devices can include any of a variety of devices, including desktop computers, notebook computers (i.e., laptops), tablet computers, set-top boxes, mobile phones (such as so-called "smart" phones), so-called "smart" boards, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source and destination devices can be configured for wireless communication.

[0156] The destination device can receive encoded video data to be decoded via computer-readable media. Computer-readable media can include any type of media or device capable of moving encoded video data from the source device to the destination device. In one example, computer-readable media can include communication media enabling the source device to directly and instantly transmit encoded video data to the destination device. The encoded video data can be modulated according to communication standards such as wireless communication protocols, and the encoded video data can be transmitted to the destination device. Communication media can include any wireless or wired communication media, such as radio frequency (RF) spectrum or one or more physical transmission lines. Communication media can form part of a packet-based network, such as a local area network, wide area network, or global area network (such as the Internet). Communication media can include routers, switches, base stations, or any other means that can be used to facilitate communication from the source device to the destination device.

[0157] In some instances, encoded data can be output from an output interface to a storage device. Similarly, encoded data can be accessed from a storage device via an input interface. The storage device can include any of a variety of distributed or locally accessible data storage media, such as hard drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. In other instances, the storage device can correspond to a file server or another intermediate storage device that can store encoded video generated by a source device. The destination device can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing encoded video data and sending it to the destination device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device can access the encoded video data via any standard data connection, including an Internet connection. The connection can include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on a file server. Transmission of the encoded video data from the storage device can be streaming, downloading, or a combination thereof.

[0158] The technologies disclosed herein are not necessarily limited to wireless applications or setups. These technologies can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as HTTP-based Dynamic Self-Adjusting Streaming (DASH)), digital video encoded onto data storage media, decoding digital video stored on data storage media, or other applications. In some instances, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video replay, video broadcasting, and / or video telephony.

[0159] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device may include an input interface, a video decoder, and a display device. The video encoder of the source device may be configured to apply the techniques disclosed herein. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device may receive video data from an external video source such as an external camera. Similarly, the destination device may interface with an external display device instead of including an integrated display device.

[0160] The example system described above is merely one example. Techniques for parallel processing of video data can be implemented by any digital video encoding and / or decoding device. Although, in general, the techniques disclosed herein are implemented by video encoding devices, these techniques can also be implemented by video encoders / decoders commonly referred to as "CODECs." Furthermore, the techniques disclosed herein can also be implemented by video preprocessors. The source device and destination device are merely examples of such decoding devices: where the source device generates decoded video data for transmission to the destination device. In some instances, the source device and destination device can operate in a substantially symmetrical manner, such that each of these devices includes video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video replay, video broadcasting, or video telephony.

[0161] A video source may include video capturing devices, such as a camera, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. Alternatively, a video source may generate computer graphics-based data as source video, or a combination of real-time video, archived video, and computer-generated video. In some cases, if the video source is a camera, the source device and the destination device may form a so-called camera phone or video phone. However, as mentioned above, the techniques described in this disclosure are generally applicable to video decoding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by a video encoder. The encoded video information can be output to a computer-readable medium via an output interface.

[0162] As mentioned, computer-readable media may include temporary media such as wireless broadcasting or wired network transmission, or storage media such as hard disks, flash memory drives, compressed optical discs, digital video discs, Blu-ray discs (i.e., non-temporary storage media), or other computer-readable media. In some instances, a network server (not shown) may, for example, receive encoded video data from a source device via network transmission and provide the encoded video data to a destination device. Similarly, a computing device in a media production facility, such as a disc stamping facility, may receive encoded video data from a source device and produce an optical disc containing the encoded video data. Therefore, in various instances, computer-readable media can be understood to include one or more computer-readable media of various forms.

[0163] The input interface of the destination device receives information from computer-readable media. The information from the computer-readable media may include syntax information defined by the video encoder (which is also used by the video decoder), including syntax elements describing the characteristics and / or processing of blocks and other decoding units (e.g., groups of pictures (GOPs)). The display device displays the decoded video data to the user and may include any of a variety of display devices, such as cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device. Various examples of this disclosure have been described.

[0164] Specific details of the encoding device 104 and the decoding device 112 are illustrated in Figures 17 and 18, respectively. Figure 17 is a block diagram illustrating an example encoding device 104 that can implement one or more of the techniques described in this disclosure. The encoding device 104 can, for example, generate the syntax structures described herein (e.g., syntax structures of VPS, SPS, PPS, or other syntax elements). The encoding device 104 can perform intra-frame predictive decoding and inter-frame predictive decoding of video blocks within a video slice. As previously described, intra-frame decoding relies at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter-frame decoding relies at least in part on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames in a video sequence. Intra-frame mode (I-mode) can refer to any of several spatial-based compression modes. Inter-frame modes such as one-way prediction (P-mode) or two-way prediction (B-mode) can refer to any of several time-based compression modes.

[0165] The encoding device 104 includes a segmentation unit 35, a prediction processing unit 41, a filter unit 63, an image memory 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an in-frame prediction processing unit 46. For video block reconstruction, the encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62. The filter unit 63 is intended to represent one or more loop filters, such as a deblocking filter, an in-loop filter (ALF), and a sample self-adjusting offset (SAO) filter. Although the filter unit 63 is shown as an in-loop filter in FIG. 17, in other configurations, the filter unit 63 may be implemented as a post-loop filter. The post-processing device 57 may perform additional processing on the encoded video data generated by the encoding device 104. In some cases, the techniques of this disclosure may be implemented by the encoding device 104. However, in other cases, one or more of the techniques disclosed herein may be implemented by the post-processing device 57.

[0166] As shown in FIG. 17, the encoding device 104 receives video data, and the segmentation unit 35 segments the data into video blocks. Such segmentation may also include, for example, segmentation into slices, segments, tiles, or other larger units based on the quadtree structure of LCUs and CUs, as well as video block segmentation. The encoding device 104 is generally shown as the components for encoding video blocks within a video slice to be encoded. The slice may be divided into multiple video blocks (and may be divided into a set of video blocks referred to as tiles). The prediction processing unit 41 may select one of a plurality of possible decoding modes for the current video block based on error results (e.g., decoding rate and distortion level, etc.), such as one of a plurality of intra-frame prediction decoding modes or one of a plurality of inter-frame prediction decoding modes. The prediction processing unit 41 can provide the obtained block decoded within a frame or block decoded between frames to the summer 50 to generate residual block data, and provide it to the summer 62 to reconstruct the encoded block for use as a reference image.

[0167] The in-frame prediction processing unit 46 within the prediction processing unit 41 can perform in-frame prediction decoding of the current video block relative to one or more adjacent blocks in the same frame or slice as the current video block to be decoded, to provide spatial compression. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-frame prediction decoding of the current video block relative to one or more prediction blocks in one or more reference images, to provide temporal compression.

[0168] The motion estimation unit 42 can be configured to determine the inter-frame prediction mode for video slices based on a predetermined pattern for the video sequence. The predetermined pattern can designate video slices in the sequence as P slices, B slices, or GPB slices. The motion estimation unit 42 and the motion compensation unit 44 can be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process of generating motion vectors that estimate the motion of the video block. The motion vectors can, for example, indicate the displacement of the prediction unit (PU) of the video block within the current video frame or picture relative to the prediction block within the reference picture.

[0169] A predicted block is a block found to closely match the PU of the video block to be decoded in terms of primitive difference, which can be determined by the sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some instances, the encoding device 104 can calculate values ​​for sub-integer primitive positions of a reference image stored in the image memory 64. For example, the encoding device 104 can interpolate values ​​for quarter-priority primitive positions, eighth-priority primitive positions, or other fractional primitive positions of the reference image. Thus, the motion estimation unit 42 can perform motion search relative to full primitive positions and fractional primitive positions and output motion vectors with fractional primitive precision.

[0170] The motion estimation unit 42 calculates a motion vector for the PU by comparing the position of the PU in the video block of the inter-frame decoded slice with the position of the predicted block in the reference image. Reference images can be selected from either a first reference image list (list 0) or a second reference image list (list 1), each of which is identified as one or more reference images stored in the image memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.

[0171] Motion compensation performed by the motion compensation unit 44 may involve extracting or generating prediction blocks based on motion vectors determined by motion estimation, possibly interpolating sub-primitive precision. After receiving the motion vector of the PU for the current video block, the motion compensation unit 44 can locate the prediction block pointed to by the motion vector in the reference image list. The encoding device 104 forms a residual video block by subtracting the primitive values ​​of the prediction block from the primitive values ​​of the current video block being decoded, thereby forming a primitive difference. The primitive difference forms residual data for the block and may include both luminance difference components and chrominance difference components. The summer 50 represents one or more components performing the subtraction operation. The motion compensation unit 44 may also generate syntax elements associated with video blocks and video slices for use by the decoding device 112 when decoding video blocks of video slices.

[0172] The intra-frame prediction processing unit 46 can perform intra-frame prediction on the current block as an alternative to the inter-frame prediction performed by the motion estimation unit 42 and the motion compensation unit 44, as described above. Specifically, the intra-frame prediction processing unit 46 can determine the intra-frame prediction mode to be used for encoding the current block. In some instances, the intra-frame prediction processing unit 46 can use various intra-frame prediction modes to encode the current block, for example, during individual encoding paths, and the intra-frame prediction processing unit 46 can select a suitable intra-frame prediction mode from the tested modes. For example, the intra-frame prediction processing unit 46 can use rate-distortion analysis for various tested intra-frame prediction modes to calculate rate-distortion values, and can select the intra-frame prediction mode with the best rate-distortion characteristics from the tested modes. Rate-distortion analysis typically determines the amount of distortion (or error) between the encoded block and the original uncoded block that was encoded to produce the encoded block, as well as the bit rate (i.e., the number of bits) used to produce the encoded block. The in-frame prediction processing unit 46 can calculate the ratio based on the distortion and rate for various encoded blocks to determine which in-frame prediction mode exhibits the optimal rate-distortion value for those blocks.

[0173] In any case, after selecting an intra-frame prediction mode for a block, the intra-frame prediction processing unit 46 may provide information indicating the selected intra-frame prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra-frame prediction mode. The encoding device 104 may include in the transmitted bitstream configuration data definitions for encoding contexts for various blocks, and indications of the most probable intra-frame prediction mode to be used for each of these contexts, an intra-frame prediction mode index table, and a modified intra-frame prediction mode index table. The bitstream configuration data may include a plurality of intra-frame prediction mode index tables and a plurality of modified intra-frame prediction mode index tables (also referred to as encoded character mapping tables).

[0174] After the prediction processing unit 41 generates a prediction block for the current video block via inter-frame prediction or intra-frame prediction, the encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block can be included in one or more TUs and applied to the transform processing unit 52. The transform processing unit 52 uses a transform (such as a discrete cosine transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients. The transform processing unit 52 can transform the residual video data from the primitive domain to the transform domain (such as the frequency domain).

[0175] The transform processing unit 52 may send the obtained transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameters. In some instances, the quantization unit 54 may perform a scan of a matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform this scan.

[0176] After quantization, the entropy coding unit 56 performs entropy coding on the quantized transform coefficients. For example, the entropy coding unit 56 can perform context-adjustable variable-length decoding (CAVLC), context-adjustable binary arithmetic decoding (CABAC), syntax-based context-adjustable binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, or another entropy coding technique. After entropy coding by the entropy coding unit 56, the encoded bitstream can be sent to the decoding device 112, or archived for later transmission or retrieved by the decoding device 112. The entropy coding unit 56 can also perform entropy coding on the motion vectors and other syntax elements used for the current video slice being decoded.

[0177] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual block in the primitive domain for use as a reference block in a reference image later. The motion compensation unit 44 can calculate the reference block by adding the residual block to a predicted block of one of the reference images in the reference image list. The motion compensation unit 44 can also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer primitive values ​​for motion estimation. The summer 62 adds the reconstructed residual block to the motion-compensated predicted block generated by the motion compensation unit 44 to produce a reference block for storage in the image memory 64. The reference block can be used by the motion estimation unit 42 and the motion compensation unit 44 for inter-frame prediction of blocks in subsequent video frames or images.

[0178] Encoding device 104 can perform any of the techniques described herein. Some of the techniques of this disclosure have been generally described with respect to encoding device 104, but as mentioned above, some of the techniques of this disclosure can also be implemented by post-processing device 57.

[0179] The encoding device 104 of FIG17 represents an instance of a video encoder configured to perform one or more transform decoding techniques described herein. The encoding device 104 may perform any of the techniques described herein (including the process described above with respect to FIG18).

[0180] FIG18 is a block diagram illustrating an example decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, a filter unit 91, and an image memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an in-frame prediction processing unit 84. In some instances, the decoding device 112 can perform a decoding path that typically interacts with the encoding path described relative to the encoding device 104 from FIG17.

[0181] During the decoding process, decoding device 112 receives an encoded video bitstream sent by encoding device 104, which represents video blocks of encoded video slices and associated syntax elements. In some instances, decoding device 112 may receive the encoded video bitstream from encoding device 104. In some instances, decoding device 112 may receive the encoded video bitstream from network entity 79 (such as a server, a media-aware network element (MANE), a video editor / stitcher, or other such device configured to implement one or more of the techniques described above). Network entity 79 may or may not include encoding device 104. Before sending the encoded video bitstream to decoding device 112, network entity 79 may implement some of the techniques described in this disclosure. In some video decoding systems, network entity 79 and decoding device 112 may be part of a separate device, while in other cases, the functionality described with respect to network entity 79 may be performed by the same device that includes decoding device 112.

[0182] The entropy decoding unit 80 of the decoding device 112 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 can receive syntax elements at the video slice level and / or video block level. The entropy decoding unit 80 can process and parse both fixed-length and variable-length syntax elements in more parameter sets such as VPS, SPS, and PPS.

[0183] When a video slice is decoded into a slice that has been decoded within a frame (I), the intra-frame prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for video blocks for the current video slice based on the intra-frame prediction mode notified by a signal and data from previously decoded blocks from the current frame or image. When a video frame is decoded into a slice that has been decoded between frames (i.e., B, P, or GPB), the motion compensation unit 82 of the prediction processing unit 81 generates prediction blocks for video blocks for the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit 80. Prediction blocks can be generated from one of the reference images in the reference image list. The decoding device 112 can construct a reference frame list, list 0, and list 1 based on the reference images stored in the image memory 92 using a preset construction technique.

[0184] The motion compensation unit 82 determines prediction information for video blocks in the current video slice by parsing motion vectors and other syntax elements, and uses the prediction information to generate prediction blocks for the current video block being decoded. For example, the motion compensation unit 82 may use one or more syntax elements in the parameter set to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) for decoding video blocks in the video slice, the inter-frame prediction slice type (e.g., B slice, P slice, or GPB slice), construction information for one or more reference picture lists for the slice, motion vectors for each inter-frame encoded video block in the slice, inter-frame prediction state for each inter-frame decoded video block in the slice, and other information for decoding video blocks in the current video slice.

[0185] The motion compensation unit 82 can also perform interpolation based on an interpolation filter. The motion compensation unit 82 can use an interpolation filter, such as that used by the encoding device 104 during the encoding of a video block, to calculate the interpolated value of the sub-integer primitives for the reference block. In the above case, the motion compensation unit 82 can determine the interpolation filter used by the encoding device 104 based on the received syntax elements, and can use the interpolation filter to generate a prediction block.

[0186] The inverse quantization unit 86 inversely quantizes or dequantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include using quantization parameters calculated by the encoding device 104 for each video block in the video slice to determine the degree of quantization and, similarly, the degree of inverse quantization that should be applied. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT or other suitable inverse transform), an inverse integer transform, or a conceptually similar inverse transform process to the transform coefficients to produce residual blocks in the primitive domain.

[0187] After the motion compensation unit 82 generates a prediction block for the current video block based on motion vectors and other syntax elements, the decoding device 112 forms a decoded video block by summing the residual block from the inverse transform processing unit 88 with the corresponding prediction block generated by the motion compensation unit 82. The summer 90 represents one or more components performing this summation operation. If needed, loop filters (in or after the decoding loop) can also be used to smooth primitive transitions or otherwise improve video quality. The filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an auto-adjusting loop filter (ALF), and a sample auto-adjusting offset (SAO) filter. Although the filter unit 91 is shown as an in-loop filter in FIG. 18, in other configurations, the filter unit 91 can be implemented as a post-loop filter. The decoded video block in a given frame or picture is stored in picture memory 92, which stores a reference picture for subsequent motion compensation. Image memory 92 also stores the decoded video for later presentation on a display device, such as video destination device 122 shown in Figure 1.

[0188] The decoding device 112 of FIG18 represents an instance of a video decoder configured to perform one or more of the transform decoding techniques described herein. The decoding device 112 can perform any of the techniques described herein, including process 1900 described above with respect to FIG18.

[0189] In the foregoing description, various aspects of this application have been described with reference to specific examples of this application; however, those skilled in the art will recognize that the subject matter of this application is not limited thereto. Therefore, although illustrative examples of this application have been described in detail herein, it is to be understood that the concepts described herein may be embodied and employed differently in other ways, and the appended claims are intended to be construed as including such variations, except those limited by prior art. Various features and aspects of the subject matter described above may be used individually or collectively. Furthermore, without departing from the broader spirit and scope of this specification, the examples may be used in any number of environments and applications other than those described herein. Therefore, the specification and drawings are to be considered illustrative rather than restrictive. For illustrative purposes, methods have been described in a particular order. It should be understood that, in alternative embodiments, these methods may be performed in a different order than that described.

[0190] Those skilled in the art will understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.

[0191] When a component is described as being "configured" to perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0192] A request language or other language that specifies "at least one" and / or "one or more" of a set indicates that one or more members of the set (in any combination) satisfy the request. For example, a request language specifying "at least one of A and B" means A, B, or A and B. In another instance, a request language specifying "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The specification of "at least one" and / or "one or more" of a language set does not limit the set to items listed in that set. For example, a request language specifying "at least one of A and B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0193] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the examples disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate the interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0194] The techniques described herein can also be implemented using electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices, mobile phones, or integrated circuit devices with multiple uses (including applications in wireless communication devices, mobile phones, and other devices). Any feature described as a module or component can be implemented together in an integrated logic device or separately as an individual but interoperable logic device. If implemented in software, such techniques can be implemented at least in part by a computer-readable data storage medium including program code that, when executed, performs one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging materials. Computer-readable media can include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electronically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Alternatively or additionally, such technologies can be implemented at least in part by computer-readable communication media (such as propagated signals or waves) that carry or transmit program code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer.

[0195] The code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or individual logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, alternatively, the processor may be any known processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein. Additionally, in some cases, the functionality described herein may be provided within a dedicated software or hardware module configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).

[0196] The descriptive nature of this disclosure includes:

[0197] State 1: An apparatus for processing video data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain a first plurality of prediction candidates associated with the video data; generate one or more first groups of prediction candidates by applying a first grouping method to the first plurality of prediction candidates at least in part; reorder the one or more first groups of prediction candidates; generate a first candidate list based on the reordered one or more first groups of prediction candidates; and generate one or more second groups of prediction candidates by applying a second grouping method to the first candidate list at least in part.

[0198] State 2: The apparatus as described in State 1, wherein the at least one processor is configured to: reorder the one or more second sets of prediction candidates; and generate a second candidate list based on the reordered one or more second sets of prediction candidates.

[0199] State 3: The apparatus as described in any one of State 1 or 2, wherein the second candidate list is a merge candidate list for the merge mode.

[0200] State 4: The apparatus as described in any one of states 1 to 3, wherein the first grouping method is based on a plurality of candidate types associated with the first plurality of predicted candidates.

[0201] State 5: The apparatus as described in State 4, wherein, in order to generate the first candidate list, the at least one processor is configured to include one or more first group of predicted candidates in the first candidate list in a predefined order based on the plurality of candidate types.

[0202] State 6: The apparatus as described in any one of States 4 or 5, wherein the plurality of candidate types includes at least one of the following: Spatial Motion Vector Predictor (SMVP) type, Temporal Motion Vector Predictor (TMVP) type, Non-adjacent Temporal Motion Vector Predictor (NA-TMVP) candidate, History-based Motion Vector Predictor (HMVP) candidate, Sub-block-based Temporal Motion Vector Predictor (SbTMVP) candidate, Inherited Affine Merging (I-AffineMVP) candidate, and Constructed Affine Merging (C-AffineMVP) candidate.

[0203] State 7: The apparatus as described in any one of States 2 to 6, wherein the second grouping method is based on a candidate index.

[0204] State 8: The apparatus of any one of states 1 to 7, wherein the at least one processor is configured to reorder the one or more first group of prediction candidates based on cost values.

[0205] State 9: The apparatus as described in State 8, wherein the at least one processor is configured to reorder the one or more first group of prediction candidates by powers of the cost values.

[0206] Style 10: The apparatus as described in any one of Style 8 or 9, wherein the cost values ​​are based on template matching.

[0207] State 11: The apparatus of any one of states 1 to 10, wherein the at least one processor is configured to discard at least one candidate from one or more first groups of predicted candidates before generating the first candidate list.

[0208] State 12: The apparatus of any one of states 2 to 11, wherein the at least one processor is configured to generate a prediction of the current block for the video data based on the second candidate list.

[0209] State 13: The apparatus as described in State 12, wherein the at least one processor is configured to decode the current block of the video data based on the prediction.

[0210] Sample 14: The apparatus as described in Sample 12, wherein the at least one processor is configured to encode the current block of the video data based on the prediction.

[0211] Session 15: The apparatus of any one of Symposia 1 to 14 also includes: a display device coupled to the at least one processor and configured to display an image from the video data.

[0212] Style 16: The apparatus of any one of Styles 1 to 15 also includes: one or more wireless interfaces coupled to the at least one processor, the one or more wireless interfaces including one or more baseband processors and one or more transceivers.

[0213] Style 17: The apparatus of any one of Styles 1 to 16 also includes: at least one camera configured to capture one or more frames associated with the video data.

[0214] Sample 18: A method for decoding video data, the method comprising: obtaining a first plurality of prediction candidates associated with the video data; generating one or more first groups of prediction candidates by applying a first grouping method to the first plurality of prediction candidates at least in part; reordering the one or more first groups of prediction candidates; generating a first candidate list based on the reordered one or more first groups of prediction candidates; and generating one or more second groups of prediction candidates by applying a second grouping method to the first candidate list at least in part.

[0215] State 19: The method as described in State 18 also includes: reordering the one or more second group of prediction candidates; and generating a second candidate list based on the reordered one or more second group of prediction candidates.

[0216] State 20: The method as described in any one of States 18 or 19, wherein the second candidate list is a merge candidate list for the merge mode.

[0217] Style 21: The method of any one of styles 18 to 20, wherein the first grouping method is based on a plurality of candidate types associated with the first plurality of predicted candidates.

[0218] State 22: The method as described in State 21, wherein generating the first candidate list includes adding one or more first group of predicted candidates to the first candidate list in a predefined order based on the plurality of candidate types.

[0219] State 23: The method as described in any one of States 21 or 22, wherein the plurality of candidate types includes at least one of the following: Spatial Motion Vector Predictor (SMVP) type, Temporal Motion Vector Predictor (TMVP) type, Non-adjacent Temporal Motion Vector Predictor (NA-TMVP) candidate, History-based Motion Vector Predictor (HMVP) candidate, Sub-block-based Temporal Motion Vector Predictor (SbTMVP) candidate, Inherited Affine Merging (I-AffineMVP) candidate, and Constructed Affine Merging (C-AffineMVP) candidate.

[0220] Style 24: The method of any one of styles 19 to 23, wherein the second grouping method is based on the candidate index.

[0221] Style 25: The method of any one of styles 18 to 24, wherein the one or more first group of prediction candidates are reordered based on cost values.

[0222] State 26: The method as described in State 25, wherein the one or more first group of prediction candidates are reordered based on the power of the cost values.

[0223] Style 27: The method as described in any one of Style 25 or 26, wherein such cost values ​​are based on template matching.

[0224] Style 28: The method of any one of styles 18 to 27 also includes: discarding at least one candidate from one or more first group of predicted candidates before generating the first candidate list.

[0225] Style 29: The method of any one of styles 19 to 28 also includes: generating a prediction of the current block for the video data based on the second candidate list.

[0226] State 30: The method described in the same way as in State 29 also includes: decoding the current block of the video data based on the prediction.

[0227] State 31: The method described in State 29 also includes: encoding the current block of the video data based on the prediction.

[0228] Speech 32: A non-transitory computer-readable storage medium including instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform any one of the specifications 1 to 31.

[0229] Format 33: An apparatus for processing video data, comprising one or more components for performing the operations described in any one of Formats 1 to 31.

[0230] Speech 32: A non-transitory computer-readable storage medium including instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform the operation as described in any one of Specifications 1 to 31.

[0231] Style 33: An apparatus for processing video data, comprising one or more components for performing the operations described in any one of Styles 1 to 31.

[0232] Sample 34: An apparatus for processing video data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain a first plurality of prediction candidates associated with the video data; determine a first group of prediction candidates by at least partially applying a first grouping method to the first plurality of prediction candidates; reorder the first group of prediction candidates; select a first merged candidate from the reordered first group of prediction candidates; and add the first merged candidate to a candidate list.

[0233] State 35: The apparatus as claimed in claim 34, wherein the at least one processor is configured to determine a second group of predicted candidates at least in part by applying a second grouping method to the candidate list.

[0234] State 36: The apparatus as claimed in claim 35, wherein the at least one processor is configured to: reorder the second set of prediction candidates; select a second merged candidate from the reordered second set of prediction candidates; and add the second merged candidate to the candidate list.

[0235] State 37: The apparatus as described in claim 34, wherein the candidate list is a merge candidate list for the merge mode.

[0236] Format 38: The apparatus as claimed in claim 34, wherein the at least one processor is configured to: determine that the first merge candidate is not in the candidate list; and add the first merge candidate to the candidate list based on the determination that the first merge candidate is not in the candidate list.

[0237] Specimen 39: The apparatus as claimed in claim 34, wherein the first grouping method is based on a plurality of candidate types associated with the first plurality of predicted candidates.

[0238] State 40: The apparatus as claimed in claim 39, wherein the candidate list includes the first merged candidate in a predefined order based on the plurality of candidate types.

[0239] State 41: The apparatus as claimed in claim 40, wherein the plurality of candidate types includes at least one of the following: Spatial Motion Vector Predictor (SMVP) type, Temporal Motion Vector Predictor (TMVP) type, Non-adjacent Temporal Motion Vector Predictor (NA-TMVP) candidate, History-based Motion Vector Predictor (HMVP) candidate, Sub-block-based Temporal Motion Vector Predictor (SbTMVP) candidate, Inherited Affine Merging (I-AffineMVP) candidate, or Constructed Affine Merging (C-AffineMVP) candidate.

[0240] State 42: The apparatus as described in claim 34, wherein the first grouping method is one of a time motion vector predictor (TMVP) or a non-adjacent time motion vector predictor (NA-TMVP).

[0241] Format 43: The apparatus as claimed in claim 34, wherein the at least one processor is configured to reorder the first set of prediction candidates based on cost values.

[0242] Format 44: The apparatus as claimed in claim 43, wherein the at least one processor is configured to reorder the first set of prediction candidates by powers of the cost values.

[0243] Style 45: The apparatus as described in claim 43, wherein such cost values ​​are based on template matching.

[0244] Format 46: The apparatus as claimed in claim 34, wherein the at least one processor is configured to discard at least one candidate from the reordered first set of predicted candidates before adding the first merged candidate to the candidate list.

[0245] Format 47: The apparatus as claimed in claim 34, wherein the at least one processor is configured to: generate a prediction of the current block for the video data based on the candidate list.

[0246] Format 48: The apparatus as claimed in claim 47, wherein the at least one processor is configured to decode the current block of the video data based on the prediction.

[0247] Format 49: The apparatus as claimed in claim 47, wherein the at least one processor is configured to encode the current block of the video data based on the prediction.

[0248] Format 50: The apparatus as claimed in claim 34 also includes: a display device coupled to the at least one processor and configured to display an image from the video data.

[0249] Style 51: The apparatus as described in claim 34 also includes: one or more wireless interfaces coupled to the at least one processor, the one or more wireless interfaces including one or more baseband processors and one or more transceivers.

[0250] Style 52: The apparatus as described in claim 34 also includes: at least one camera configured to capture one or more frames associated with the video data.

[0251] State 53: The apparatus as claimed in claim 34, wherein the first set of prediction candidates includes fewer prediction candidates than the first plurality of prediction candidates.

[0252] State 54: A method for decoding video data, the method comprising: obtaining a first plurality of prediction candidates associated with the video data; determining a first group of prediction candidates by at least partly applying a first grouping method to the first plurality of prediction candidates; reordering the first group of prediction candidates; selecting a first merged candidate from the reordered first group of prediction candidates; and adding the first merged candidate to a candidate list.

[0253] State 55: The method described in claim 54 also includes: determining a second group of predicted candidates by at least partly applying a second grouping method to the candidate list.

[0254] State 56: The method of claim 55 also includes: reordering the second group of prediction candidates; selecting a second merged candidate from the reordered second group of prediction candidates; and adding the second merged candidate to the candidate list.

[0255] State 57: The method as described in request 54, wherein the candidate list is a merge candidate list for the merge mode.

[0256] State 58: The method as described in request 54 also includes: determining that the first merge candidate is not in the candidate list; and adding the first merge candidate to the candidate list based on the determination that the first merge candidate is not in the candidate list.

[0257] State 59: The method as described in request 54, wherein the first grouping method is based on a plurality of candidate types associated with the first plurality of predicted candidates.

[0258] State 60: The method as described in request 59, wherein the candidate list includes the first merged candidate in a predefined order based on the plurality of candidate types.

[0259] State 61: The method as described in request 60, wherein the plurality of candidate types includes at least one of the following: Spatial Motion Vector Predictor (SMVP) type, Temporal Motion Vector Predictor (TMVP) type, Non-adjacent Temporal Motion Vector Predictor (NA-TMVP) candidate, History-based Motion Vector Predictor (HMVP) candidate, Sub-block-based Temporal Motion Vector Predictor (SbTMVP) candidate, Inherited Affine Merging (I-AffineMVP) candidate, or Constructed Affine Merging (C-AffineMVP) candidate.

[0260] State 62: The method as described in request 54 also includes: reordering the first group of prediction candidates based on cost values.

[0261] State 63: The method as described in request 62, wherein the at least one processor is configured to reorder the first group of prediction candidates by powers of the cost values.

[0262] Style 64: The method described in request 62, wherein such cost values ​​are based on template matching.

[0263] State 65: The method of claim 54 also includes: discarding at least one candidate from the reordered first group of predicted candidates before adding the first merged candidate to the candidate list.

[0264] State 66: The method described in request 54 also includes: generating a prediction of the current block for the video data based on the candidate list.

[0265] State 67: The method as described in claim 66, wherein the at least one processor is configured to decode the current block of the video data based on the prediction.

[0266] State 68: The method as described in request 66, wherein the at least one processor is configured to encode the current block of the video data based on the prediction.

[0267] State 69: The method described in claim 54 also includes: displaying an image from the video data.

[0268] State 70: The method described in claim 54 also includes: capturing one or more frames associated with the video data.

[0269] State 71: The method as described in claim 54, wherein the first set of prediction candidates includes fewer prediction candidates than the first plurality of prediction candidates.

[0270] Sample 72: A non-transitory computer-readable medium having instructions stored thereon, which, when executed by at least one processor, cause the at least one processor to perform the following operations: obtain a first plurality of prediction candidates associated with video data; determine a first group of prediction candidates by at least partly applying a first grouping method to the first plurality of prediction candidates; reorder the first group of prediction candidates; select a first merged candidate from the reordered first group of prediction candidates; and add the first merged candidate to a candidate list.

[0271] State 73: Non-transitory computer-readable medium as described in claim 72, wherein the instructions also cause the at least one processor to perform the following operation: determine a second group of predicted candidates by applying a second grouping method to the candidate list, at least in part.

[0272] State 74: Non-transitory computer-readable medium as claimed in claim 73, wherein the instructions also cause the at least one processor to perform the following operations: reorder the second set of prediction candidates; select a second merged candidate from the reordered second set of prediction candidates; and add the second merged candidate to the candidate list.

[0273] Style 75: Non-transitory computer-readable media as described in claim 72, wherein the candidate list is a merge candidate list for the merge mode.

[0274] State 76: Non-transitory computer-readable medium as described in claim 72, wherein the instructions also cause the at least one processor to: determine that the first merge candidate is not in the candidate list; and add the first merge candidate to the candidate list based on the determination that the first merge candidate is not in the candidate list.

[0275] State 77: Non-transitory computer-readable media as described in claim 72, wherein the first grouping method is based on a plurality of candidate types associated with the first plurality of predicted candidates.

[0276] Format 78: Non-transitory computer-readable media as described in claim 77, wherein the candidate list includes the first merged candidate in a predefined order based on the plurality of candidate types.

[0277] State 79: Non-transitory computer-readable media as described in claim 78, wherein the plurality of candidate types includes at least one of the following: Spatial Motion Vector Predictor (SMVP) type, Temporal Motion Vector Predictor (TMVP) type, Non-adjacent Temporal Motion Vector Predictor (NA-TMVP) candidate, History-based Motion Vector Predictor (HMVP) candidate, Sub-block-based Temporal Motion Vector Predictor (SbTMVP) candidate, Inherited Affine Merging (I-AffineMVP) candidate, or Constructed Affine Merging (C-AffineMVP) candidate.

[0278] Sample 80: Non-transitory computer-readable medium as described in claim 72, wherein the first grouping method is one of a time motion vector predictor (TMVP) or a non-adjacent time motion vector predictor (NA-TMVP).

[0279] State 81: Non-transitory computer-readable medium as described in claim 72, wherein the instructions also cause the at least one processor to perform the following operation: reorder the first group of prediction candidates based on cost values.

[0280] Format 82: Non-transitory computer-readable medium as described in claim 81, wherein the instructions also cause the at least one processor to perform the following operation: reorder the first group of prediction candidates based on cost values.

[0281] State 83: Non-transitory computer-readable medium as described in claim 81, wherein the instructions also cause the at least one processor to reorder the first group of prediction candidates by exponentiation of the cost values.

[0282] State 84: Non-transitory computer-readable medium as described in claim 72, wherein the instructions also cause the at least one processor to perform the following operation: discard at least one candidate from the reordered first group of predicted candidates before adding the first merge candidate to the candidate list.

[0283] State 85: Non-transitory computer-readable media as described in claim 72, wherein the instructions also cause the at least one processor to perform the following operation: generate a prediction of the current block for the video data based on the candidate list.

[0284] Format 86: Non-transitory computer-readable media as described in claim 86, wherein the instructions also cause the at least one processor to perform the following operation: decode the current block of the video data based on the prediction.

[0285] Format 87: Non-transitory computer-readable media as described in claim 86, wherein the instructions also cause the at least one processor to encode the current block of the video data based on the prediction.

[0286] State 88: Non-transitory computer-readable media as described in claim 72, wherein the instructions also cause the at least one processor to perform the following operation: display an image from the video data.

[0287] Format 89: Non-transitory computer-readable media as described in claim 72, wherein the instructions also cause the at least one processor to perform the following operation: capture one or more frames associated with the video data.

[0288] Style 90: Non-transitory computer-readable medium as described in claim 72, wherein the first set of prediction candidates includes fewer prediction candidates than the first plurality of prediction candidates. [Simplified Explanation of the Diagram]

[0011] The following describes examples of various implementation methods in detail with reference to the attached drawings:

[0012] Figure 1 is a block diagram illustrating encoding and decoding devices according to some examples;

[0013] Figure 2 is a diagram illustrating an example of the location of the spatial motion vector (MV) candidate according to the example described herein;

[0014] Figure 3 is a diagram illustrating an example of motion vector scaling for time merging candidates according to the example described herein;

[0015] Figure 4 is a diagram illustrating an example of candidate positions for time merging candidates according to the example described herein;

[0016] Figure 5 is a diagram illustrating an example of deriving spatially adjacent blocks for spatial merging candidates according to the example described herein;

[0017] Figure 6A is a diagram illustrating an example of spatially adjacent blocks used by sub-block-based temporal motion vector prediction (SbTMVP) in VVC according to the example described herein;

[0018] Figure 6B is a diagram illustrating an example of deriving the motion field of a sub-CU by applying motion shifts from spatial neighbors and scaling motion information from the corresponding co-located sub-decoding unit (CU) according to the example described herein;

[0019] Figure 7 is a diagram showing an example of the positions of candidate locations for constructing an affine merging pattern according to the example described herein;

[0020] Figure 8 is a diagram showing an example of a template and an example of a reference sample of the template in the example described herein;

[0021] Figure 9 is a diagram illustrating an example of a template and a reference sample of a template for a block having sub-block motion using motion information of the current block according to the example described herein;

[0022] Figure 10 is a diagram illustrating an example of dividing a block into two parts by a geometrically positioned straight line, according to the example described herein;

[0023] Figure 11 is a diagram illustrating an example derivation of a single prediction candidate list for a geometric segmentation pattern (GEO pattern) based on the example described herein;

[0024] Figure 12 is a diagram illustrating an example of the location of a temporal motion vector prediction (TMVP) candidate based on the example described herein;

[0025] Figure 13 is a diagram illustrating an example of non-adjacent spatial adjacent blocks for deriving a non-adjacent (NA) motion vector predictor (MVP) according to the example described herein;

[0026] Figure 14 is a diagram illustrating an example of a TMVP from non-adjacent co-located blocks (e.g., C2, C3 to C11) according to the example described herein;

[0027] Figure 15 is a block diagram illustrating an example of a multi-level ARMC of various forms according to this disclosure;

[0028] Figure 16 is a flowchart illustrating various techniques for performing bit rate estimation according to the present disclosure;

[0029] Figure 17 is a block diagram illustrating a video encoding device according to some examples; and

[0030] Figure 18 is a block diagram showing a video decoding device according to some examples. [Biomaterial Storage]

[0290] Domestic storage information (please note in order of storage institution, date, and number): None. International storage information (please note in order of storage country, institution, date, and number): None.

Claims

1. An apparatus for processing video data, the apparatus comprising: At least one memory cell; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain a plurality of prediction candidates associated with video data, the plurality of prediction candidates including at least a first group of prediction candidates and a second group of prediction candidates, the first group of prediction candidates including N1 prediction candidates, the second group of prediction candidates including N2 prediction candidates, wherein N1 and N2 are positive integers, and the N1 prediction candidates are different from the N2 prediction candidates; determine the first group of prediction candidates including the N1 prediction candidates from the plurality of prediction candidates; reorder the N1 prediction candidates included in the first group of prediction candidates based on a corresponding cost value associated with each of the N1 prediction candidates to determine a first group of reordered prediction candidates; select a first set of prediction candidates including M1 prediction candidates from the first group of reordered prediction candidates, wherein M1 is a positive integer less than N1; and add only the first set of prediction candidates from the first group of reordered prediction candidates to a candidate list.

2. The apparatus of claim 1, wherein the at least one processor is configured to: determine, from the plurality of prediction candidates, the second set of prediction candidates including the N2 prediction candidates.

3. The apparatus of claim 2, wherein the at least one processor is configured to: reorder the N2 prediction candidates in the second set of prediction candidates to determine a second set of reordered prediction candidates; select a second set of prediction candidates including M2 prediction candidates from the second set of reordered prediction candidates, where M2 is a positive integer; and add only the second set of prediction candidates from the second set of reordered prediction candidates to the candidate list.

4. The apparatus as claimed in claim 1, wherein the candidate list is a merge candidate list for a merge mode.

5. The apparatus of claim 3, wherein the at least one processor is configured to determine the first set of prediction candidates and the second set of prediction candidates based on a plurality of candidate types associated with the first plurality of prediction candidates.

6. The apparatus of claim 5, wherein the first set of prediction candidates consists of temporal motion vector predictor (TMVP) candidates, the second set of prediction candidates consists of non-adjacent temporal motion vector predictor (NA-TMVP) candidates, and wherein the at least one processor is configured to add the first set of prediction candidates to the candidate list before adding the second set of prediction candidates to the candidate list.

7. The apparatus of claim 1, wherein the at least one processor is configured to reorder the N1 prediction candidates by exponentiation of the corresponding cost value of each of the N1 prediction candidates.

8. The apparatus as claimed in claim 7, wherein each corresponding cost value is determined based on template matching.

9. The apparatus of claim 1, wherein the at least one processor is configured to: generate a prediction for a current block of the video data based on the candidate list.

10. The apparatus of claim 9, wherein the at least one processor is configured to: decode the current block of the video data based on the prediction.

11. The apparatus of claim 9, wherein the at least one processor is configured to: encode the current block of the video data based on the prediction.

12. The apparatus as described in claim 1, also comprising: A display device coupled to the at least one processor and configured to display an image from the video data.

13. The apparatus as described in claim 1, also comprising: One or more wireless interfaces coupled to the at least one processor, the one or more wireless interfaces including one or more baseband processors and one or more transceivers.

14. The apparatus as described in claim 1, also comprising: At least one camera is configured to capture one or more frames associated with the video data.

15. The apparatus of claim 1, wherein the first set of prediction candidates includes fewer prediction candidates than the plurality of prediction candidates.

16. A method for decoding video data, the method comprising the steps of: obtaining a plurality of prediction candidates associated with the video data, the plurality of prediction candidates including at least a first group of prediction candidates and a second group of prediction candidates, the first group of prediction candidates including N1 prediction candidates, the second group of prediction candidates including N2 prediction candidates, wherein N1 and N2 are positive integers, and the N1 prediction candidates are different from the N2 prediction candidates; determining the first group of prediction candidates including the N1 prediction candidates from the plurality of prediction candidates; reordering the N1 prediction candidates included in the first group of prediction candidates based on a corresponding cost value associated with each of the N1 prediction candidates to determine a first group of reordered prediction candidates; selecting a first set of prediction candidates including M1 prediction candidates from the first group of reordered prediction candidates, wherein M1 is a positive integer less than N1; and adding only the first set of prediction candidates from the first group of reordered prediction candidates to a candidate list.

17. The method of claim 16 also includes the step of: determining the second set of prediction candidates, which includes the N2 prediction candidates, from the plurality of prediction candidates.

18. The method of claim 17 further includes the steps of: reordering the N2 prediction candidates in the second group of prediction candidates to determine the second group of reordered prediction candidates; selecting a second set of prediction candidates including M2 prediction candidates from the second group of reordered prediction candidates, where M2 is a positive integer; and adding only the second set of prediction candidates from the second group of reordered prediction candidates to the candidate list.

19. The method as described in claim 16, wherein the candidate list is a merge candidate list for a merge mode.

20. The method of claim 18 further includes the step of: determining the first group of prediction candidates and the second group of prediction candidates based on the plurality of candidate types associated with the plurality of prediction candidates.

21. The method of claim 20, wherein the first set of prediction candidates consists of temporal motion vector predictor (TMVP) candidates, the second set of prediction candidates consists of non-adjacent temporal motion vector predictor (NA-TMVP) candidates, and the method further comprises the step of adding the first set of prediction candidates to the candidate list before adding the second set of prediction candidates to the candidate list.

22. The method of claim 16 also includes the step of: reordering the N1 prediction candidates by exponentiation of the corresponding cost value of each of the N1 prediction candidates.

23. The method as described in request item 16, wherein each corresponding cost value is determined based on template matching.

24. The method of claim 16 also includes the step of: generating a prediction for a current block of the video data based on the candidate list.

25. The method of claim 16, wherein the first set of prediction candidates includes fewer prediction candidates than the plurality of prediction candidates.

Citation Information

Patent Citations

  • Method and Apparatus of Reordering Motion Vector Prediction Candidate Set for Video Coding

    US20200068218A1

  • Method and apparatus for further improved context design for prediction mode and coded block flag (CBF)

    US20200186792A1