History-based motion vector prediction

By employing a motion vector prediction method based on affine motion patterns in video coding, the motion vector at a predetermined position is estimated using the control point motion vectors of the block. This extends the historical motion vector predictor table, solves the problem of insufficient coding efficiency in existing technologies, and achieves more efficient video data compression.

CN114503584BActive Publication Date: 2026-05-08QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2020-09-23
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing video coding technologies struggle to effectively utilize affine motion patterns to extend historical motion vector predictions when compressing video data, resulting in insufficient coding efficiency and failing to meet the demands of high-quality video services.

Method used

A motion vector prediction method based on affine motion patterns is adopted. By determining the motion vector of the control point of the block, the motion vector at the predetermined position is estimated, and a historical motion vector predictor table is used for prediction. This expands the types of motion information to improve coding efficiency.

Benefits of technology

It improves the efficiency and performance of video encoding, better meets the needs of high-quality video services, and reduces the burden of video data processing and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114503584B_ABST
    Figure CN114503584B_ABST
Patent Text Reader

Abstract

Systems, methods, and computer-readable media for updating a history-based motion vector table are provided. In some examples, a method can include obtaining one or more blocks of video data, determining a first motion vector derived from a first control point of a block of the one or more blocks, the block being encoded using an affine mode of motion, determining a second motion vector derived from a second control point of the block, estimating a third motion vector for a predetermined location within the block based on the first motion vector and the second motion vector, and populating a history-based motion vector predictor (HMVP) table with the third motion vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In summary, this disclosure relates to video encoding and compression, and more specifically, this disclosure relates to historical motion vector prediction. Background Technology

[0002] Many devices and systems allow video data to be processed and output for consumption. Digital video data comprises vast amounts of data to meet the needs of consumers and video providers. For example, consumers of video data expect the highest quality video with high fidelity, resolution, frame rate, and so on. As a result, the large amount of video data required to meet these needs places a burden on the communication networks and equipment that process and store video data.

[0003] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include Basic Video Coding (EVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), MPEG-2 Part 2 Coding (MPEG stands for Moving Picture Experts Group), VP9, ​​and Open Media Consortium (AOMedia) Video 1 (AV1), etc. Video coding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that take advantage of redundancy present in video images or sequences. Generally, the goal of video coding techniques is to compress video data to a form using a lower bitrate while avoiding or minimizing video quality degradation. As more and more video services become available, there is a need for coding techniques with better coding efficiency and performance. Summary of the Invention

[0004] Systems, methods, and computer-readable media for providing history-based motion vector prediction are disclosed. According to at least one example, a method for history-based motion vector prediction is provided. The method may include: obtaining one or more blocks of video data; determining a first motion vector derived from a first control point of a block in the one or more blocks, the block being encoded using an affine motion pattern; determining a second motion vector derived from a second control point of the block; estimating a third motion vector for a predetermined position within the block based on the first motion vector and the second motion vector; and populating a history-based motion vector predictor (HMVP) table using the third motion vector.

[0005] According to at least one example, an apparatus for history-based motion vector prediction is provided. In some examples, the apparatus may include: a memory; and one or more processors coupled to the memory, the one or more processors being configured to: acquire one or more blocks of video data; determine a first motion vector derived from a first control point of a block in the one or more blocks, the block being encoded using an affine motion pattern; determine a second motion vector derived from a second control point of the block; estimate a third motion vector for a predetermined position within the block based on the first motion vector and the second motion vector; and populate a history-based motion vector predictor (HMVP) table using the third motion vector.

[0006] According to at least one example, a non-transitory computer-readable storage medium is provided for history-based motion vector prediction. The non-transitory computer-readable medium may include instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: obtain one or more blocks of video data; determine a first motion vector derived from a first control point of a block in the one or more blocks, the block being encoded using an affine motion pattern; determine a second motion vector derived from a second control point of the block; estimate a third motion vector for a predetermined position within the block based on the first motion vector and the second motion vector; and populate a history-based motion vector predictor (HMVP) table using the third motion vector.

[0007] According to at least one example, an apparatus is provided for generating a fuzzy control interface for history-based motion vector prediction. The apparatus may include units for performing the following operations: acquiring one or more blocks of video data; determining a first motion vector derived from a first control point of one or more blocks, the blocks being encoded using an affine motion pattern; determining a second motion vector derived from a second control point of the blocks; estimating a third motion vector for a predetermined position within the blocks based on the first and second motion vectors; and populating a history-based motion vector predictor (HMVP) table using the third motion vector.

[0008] In some examples, the methods, computer-readable media, and apparatus described above may add one or more HMVP candidates from the HMVP table to at least one of an advanced motion vector prediction (AMVP) candidate list, a merged pattern candidate list, and a motion vector prediction predictor for encoding using the affine motion pattern.

[0009] In some examples, the predetermined position may include the center of the block. In some examples, the first control point may include the upper left control point, and the second control point may include the upper right control point. In some cases, the third motion vector may be further estimated based on the control point motion vector associated with the bottom control point.

[0010] In some cases, estimating the third motion vector may include: determining the rate of change between the first and second motion vectors based on the difference between the first and second motion vectors; and multiplying the rate of change by a multiplication factor corresponding to the predetermined position. In some examples, the predetermined position may include the center of the block, and the multiplication factor may include half of the width and / or height of the block. In some examples, the rate of change may include a rate of change per unit, and each unit of the rate of change per unit may include a sample, a sub-block, and / or a pixel. In some cases, the multiplication factor may include the number of samples between the predetermined position and the boundary of the block.

[0011] In some examples, the third motion vector is based on a first motion change between the horizontal components of the first motion vector and the horizontal components of the second motion vector, and a second motion change between the vertical components of the first motion vector and the vertical components of the second motion vector.

[0012] In some examples, the third motion vector may include a translational motion vector generated based on affine motion information associated with the block. In some cases, the third motion vector may include motion information associated with one or more sub-blocks of the block, and at least one of the one or more sub-blocks corresponds to the predetermined position.

[0013] In some cases, the HMVP table and / or the third motion vector can be used for motion prediction of additional blocks.

[0014] In some aspects, each of the above-described devices is or includes a camera, a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, an autonomous vehicle, an encoder, a decoder, or other device. In some aspects, the device includes one or more cameras for capturing one or more videos and / or images. In some aspects, the device also includes a display for displaying one or more videos and / or images. In some aspects, the above-described device may include one or more sensors.

[0015] This invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to define the scope of the claimed subject matter. The subject matter should be understood by referring to appropriate portions of the entire specification, any or all of the drawings, and each claim.

[0016] The foregoing, as well as other features and embodiments, will become more apparent upon reference to the following description, claims, and drawings. Attached Figure Description

[0017] To describe in detail the various advantages and features of this disclosure that can be obtained therein, a more specific description of the above principles will be provided with reference to specific embodiments illustrated in the accompanying drawings. It is to be understood that these drawings depict only exemplary embodiments of this disclosure and are not intended to limit its scope. The principles herein are described and explained with additional specificity and detail through the use of the drawings, in which:

[0018] Figure 1 This is a block diagram illustrating examples of encoding and decoding devices based on some examples;

[0019] Figure 2A This is a conceptual diagram illustrating exemplary spatially adjacent motion vector candidates for merging patterns, based on some examples;

[0020] Figure 2B This is a conceptual diagram illustrating exemplary spatially adjacent motion vector candidates for an advanced motion vector prediction (AMVP) pattern, based on some examples.

[0021] Figure 3A This is a conceptual diagram illustrating exemplary Time Motion Vector Predictor (TMVP) candidates based on some examples;

[0022] Figure 3B This is a conceptual diagram illustrating an example of scaling motion vectors based on some examples;

[0023] Figure 4 This is a diagram illustrating an exemplary history-based motion vector predictor (HMVP) table based on some examples;

[0024] Figure 5 This is a diagram illustrating examples of extracting non-adjacent space merging candidates based on some examples;

[0025] Figure 6A This is a diagram illustrating examples of spatial and temporal locations used in MVP prediction, based on several examples.

[0026] Figure 6B This is a diagram illustrating an example of the access order for a spatial MVP (S-MVP) based on some examples;

[0027] Figure 6C This illustrates spatial inverse pattern substitution based on some examples (with) Figure 6B The order of the examples in the diagram is compared to the order of the examples in the diagram.

[0028] Figure 7 This is a diagram illustrating an example of a simplified affine motion model for the current block, based on some examples;

[0029] Figure 8 This is a diagram illustrating an example of the motion vector field of sub-blocks of a block, based on some examples;

[0030] Figure 9 This is a diagram illustrating examples of motion vector prediction in affine inter-frame (AF_INTER) mode based on some examples;

[0031] Figure 10A and Figure 10B This is a diagram illustrating examples of motion vector prediction in affine merging (AF_MERGE) mode based on some examples;

[0032] Figure 11 This is a diagram illustrating an example of an affine motion model for the current block, based on some examples;

[0033] Figure 12 This is a diagram illustrating another example of an affine motion model for the current block, based on some examples;

[0034] Figure 13 This is a diagram showing examples of the current block and candidate blocks based on some examples;

[0035] Figure 14 This is a diagram illustrating examples of the current block, the control point of the current block, and candidate blocks, based on some examples;

[0036] Figure 15 This is a diagram showing exemplary motion vectors estimated for an affine coded block and corresponding to the center position of the coded block, based on some examples;

[0037] Figure 16 This is a flowchart illustrating an exemplary process for updating a history-based motion prediction table using motion information generated from affine coded blocks, based on some examples.

[0038] Figure 17 This is a block diagram illustrating an exemplary encoding device according to some examples; and

[0039] Figure 18 This is a block diagram illustrating an exemplary video decoding device according to some examples. Detailed Implementation

[0040] Certain aspects and embodiments of this disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments can be applied independently, and some can be applied in combination. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments may be practiced without these specific details. The accompanying drawings and description are not intended to be limiting.

[0041] The following description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the subsequent description of these exemplary embodiments will provide those skilled in the art with a feasible description for implementing the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0042] Video encoding devices implement video compression techniques to efficiently encode and decode video data. Video compression techniques may include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques to reduce or remove redundancy inherent in a video sequence. A video encoder can divide each frame of the original video sequence into multiple rectangular regions, which are called video blocks or coding units (described in more detail below). These video blocks can be encoded using specific prediction modes.

[0043] Video blocks can be divided into one or more smaller blocks in one or more ways. Blocks may include coded tree blocks, prediction blocks, transform blocks, and / or other suitable blocks. Unless otherwise specified, the reference to “block” generally refers to such a video block (e.g., coded tree block, coded block, prediction block, transform block, or other suitable block or sub-block, as will be understood by one of ordinary skill in the art). Furthermore, each of these blocks may also be interchangeably referred to herein as a “unit” (e.g., coded tree unit (CTU), coded unit, prediction unit (PU), transform unit (TU), etc.). In some cases, a unit may refer to a coded logic unit encoded in the bitstream, while a block may refer to a portion of the process targeted in the video frame buffer.

[0044] For inter-frame prediction mode, the video encoder searches for blocks similar to the encoded block in a frame (or picture) located at another time position (called a reference frame or reference picture). The video encoder can restrict the search to a certain spatial displacement from the block to be encoded. The best match can be located using two-dimensional (2D) motion vectors that include horizontal and vertical displacement components. For intra-frame prediction mode, the video encoder can use spatial prediction techniques to form a prediction block based on data from previously encoded adjacent blocks within the same picture.

[0045] Video encoders can determine prediction errors. For example, a prediction can be determined as the difference between the pixel values ​​in the block being encoded and the predicted block. Prediction errors can also be referred to as residuals. Video encoders can also apply transform coding (e.g., using the form of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), or other suitable transforms) to the prediction error to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and motion vectors can be represented using syntax elements and, together with control information, form the coded representation of the video sequence. In some cases, video encoders can entropy-encode the syntax elements, further reducing the number of bits required for their representation.

[0046] The video decoder can use the syntax elements and control information discussed above to construct prediction data (e.g., prediction blocks) for decoding the current frame. For example, the video decoder can add the prediction blocks to the compressed prediction error. The video decoder can determine the compressed prediction error by weighting the transform basis function using quantization coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.

[0047] As described in more detail below, this document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media for using affine motion information for historical motion vector prediction. The techniques described herein can be applied to one or more of a variety of block-based video coding techniques (where video is reconstructed on a block-by-block basis). For example, the techniques described herein can be applied to any existing video codec (e.g., High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), or other suitable existing video codecs), and / or can be efficient coding tools for any video coding standard under development and / or future video coding standards, such as Basic Video Coding (EVC), Universal Video Coding (VVC), Joint Exploratory Model (JEM), VP9, ​​AV1, and / or other video coding standards under development or to be developed.

[0048] In some examples, the method described in this paper can be used to generate translational motion vectors based on affine coded blocks. An affine motion model can use multiple control points to derive multiple motion vectors for a block. For example, an affine motion model can generate multiple local motion vectors for sub-blocks or pixels of a block. However, in some video coding standards, the HMVP table for History-Based Motion Vector Prediction (HMVP) only includes and / or supports a single translational motion vector per block. Therefore, in some examples, to expand the types of motion information included in the HMVP table, the method in this paper can use affine motion information to approximate the translation vector for the block and include the translation vector in the HMVP table. The translation vector can represent the motion information for the block, which, once included in the HMVP table, can be used for HMVP.

[0049] In some cases, the method described in this paper can generate a motion vector for the center position of a block based on the control point motion vectors associated with the upper left control point and the upper right control point. For example, the method can calculate the rate of change between the horizontal and vertical components of the control point motion vectors associated with the upper left and upper right control points, and generate the motion vector for the center position based on this rate of change. The motion vector can represent motion information for the block and can be included in the HMVP table for future use in motion prediction.

[0050] The techniques described herein will be as follows in the following disclosure. The discussion begins with a description of exemplary systems and techniques for video coding and motion vector prediction, such as... Figures 1 to 15 As shown. Next, a description of an exemplary method for updating a history-based motion prediction table using motion information generated from affine coded blocks will follow, as... Figure 16 As shown. The discussion concludes with a description of an exemplary encoding device architecture and an exemplary decoding device architecture, as illustrated. Figure 17 and 18 As shown. This disclosure now turns to Figure 1 .

[0051] Figure 1This is a block diagram illustrating an example of a system 100 including encoding device 104 and decoding device 112. Encoding device 104 may be part of a source device, and decoding device 112 may be part of a receiving device (also referred to as a client device). The source device and / or receiving device may include electronic devices such as mobile or landline phones (e.g., smartphones, cellular phones, etc.), desktop computers, laptops or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, Internet Protocol (IP) cameras, server devices in server systems (including one or more server devices (e.g., video streaming server systems, or other suitable server systems)), head-mounted displays (HMDs), head-up displays (HUDs), smart glasses (e.g., virtual reality (VR) glasses, augmented reality (AR) glasses, or other smart glasses), or any other suitable electronic devices.

[0052] The components of system 100 may include and / or may be implemented using circuitry or other electronic hardware, which may include one or more programmable circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.

[0053] Although system 100 is shown to include certain components, those skilled in the art will understand that system 100 may include more than [other components]. Figure 1 The components shown may include more or fewer components. For example, in some cases, system 100 may also include one or more memory devices (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices) in addition to memory 108 and storage cell 118; one or more processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices) communicating with and / or electrically connected to one or more memory devices; one or more wireless interfaces for performing wireless communication (e.g., including one or more transceivers and baseband processors for each wireless interface); one or more wired interfaces for performing communication over one or more hardwired connections (e.g., serial interfaces such as Universal Serial Bus (USB) inputs, lighting connectors, and / or other wired interfaces); and / or Figure 1 Other components not shown.

[0054] The encoding techniques described herein are applicable to video encoding in a variety of multimedia applications, including streaming video (e.g., over the Internet), television broadcasting or transmission, encoding digital video for storage on data storage media, decoding digital video stored on data storage media, or other applications. In some examples, system 100 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.

[0055] Encoding device 104 (or encoder) can be used to encode video data using video coding standards or protocols to generate an encoded video bitstream. Examples of video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Video, ITU-T H.262, or ISO / IEC MPEG-2 Video, ITU-T H.263, ISO / IEC MPEG-4 Video, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its Scalable Video Coding (SVC) and Multi-View Video Coding (MVC) extensions), and High Efficiency Video Coding (HEVC) or ITU-T H.265. Various extensions of HEVC exist that handle multi-layer video coding, including range and screen content coding extensions, 3D Video Coding (3D-HEVC), Multi-View Extension (MV-HEVC), and Scalable Extension (SHVC). The ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) Joint Collaborative Working Group on Video Coding (JCT-VC) and the Joint Collaborative Working Group on 3D Video Coding Extensions (JCT-3V) have developed HEVC and its extensions.

[0056] MPEG and ITU-T VCEG also established the Joint Video Exploration Group (JVET) to explore and develop new video coding tools for next-generation video coding standards, which was named Universal Video Coding (VVC). The reference software is called the VVC Test Model (VTM). VVC aims to provide a significant improvement in compression performance compared to the existing HEVC standard to facilitate the deployment of higher-quality video services and emerging applications (e.g., 360° omnidirectional immersive multimedia, high dynamic range (HDR) video, and other examples). Basic Video Coding (EVC), VP9, ​​and the Open Media Consortium (AOMedia) Video 1 (AV1) are other video coding standards to which the technologies described in this paper can be applied.

[0057] Many of the embodiments described herein can be performed using video codecs such as EVC, VTM, VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein are also applicable to other coding standards, such as MPEG, JPEG (or other coding standards for still images), VP9, ​​AV1, their extensions, or other suitable coding standards that are already available or not yet available or under development. Therefore, although the techniques and systems described herein may be described with reference to a specific video coding standard, it will be understood by those skilled in the art that the description should not be construed as applicable only to that particular standard.

[0058] refer to Figure 1 Video source 102 can provide video data to encoding device 104. Video source 102 can be part of a source device, or it can be part of a device other than a source device. Video source 102 can include video capture devices (e.g., cameras, camera phones, video phones, etc.), video archiving units containing stored video, video servers or content providers that provide video data, video feed interfaces that receive video from video servers or content providers, computer graphics systems for generating computer graphics video data, combinations of such sources, or any other suitable video source.

[0059] Video data from video source 102 may include one or more input pictures. Pictures may also be referred to as "frames." A picture or frame is a still image that, in some cases, is part of the video. In some examples, the data from video source 102 may be still images that are not part of the video. In HEVC, VVC, and other video coding standards, a video sequence may include a series of pictures. A picture may include a three-sample array, which is represented as S. L S Cb and S Cr S L It is a two-dimensional array of brightness samples, S Cb It is a two-dimensional array of Cb chromaticity samples, and S Cr It is a two-dimensional array of chromaticity samples. Chromaticity samples may also be referred to as "chroma" samples in this text. In other cases, the image may be monochrome and may consist only of an array of luminance samples.

[0060] Encoding device 104's encoder engine 106 (or encoder) encodes video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more encoded video sequences. An encoded video sequence (CVS) includes a series of access units (AUs) that begin with an AU that has a random access point picture in the base layer and possesses certain attributes, and continue until the next AU that has a random access point picture in the base layer and possesses certain attributes, but does not include that next AU. For example, certain attributes of the random access point picture that begins the CVS may include a RASL flag equal to 1 (e.g., NoRaslOutputFlag). Otherwise, a random access point picture (where the RASL flag is equal to 0) does not begin the CVS. An access unit (AU) includes one or more encoded pictures and control information corresponding to encoded pictures that share the same output time. Encoded slices of pictures are encapsulated into data units at the bitstream level, which are called Network Abstraction Layer (NAL) units. For example, an HEVC video bitstream may include one or more CVSs that include NAL units. Each NAL unit has a NAL unit header. In one example, the header is one byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. The syntax elements in the NAL unit header use specified bits and are therefore visible to all kinds of systems and transport layers, such as transport streams, Real-Time Transport (RTP) protocols, file formats, and others.

[0061] The HEVC standard contains two types of NAL units: Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain encoded picture data that forms the encoded video bitstream. For example, a VCL NAL unit contains a sequence of bits that form the encoded video bitstream. A VCL NAL unit includes a slice or fragment of the encoded picture data (described below), while a non-VCL NAL unit includes control information related to one or more encoded pictures. In some cases, NAL units may be referred to as packets. A HEVC AU includes: VCL NAL units containing encoded picture data, and non-VCL NAL units (if any) corresponding to the encoded picture data. Among other information, non-VCL NAL units may also contain a set of parameters with high-level information related to the encoded video bitstream. For example, the parameter set may include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). Each slice or other portion of the bitstream can reference a single valid PPS, SPS, and VPS to allow the decoding device 112 to access information that can be used to decode the slice or other portion of the bitstream.

[0062] NAL units can contain bit sequences that form an encoded representation of video data (e.g., an encoded video bitstream, a CVS of a bitstream, etc.), such as encoded representations of images in a video. Encoder engine 106 generates encoded representations of images by dividing each image into multiple slices. Each slice is independent of other slices, allowing information within that slice to be encoded without depending on data from other slices within the same image. A slice includes one or more segments, comprising independent segments and (if present) one or more dependent segments that depend on previous segments.

[0063] In HEVC, slices are then divided into code tree blocks (CTBs) for luma and chroma samples. One or more CTBs for luma samples and one CTB for chroma samples, along with the syntax used for the samples, are called code tree units (CTUs). CTUs can also be called "tree blocks" or "maximum coding units" (LCUs). A CTU is the basic processing unit used for HEVC coding. A CTU can be subdivided into multiple coding units (CUs) of different sizes. A CU contains an array of luma and chroma samples called a code block (CB).

[0064] Luminance and chrominance CBs can be further subdivided into prediction blocks (PBs). A PB is a sample block of either the luminance or chrominance component that uses the same motion parameters for inter-frame prediction or intra-block copy (IBC) prediction (when available or enabled). A luminance PB and one or more chrominance PBs, along with their associated syntax, form a prediction unit (PU). For inter-frame prediction, the set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU and used for inter-frame prediction of the luminance PB and one or more chrominance PBs. Motion parameters can also be referred to as motion information. CBs can also be subdivided into one or more transform blocks (TBs). A TB represents a square block of samples of the chrominance component, to which a residual transform (e.g., in some cases, the same two-dimensional transform) is applied to encode the prediction residual signal. A transform unit (TU) represents a TB of luminance and chrominance samples along with the corresponding syntax elements. Transform coding is described in more detail below.

[0065] The size of a CU corresponds to the size of the coding mode and can be square. For example, the size of a CU can be 8x8 samples, 16x16 samples, 32x32 samples, 64x64 samples, or any other suitable size up to the corresponding CTU size. The phrase “N x N” is used herein to refer to the pixel size of a video block in both the vertical and horizontal dimensions (e.g., 8 pixels x 8 pixels). Pixels in a block can be arranged in rows and columns. In some embodiments, a block may not have the same number of pixels in the horizontal direction as it does in the vertical direction. The syntax data associated with a CU can describe, for example, the segmentation of the CU into one or more PUs. The segmentation mode can differ between CUs encoded using intra-predictive mode and inter-predictive mode. PUs can be segmented into non-square shapes. The syntax data associated with a CU can also describe, for example, the segmentation of the CU into one or more TUs according to the CTU. TUs can be square or non-square.

[0066] According to the HEVC standard, transform units (TUs) can be used to perform transforms. The TU can be different for different CUs. The size of the TU can be set based on the size of the PU within a given CU. The TU can have the same size as or smaller than the PU. In some examples, a quadtree structure called a residual quadtree (RQT) can be used to subdivide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can correspond to TUs. The pixel differences associated with the TU can be transformed to produce transform coefficients. The transform coefficients can then be quantized by the encoder engine 106.

[0067] Once the video data is segmented into Units (CUs), the encoder engine 106 uses a prediction mode to predict each Processing Unit (PU). The prediction unit or block is then subtracted from the original video data to obtain a residual (described below). For each CU, the prediction mode can be signaled within the bitstream using syntax data. Prediction modes can include intra-frame prediction (or intra-picture prediction) or inter-frame prediction (or inter-picture prediction). Intra-frame prediction utilizes the correlation between spatially adjacent samples within a picture. For example, using intra-frame prediction, each PU is predicted from neighboring image data in the same picture using, for example, DC prediction to find the average value for the PU, planar prediction to adapt a planar surface to the PU, orientation prediction to infer from neighboring data, or any other suitable prediction type. Inter-frame prediction uses temporal correlations between pictures to derive motion-compensated predictions for blocks of image samples. For example, using inter-frame prediction, each PU is predicted from image data in one or more reference pictures (in the output order before or after the current picture) using motion-compensated prediction. For example, a decision can be made at the CU level whether to use inter-picture prediction or intra-picture prediction to encode a picture region.

[0068] Encoder engine 106 and decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, the video coding apparatus (coder) (e.g., encoder engine 106 and / or decoder engine 116) segments the picture into multiple coding tree units (CTUs) (wherein, one or more CTBs of luminance samples and chrominance samples, together with the syntax used for the samples, are referred to as CTUs). The video coding apparatus can segment CTUs according to a tree structure (e.g., a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple segmentation types (e.g., the distinction between CUs, PUs, and TUs in HEVC). The QTBT structure comprises two levels: a first level segmented according to quadtree segmentation and a second level segmented according to binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0069] In the MTT partitioning structure, blocks can be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. A ternary tree partition is a partition in which a block is divided into three sub-blocks. In some examples, a ternary tree partition divides a block into three sub-blocks without partitioning the original block through a center. The partitioning type in MTT (e.g., quadtree, binary tree, and ternary tree) can be symmetric or asymmetric.

[0070] In some examples, the video encoding apparatus may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoding apparatus may use two or more QTBT or MTT structures, for example, one QTBT or MTT structure for the luma component and another QTBT or MTT structure for the two chroma components (or two QTBT and / or MTT structures for the respective chroma components).

[0071] Video coding apparatuses can be configured to use per-HEVC quadtree segmentation, QTBT segmentation, MTT segmentation, or other segmentation structures. For illustrative purposes, the description herein may refer to QTBT segmentation. However, it should be understood that the techniques of this disclosure can also be applied to video coding apparatuses configured to use quadtree segmentation or other types of segmentation.

[0072] In some examples, one or more slices of an image are assigned slice types. Slice types include intra-coded slices (I-slices), inter-coded P-slices, and inter-coded B-slices. An I-slice (intra-coded frame, independently decodable) is a slice of an image encoded solely by intra-frame prediction and is therefore independently decodable because an I-slice requires only intra-frame data to predict any prediction unit or prediction block of the slice. A P-slice (one-way prediction frame) is a slice of an image that can be encoded using both intra-frame and one-way inter-frame prediction. Each prediction unit or prediction block within a P-slice is encoded using either intra-frame or inter-frame prediction. When inter-frame prediction is applied, the prediction unit or prediction block is predicted using only one reference image, and therefore the reference sample is a reference region from only one frame. A B-slice (two-way prediction frame) is a slice of an image that can be encoded using both intra-frame and inter-frame prediction (e.g., two-way or one-way prediction). Bidirectional prediction of a B-slice's prediction unit or block can be performed from two reference images, where each image contributes a reference region, and the sample sets of the two reference regions are weighted (e.g., with equal weights or different weights) to produce the prediction signal for the bidirectional prediction block. As explained above, a slice of an image is encoded independently. In some cases, an image may be encoded as only one slice.

[0073] As mentioned above, intra-image prediction leverages the correlation between spatially adjacent samples within an image. Several intra-prediction modes exist (also referred to as "intra-frame modes"). In some examples, intra-frame prediction for luma blocks includes 35 modes, comprising planar modes, DC modes, and 33 angular modes (e.g., diagonal intra-frame prediction modes and adjacent angular modes). The 35 intra-frame prediction modes are indexed as shown in Table 1 below. In other examples, more intra-frame modes can be defined, including prediction angles that may not yet be represented by the 33 angular modes. In other examples, the prediction angles associated with angular modes may differ from those used in HEVC.

[0074] Intra-prediction mode Associated names 0 INTRA_PLANAR 1 INTRA_DC 2..34 INTRA_ANGULAR2..INTRA_ANGULAR34

[0075] Table 1 - Specification of Intra-Frame Prediction Modes and Associated Names

[0076] Inter-image prediction utilizes temporal correlations between images to derive motion-compensated predictions for image sample blocks. Using a translational motion model, the position of a block in a previously decoded image (reference image) is represented by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement of the reference block relative to the current block's position, and Δy specifies the vertical displacement of the reference block relative to the current block's position. In some cases, the motion vector (Δx, Δy) can be integer sample precision (also known as integer precision), in which case the motion vector points to an integer pixel grid (or integer pixel sampling grid) of the reference frame. In other cases, the motion vector (Δx, Δy) can have fractional sample precision (also known as fractional pixel precision or non-integer precision) to more accurately capture the motion of the underlying object without being limited to an integer pixel grid of the reference frame. The precision of the motion vector can be expressed by its quantization level. For example, the quantization level can be integer precision (e.g., 1 pixel) or fractional pixel precision (e.g., 1 / 4 pixel, 1 / 2 pixel, or other values ​​below 1 pixel). When the corresponding motion vector has fractional sample accuracy, interpolation is applied to the reference image to derive the predicted signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate values ​​at fractional positions. The previously decoded reference image is indicated by a reference index (refIdx) for a list of reference images. The motion vector and reference index can be referred to as motion parameters. Two types of inter-image prediction can be performed, including single prediction and double prediction.

[0077] In the case of inter-frame prediction using dual prediction, two sets of motion parameters (Δx0, y0, refIdx0 and Δx1, y1, refIdx1) are used to generate two motion-compensated predictions (from the same reference image or possibly from different reference images). For example, in the case of dual prediction, each prediction block uses two motion-compensated prediction signals and generates B prediction units. The two motion-compensated predictions are then combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference images that can be used in dual prediction are stored in two separate lists, denoted as list 0 and list 1, respectively. The motion parameters can be derived at the encoder using a motion estimation process.

[0078] When using single prediction for inter-frame prediction, a set of motion parameters (Δx0, y0, refIdx0) is used to generate motion-compensated predictions from a reference image. For example, in the case of single prediction, each prediction block uses at most one motion-compensated prediction signal and generates P prediction units.

[0079] The prediction unit (PU) may include data related to the prediction process (e.g., motion parameters or other appropriate data). For example, when the PU is encoded using intra-frame prediction, the PU may include data describing the intra-frame prediction mode used for the PU. As another example, when the PU is encoded using inter-frame prediction, the PU may include data defining the motion vectors used for the PU. The data defining the motion vectors used for the PU may describe, for example, the horizontal component (Δx) of the motion vector, the vertical component (Δy) of the motion vector, the resolution used for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference picture to which the motion vector points, the reference index, a list of reference pictures used for the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.

[0080] After performing prediction using intra-frame prediction and / or inter-frame prediction, encoding device 104 can then perform transform and quantization. For example, after prediction, encoder engine 106 can compute residual values ​​corresponding to the PU. Residual values ​​can include pixel differences between the current pixel block (PU) being encoded and the prediction block used to predict the current block (e.g., a predicted version of the current block). For example, after generating a prediction block (e.g., using inter-frame prediction or intra-frame prediction), encoder engine 106 can generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block comprises a set of pixel differences that quantize the differences between pixel values ​​in the current block and pixel values ​​in the prediction block. In some examples, the residual block can be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of pixel values.

[0081] Block transforms are used to transform any residual data that may remain after prediction is performed. These block transforms can be based on Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), integer transforms, wavelet transforms, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., kernels of size 32x32, 16x16, 8x8, 4x4, or other suitable sizes) can be applied to the residual data in each CU. In some examples, TUs can be used for the transform and quantization processes implemented by encoder engine 106. A given CU with one or more PUs can also include one or more TUs. As described further in detail below, residual values ​​can be transformed into transform coefficients using block transforms, and then quantized and scanned using TUs to produce serialized transform coefficients for entropy coding.

[0082] In some examples, after intra-frame or inter-frame prediction coding using the PU of the CU, the encoder engine 106 can compute residual data for the TU of the CU. The PU can include pixel data in the spatial domain (or pixel domain). As previously mentioned, the residual data can correspond to the pixel difference between a pixel in the uncoded image and the prediction value corresponding to the PU. The encoder engine 106 can form one or more TUs that include the residual data for the CU (which includes the PU), and can then transform the TUs to produce transform coefficients for the CU. The TUs can include coefficients in the transform domain after applying the block transform.

[0083] The encoder engine 106 can perform quantization of the transform coefficients. Quantization provides further compression by reducing the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient with an n-bit value can be rounded down to an m-bit value during quantization, where n is greater than m.

[0084] Once quantization is performed, the encoded video bitstream includes quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), segmentation information, and any other suitable data (such as other syntax data). The different elements of the encoded video bitstream can then be entropy-coded by encoder engine 106. In some examples, encoder engine 106 may use a predefined scan order to scan the quantized transform coefficients to produce a serialized vector that can be entropy-coded. In some examples, encoder engine 106 may perform an adaptive scan. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), encoder engine 106 may entropy-code that vector. For example, encoder engine 106 may use context-adaptive variable-length coding, context-adaptive binary arithmetic coding, syntax-based context-adaptive binary arithmetic coding, probabilistic interval segmentation entropy coding, or another suitable entropy coding technique.

[0085] The output 110 of encoding device 104 can transmit NAL units constituting the encoded video bitstream data to decoding device 112 of the receiving device on communication link 120. The input 114 of decoding device 112 can receive the NAL units. Communication link 120 may include a channel provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network may include any wireless interface or combination of wireless interfaces and may include any suitable wireless network (e.g., the Internet or other wide area networks, packet-based networks, WiFi). TM Radio frequency (RF), UWB, WiFi Direct, Cellular, Long Term Evolution (LTE), WiMax TM Wired networks can include any wired interface (e.g., fiber optic, Ethernet, powerline Ethernet, coaxial cable Ethernet, digital signal line (DSL), etc.). Various devices can be used to implement wired and / or wireless networks, such as base stations, routers, access points, bridges, gateways, switches, etc. Encoded video bitstream data can be modulated according to communication standards such as wireless communication protocols and transmitted to receiving devices.

[0086] In some examples, encoding device 104 may store encoded video bitstream data in storage unit 108. Output 110 may obtain encoded video bitstream data from encoder engine 106 or from storage unit 108. Storage unit 108 may include any of a variety of distributed or locally accessed data storage media. For example, storage unit 108 may include hard disk drives, storage disks, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. Storage unit 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. In other examples, storage unit 108 may correspond to a file server or another intermediate storage device that may store encoded video generated by a source device. In such cases, receiving device including decoding device 112 may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing encoded video data and sending such encoded video data to receiving devices. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device can access the encoded video data via any standard data connection, including an internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from storage unit 108 can be streaming, downloading, or a combination thereof.

[0087] Input 114 of decoding device 112 receives encoded video bitstream data and may provide the video bitstream data to decoder engine 116 or to storage unit 118 for later use by decoder engine 116. For example, storage unit 118 may include a DPB for storing reference pictures used in inter-frame prediction. A receiving device including decoding device 112 may receive the encoded video data to be decoded via storage unit 108. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the receiving device. The communication medium used to transmit the encoded video data may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, switch, base station, or any other means that may be used to facilitate communication from the source device to the receiving device.

[0088] Decoder engine 116 decodes the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more encoded video sequences that constitute the encoded video data. Decoder engine 116 can then rescale the encoded video bitstream data and perform an inverse transform on it. The residual data is then passed to the prediction stage of decoder engine 116. Decoder engine 116 then predicts pixel blocks (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (the residual data).

[0089] Video decoding device 112 can output decoded video to video destination device 122, which may include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, video destination device 122 may be part of a receiving device that includes decoding device 112. In some aspects, video destination device 122 may be part of a separate device, distinct from the receiving device.

[0090] In some examples, video encoding device 104 and / or video decoding device 112 may be integrated with audio encoding device and audio decoding device, respectively. Video encoding device 104 and / or video decoding device 112 may also include other hardware or software necessary for implementing the above-described encoding techniques, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. Video encoding device 104 and video decoding device 112 may be integrated as part of a combined encoder / decoder (codec) within the respective device.

[0091] Figure 1 The example system shown is merely an illustrative example that may be used herein. The techniques used to process video data using the techniques described herein can be implemented by any digital video encoding and / or decoding device. Although, in general, the techniques of this disclosure are implemented by video encoding or video decoding devices, the techniques can also be implemented by a combined video encoder-decoder, commonly referred to as a "CODEC". Furthermore, the techniques of this disclosure can also be implemented by a video preprocessor. The source device and receiving device are merely examples of such encoding devices, wherein the source device generates encoded video data for transmission to the receiving device. In some examples, the source device and receiving device may operate in a substantially symmetrical manner, such that each of these devices includes both video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0092] Extensions to the HEVC standard include the Multi-View Video Coding Extension, known as MV-HEVC, and the Scalable Video Coding Extension, known as SHVC. MV-HEVC and SHVC extensions share the concept of layered coding, where different layers are included in the encoded video bitstream. Each layer in the encoded video sequence is addressed by a unique layer identifier (ID). The layer ID can be present in the header of a NAL unit to identify the layer associated with that NAL unit. In MV-HEVC, different layers typically represent different views of the same scene in the video bitstream. In SHVC, different scalable layers are provided to represent the video bitstream with different spatial resolutions (or picture resolutions) or different reconstruction fidelities. A scalable layer can include a base layer (where layer ID = 0) and one or more enhancement layers (where layer ID = 1, 2, ..., n). The base layer may conform to the first version of the HEVC profile and represents the lowest available layer in the bitstream. Compared to the base layer, enhancement layers have increased spatial resolution, temporal resolution, or frame rate and / or reconstruction fidelity (or quality). Enhancement layers are organized hierarchically and may depend on (or not depend on) lower layers. In some examples, a single-standard codec can be used to encode different layers (e.g., using HEVC, SHVC, or other encoding standards to encode all layers). In other examples, multi-standard codecs can be used to encode different layers. For example, AVC can be used to encode the base layer, while one or more enhancement layers can be encoded using the MV-HEVC extension of the SHVC and / or HEVC standards.

[0093] As described above, for each block, a set of motion information (also referred to herein as motion parameters) may be available. This set of motion information may contain motion information for both the forward and backward prediction directions. Here, the forward and backward prediction directions can be two prediction directions in a bidirectional prediction mode, and the terms "forward" and "backward" do not necessarily have geometric meaning. Instead, forward and backward may correspond to a reference image list 0 (RefPicList0) and a reference image list 1 (RefPicList1) for the current image, slice, or block. In some examples, when only one reference image list is available for an image, slice, or block, only RefPicList0 is available, and the motion information for each block of the slice is always forward. In some examples, RefPicList0 includes reference images that are temporally preceding the current image, while RefPicList1 includes reference images that are temporally following the current image. In some cases, motion vectors are used in the decoding process along with associated reference indices. Such motion vectors with associated reference indices are represented as a single prediction set of motion information.

[0094] For each prediction direction, motion information may include a reference index and a motion vector. In some cases, for simplicity, the motion vector may have associated information, based on which it can be assumed that the motion vector has an associated reference index. The reference index can be used to identify a reference image in the current list of reference images (RefPicList0 or RefPicList1). The motion vector may have horizontal and vertical components, providing the offset from the coordinate position in the current image to the coordinate position in the reference image identified by the reference index. For example, the reference index may indicate a specific reference image that should be used for a block in the current image, and the motion vector may indicate where the best-matching block in the reference image (the block that best matches the current block) is located in the reference image.

[0095] Picture order counts (POCs) can be used in video coding standards to identify the display order of pictures. Although it is possible for two pictures within an encoded video sequence to have the same POC value, it is generally not the case that two pictures with the same POC will appear in an encoded video sequence. When multiple encoded video sequences exist in a bitstream, pictures with the same POC value are likely to be closer to each other in terms of decoding order. The POC value of a picture can be used for reference picture list construction (such as reference picture set derivation in HEVC) and / or motion vector scaling, among other things.

[0096] In H.264 / AVC, each inter-frame macroblock (MB) can be partitioned in four different ways: one 16x16 macroblock partition; two 16x8 macroblock partitions; two 8x16 macroblock partitions; and four 8x8 macroblock partitions, as well as other partitions. Different macroblock partitions within a macroblock can have different reference index values ​​for each prediction direction (e.g., different reference index values ​​for RefPicList0 and RefPicList1).

[0097] In some cases, when a macroblock is not divided into four 8x8 macroblock partitions, it may have only one motion vector per macroblock partition in each prediction direction. In other cases, when a macroblock is divided into four 8x8 macroblock partitions, each 8x8 macroblock partition can be further subdivided into sub-blocks, and each sub-block may have a different motion vector in each prediction direction. An 8x8 macroblock partition can be divided into sub-blocks in different ways, including: one 8x8 sub-block; two 8x4 sub-blocks; two 4x8 sub-blocks; and four 4x4 sub-blocks; and other sub-blocks. Each sub-block may have a different motion vector in each prediction direction. Therefore, motion vectors can exist at a level equal to or higher than that of the sub-blocks.

[0098] In HEVC, the largest coding unit in a slice is called a coding tree block (CTB) or coding tree unit (CTU). A CTB contains a quadtree, where the nodes are coding units. In the HEVC master profile, the size of a CTB can range from 16x16 pixels to 64x64 pixels. In some cases, a CTB size of 8x8 pixels is supported. A CTB can be recursively divided into coding units (CUs) using a quadtree approach. CUs can have the same size as the CTB and can be as small as 8x8 pixels. In some cases, a mode (e.g., intra-prediction mode or inter-prediction mode) can be used to encode each coding unit. When using inter-prediction mode to inter-code CUs, a CU can be further divided into two or four prediction units (PUs), or when further division is not applicable, a CU can be considered as a single PU. When two PUs exist within a CU, the two PUs can be rectangles of half the size, or two rectangles that are 1 / 4 or 3 / 4 the size of the CU.

[0099] When a CU is inter-coded, a set of motion information can exist for each PU, which can be derived using a unique inter-prediction mode. For example, each PU can be encoded using an inter-prediction mode to derive the set of motion information. In some cases, when intra-prediction modes are used to intra-code the CU, the PU shape can be 2Nx2N or NxN. Within each PU, a single intra-prediction mode is encoded (while the chroma prediction mode is signaled at the CU level). In some cases, an NxN intra-PU shape is allowed when the current CU size is equal to the minimum CU size defined in the SPS.

[0100] For motion prediction in HEVC, there are two inter-frame prediction modes for prediction units (PUs): merge mode and Advanced Motion Vector Prediction (AMVP) mode. Special cases considered as merge are skipped. In either AMVP or merge mode, a candidate list of motion vectors (MVs) for multiple motion vector predictors can be maintained. In merge mode, the motion vector of the current PU and its reference index are generated by selecting a candidate from the MV candidate list.

[0101] In some examples, the MV candidate list contains up to five candidates for the merge mode and two candidates for the AMVP mode. In other examples, different numbers of candidates can be included in the MV candidate list for the merge mode and / or AMVP mode. The merge candidate can contain a set of motion information. For example, the set of motion information can include motion vectors corresponding to two lists of reference images (list 0 and list 1) and reference indices. If the merge candidate is identified by a merge index, the reference image is used for the prediction of the current block and to determine the associated motion vector. However, in AMVP mode, for each potential prediction direction from list 0 or list 1, the reference index needs to be explicitly signaled along with the MV predictor (MVP) index of the MV candidate list, because AMVP candidates only contain motion vectors. In AMVP mode, the predicted motion vectors can be further refined.

[0102] Merged candidates can correspond to a complete set of motion information, while AMVP candidates can contain a motion vector and a reference index for a specific prediction direction. Candidates for both modes can be derived similarly from the same spatially and temporally adjacent blocks.

[0103] In some examples, the merge mode allows inter-frame predicted PUs to inherit one or more motion vectors, prediction directions, and one or more reference image indices from inter-frame predicted PUs that include motion data locations selected from a set of spatially adjacent motion data locations and one of two temporally co-located motion data locations. In AMVP mode, one or more motion vectors of the PU can be predicted and encoded relative to one or more motion vector predictors (MVPs) from an AMVP candidate list constructed by the encoder. In some cases, for unidirectional inter-frame prediction of the PU, the encoder can generate a single AMVP candidate list. In some cases, for bidirectional prediction of the PU, the encoder can generate two AMVP candidate lists, one using motion data from spatially and temporally adjacent PUs from the forward prediction direction, and one using motion data from spatially and temporally adjacent PUs from the backward prediction direction.

[0104] Candidates for both modes can be derived from spatially and / or temporally adjacent blocks. For example, Figure 2A and Figure 2B Includes a conceptual diagram showing spatially adjacent candidates. Figure 2A Candidate spatially adjacent motion vectors (MVs) for merging patterns are shown. Figure 2BSpatial neighbor motion vector (MV) candidates for AMVP mode are shown. Although the methods for generating candidates based on blocks differ for merging and AMVP modes, spatial MV candidates are derived based on neighboring blocks for a specific PU (PU0).

[0105] In merging mode, the encoder can form a merging candidate list by considering merging candidates from various motion data locations.

[0106] For example, such as Figure 2A As shown, regarding in Figure 2A The spatially adjacent motion data positions, represented by the numbers 0-4, can be used to deduce multiple...

[0107] There are four spatial MV candidates. In the merged candidate list, the MV candidates can be ordered according to the numbers 0-4. For example, the positions and order can include: left position (0), top position (1), top right position (2), bottom left position (3), and top left position (4). Figure 2A In this context, block 200 includes PU0 202 and PU1 204. In some examples, when the video encoding device encodes motion information for PU0 202 using a merging mode, the video encoding device may add motion information from spatially adjacent blocks 210-218 to the candidate list in the order described above.

[0108] exist Figure 2B In the AVMP mode shown, adjacent blocks can be divided into two groups: the left group including blocks 0 and 1, and the upper group including blocks 2, 3, and 4. Figure 2B In this context, blocks 0, 1, 2, 3, and 4 are labeled as blocks 230, 232, 234, 236, and 238, respectively. Here, block 220 includes PU0 222 and PU1 224, and blocks 230, 232, 234, 236, and 238 represent spatial neighbors of PU0 222. For each group, potential candidates whose references in adjacent blocks are identical to the reference images indicated by the signaled reference index have the highest priority and are selected to form the final candidates for that group. It is possible that none of the adjacent blocks contain motion vectors pointing to the same reference image. Therefore, if no such candidate can be found, the first available candidate can be scaled to form the final candidate, thus compensating for temporal distance differences.

[0109] Figure 3A and Figure 3B Includes a conceptual diagram illustrating the prediction of time motion vectors. Figure 3A An exemplary CU 300 including PU0 302 and PU1 304 is shown. PU0 302 includes a central block 310 for PU0 302 and a lower right block 306 for PU0 302. Figure 3A An external block 308 is also shown, and the motion information for the external block 308 can be predicted from the motion information of PU0 302, as discussed below. Figure 3B The current image 342 shows the current block 326, which includes motion information to be predicted for it. Figure 3B Also shown is a co-image 330 for the current image 342 (including a co-image block 324 for the current block 326), a current reference image 342, and a co-image reference 332. The co-image block 324 is predicted using a co-image motion vector 320, which is used as a temporal motion vector predictor (TMVP) 322 for motion information of block 326.

[0110] The video encoding device adds a Temporal Motion Vector Predictor (TMVP) candidate (e.g., TMVP 322) (if enabled and available) to the MV candidate list after any spatial motion vector candidate. The process of deriving motion vectors for a TMVP candidate is the same for both merge and AMVP modes. However, in some cases, the target reference index for a TMVP candidate is always set to zero in merge mode.

[0111] like Figure 3A As shown, the primary block position for TMVP candidate derivation is the lower right block 306 outside of co-located PU 304, to compensate for deviations from the blocks above and to the left used to generate spatially adjacent candidates. However, if block 306 is located outside the current CTB (or LCU) row (e.g., as...), Figure 3A (as shown in block 308) or if the motion information for block 306 is unavailable, then the block is replaced by the center block 310 of PU 302.

[0112] refer to Figure 3B Motion vectors for TMVP candidate 322 can be derived from co-location blocks 324 of co-location image 330, indicated at the slice level. Similar to the direct temporal mode in AVC, motion vector scaling can be performed on the motion vectors of the TMVP candidates to compensate for distance differences between the current image 342 and the current reference image 340, and between the co-location image 330 and the co-location reference image 332. That is, motion vector 320 can be scaled based on the distance differences between the current image (e.g., current image 342) and the current reference image (e.g., current reference image 340), and between the co-location image (e.g., co-location image 330) and the co-location reference image (e.g., co-location reference image 332) to produce TVMP candidate 322.

[0113] Other aspects of motion prediction are also covered in HEVC, VVC, and other video coding standards. One aspect, for example, includes motion vector scaling. In motion vector scaling, it is assumed that the value of a motion vector is proportional to the distance between images at rendering time. In some examples, a first motion vector may be associated with two images (including a first reference image and a first containing image that includes the first motion vector). A second motion vector can be predicted using the first motion vector. To predict the second motion vector, a first distance between the first containing image and the first reference image can be calculated based on the Picture Order Count (POC) values ​​associated with the first reference image and the first containing image of the first motion.

[0114] A second reference image and a second contained image can be associated with a second motion vector to be predicted, wherein the second reference image may be different from the first reference image, and the second contained image may be different from the first contained image. A second distance between the second reference image and the second contained image can be calculated based on the POC values ​​associated with the second reference image and the second contained image, wherein the second distance may be different from the first distance. To predict the second motion vector, the first motion vector can be scaled based on the first distance and the second distance. For spatially adjacent candidates, the first contained image and the second contained image of the first motion vector and the second motion vector can be the same, while the first reference image and the second reference image can be different. In some examples, motion vector scaling can be applied to TMVP and AMVP patterns for spatially and temporally adjacent candidates.

[0115] Another aspect of motion prediction involves the generation of artificial motion vector candidates. For example, if the list of motion vector candidates is incomplete, artificial motion vector candidates are generated and inserted at the end of the list until all candidates are obtained. In the merging mode, there are two types of artificial motion vector candidates: a first type, which includes combined candidates derived only for B-slices; and a second type, which includes zero candidates for AMVP only (if the first type does not provide enough artificial candidates). For each pair of candidates already in the motion vector candidate list and with the necessary motion information, a bidirectional combined motion vector candidate can be derived by referencing the motion vector of the first candidate in the image in list 0 and the motion vector of the second candidate in the image in list 1.

[0116] Another aspect of the merge and AMVP pattern involves a pruning process for candidate insertion. For example, candidates from different blocks might happen to be identical, which reduces the efficiency of merging and / or AMVP candidate lists. A pruning process can be applied to address this issue. The pruning process involves comparing candidates with those already existing in the current candidate list to avoid inserting identical or duplicate candidates. To reduce the complexity of comparisons, the pruning process can be performed on fewer potential candidates than those to be inserted into the candidate list.

[0117] In some examples, enhanced motion vector prediction can be implemented. For instance, video coding standards such as VVC specify inter-frame coding tools that allow for the derivation or refinement of candidate lists for motion vector predictions or merge predictions for the current block. Examples of such methods are described below.

[0118] History-based motion vector prediction (HMVP) is a motion vector prediction method that allows each block to find its MV predictors not only from its immediate neighboring motion fields with causal relationships, but also from a list of previously decoded MVs. For example, using HMVP, in addition to the MV predictors from its immediate neighboring motion fields with causal relationships, one or more MV predictors for the current block can be obtained or predicted from a list of previously decoded MVs. The MV predictors in the list of previously decoded MVs are called HMVP candidates. HMVP candidates can include motion information associated with inter-coded blocks. An HMVP table with multiple HMVP candidates can be maintained during the encoding and / or decoding process for a slice. In some examples, the HMVP table can be updated dynamically. For example, after decoding an inter-coded block, the HMVP table can be updated by adding the associated motion information of the decoded inter-coded block as a new HMVP candidate. In some examples, the HMVP table can be cleared when a new slice is encountered.

[0119] In some cases, whenever an inter-frame coded block exists, the associated motion information can be inserted into the table as a new HMVP candidate in a first-in, first-out (FIFO) manner. Constrained FIFO rules can be applied. When inserting an HMVP into the table, a redundancy check can first be applied to see if the same HMVP already exists in the table. If found, that particular HMVP can be removed from the table, and then all HMVP candidates can be moved.

[0120] In some examples, HMVP candidates can be used during the merge candidate list construction process. In some cases, all HMVP candidates from the last entry to the first entry in the table are inserted after the TMVP candidates. Pruning can be applied to HMVP candidates. The merge candidate list construction process can be terminated once the total number of available merge candidates reaches the maximum allowed merge candidates signaled.

[0121] In some examples, HMVP candidates can be used during the AMVP candidate list construction process. In some cases, the motion vectors of the last K HMVP candidates in the table are inserted after the TMVP candidates. In some implementations, only HMVP candidates with the same reference image as the AMVP target reference image are used to construct the AMVP candidate list. Pruning can be applied to HMVP candidates.

[0122] Figure 4 This is a block diagram illustrating an example of HMVP table 400. In some examples, HMVP table 400 may be implemented as a storage device and / or structure managed using a first-in, first-out (FIFO) rule. For example, HMVP candidates, including MV predictors, may be stored in HMVP table 400. HMVP candidates may be stored in the order in which they are encoded or decoded. In one example, the order in which HMVP candidates are stored in HMVP table 400 may correspond to the time in which the HMVP candidates are constructed. For example, when implemented in a decoder such as decoding device 112, HMVP candidates may be constructed to include motion information of decoded inter-frame coded blocks. In some examples, one or more HMVP candidates from HMVP table 400 may include motion vector predictors, which can be used to predict motion vectors for the current block to be decoded. In some examples, one or more HMVP candidates may include one or more such previously decoded blocks, which may be stored in one or more entries of HMVP table 400 in a FIFO manner in the order in which they are decoded.

[0123] HMVP candidate index 402 is shown as being associated with HMVP table 400. HMVP candidate index 402 may identify one or more entries in HMVP table 400. According to an illustrative example, HMVP candidate index 402 is shown as including index values ​​0 to 4, where each index value of HMVP candidate index 402 is associated with a corresponding entry. In other examples, HMVP table 400 may include references to... Figure 4The entries shown and described are compared to the number of entries. When constructing HMVP candidates, they are filled into HMVP table 400 in a FIFO manner. For example, as HMVP candidates are decoded, they are inserted into HMVP table 400 at one end, and entries are sequentially moved through HMVP table 400 until they leave HMVP table 400 from the other end. Therefore, in some examples, memory structures such as shift registers can be used to implement HMVP table 400.

[0124] In one example, index 0 could point to the first entry of HMVP table 400, which could correspond to the first end of HMVP table 400 where an HMVP candidate is inserted. Correspondingly, index 4 could point to the second entry of HMVP table 400, which could correspond to the second end of HMVP table 400 where an HMVP candidate is removed from or cleared from HMVP table 400. Therefore, the HMVP candidate inserted at the first entry at index 0 can traverse HMVP table 400 to make room for newer or more recently decoded HMVP candidates until an HMVP candidate reaches the second entry at index 4. Thus, among the HVMP candidates present in HMVP table 400 at any given time, the HMVP candidate in the second entry at index 4 could be the oldest or earliest, while the HMVP candidate in the first entry at index 0 could be the newest or most recent. Typically, the HMVP candidate in the second entry can be an older or earlier constructed HMVP candidate compared to the one in the first entry.

[0125] exist Figure 4 In the accompanying drawings, reference numerals 400A, 400B, and 400C are used to identify different states of the HMVP table 400. Referring to state 400A, HMVP candidates HMVP0 through HMVP4 are shown as entries existing at corresponding index values ​​4 through 0 in the HMVP table 400. For example, HMVP0 could be the oldest or earliest HMVP candidate inserted into the first entry at index value 0 in the HMVP table 400. HMVP0 can be sequentially shifted to make room for earlier inserted and newer HMVP candidates HMVP1 through HMVP4 until HMVP0 reaches the second entry at index value 4 as shown in state 400A. Accordingly, HMVP4 could be the most recent HMVP candidate to be inserted into the first entry at index value 0. Therefore, HMVP0 is an older or earlier HMVP candidate in the HMVP table 400 relative to HMVP4.

[0126] In some examples, one or more of the HMVP candidates HMVP0 through HMVP4 may include potentially redundant motion vector information. For example, a redundant HMVP candidate may include motion vector information that is identical to that in one or more other HMVP candidates stored in HMVP table 400. Since the motion vector information of a redundant HMVP candidate can be obtained from one or more other HMVP candidates, storing the redundant HMVP candidate in HMVP table 400 can be avoided. By avoiding storing redundant HMVP candidates in HMVP table 400, the resources of HMVP table 400 can be utilized more efficiently. In some examples, a redundancy check can be performed before storing an HMVP candidate in HMVP table 400 to determine whether the HMVP candidate would be redundant (e.g., the motion vector information of an HMVP candidate can be compared with the motion vector information of other already stored HMVP candidates to determine if a match exists).

[0127] In some examples, state 400B of HMVP table 400 is a conceptual illustration of the redundancy check described above. In some examples, HMVP candidates can be populated in HMVP table 400 during decoding, and the redundancy check can be performed periodically instead of as a threshold test before storing HMVP candidates. For example, as shown in state 400B, HMVP candidates HMVP1 and HMVP3 can be identified as redundant candidates (e.g., their motion information is the same as the motion information of one of the other HMVP candidates in HMVP table 400). Redundant HMVP candidates HMVP1 and HMVP3 can be removed, and the remaining HMVP candidates can be shifted accordingly.

[0128] For example, as shown in state 400C, HMVP candidates HMVP2 and HMVP4 are shifted toward higher index values ​​corresponding to older entries, while HMVP0, already in the second entry at the end of HMVP table 400, is shown as not shifted further. In some examples, shifting HMVP candidates HMVP2 and HMVP4 can free up space in HMVP table 400 for newer HMVP candidates. Thus, new HMVP candidates HMVP5 and HMVP6 are shown as shifted into HMVP table 400, where HMVP6 is the most recent or includes recently decoded motion vector information and is stored in the first entry at index 0.

[0129] In some examples, one or more HMVP candidates from HMVP table 400 can be used to construct an additional candidate list that can be used for motion prediction of the current block. For example, one or more HMVP candidates from HMVP table 400 can be added to the merged candidate list, for example, as additional merged candidates. In some examples, one or more HMVP candidates from the same HMVP table 400 or another such HMVP table can be added to the Advanced Motion Vector Prediction (AMVP) candidate list, for example, as additional AMVP predictors.

[0130] For example, during the construction of the merge candidate list, some or all of the HMVP candidates stored in the entries of HMVP table 400 can be inserted into the merge candidate list. In some examples, inserting HMVP candidates into the merge candidate list may include inserting HMVP candidates after the Time Motion Vector Predictor (TMVP) candidates in the merge candidate column. (See previous reference...) Figure 3A and Figure 3B As discussed, if a TMVP candidate is enabled and available, it can be added to the MV candidate list after the spatial motion vector candidate.

[0131] In some examples, the pruning process described above can be applied to HMVP candidates when constructing the merge candidate list. For instance, once the total number of merge candidates in the merge candidate list reaches the maximum allowed number of merge candidates, the merge candidate list construction process can be terminated, and no more HMVP candidates can be inserted into the merge candidate list. The maximum allowed number of merge candidates in the merge candidate list can be a predetermined number or a number that can be signaled from the encoder to the decoder to construct the merge candidate list.

[0132] In some examples of constructing a merge candidate list, one or more other candidates can be inserted into the merge candidate list. In some examples, motion information from previously encoded blocks that are not adjacent to the current block can be used for more efficient motion vector prediction. For example, non-adjacent spatial merge candidates can be used in constructing the merge candidate list. In some cases, the construction of non-adjacent spatial merge candidates (e.g., as described in JVET-K0228, which is incorporated herein by reference in its entirety and for all purposes) involves deriving new spatial candidates from two adjacent non-adjacent locations (e.g., from the nearest non-adjacent block on the left / above, such as...). Figure 5(As shown and discussed below). These blocks can be limited to a maximum distance of one CTU to the current block. The extraction process for non-adjacent candidates begins by tracing back the previously decoded block in the vertical direction. Vertical backtracking stops when an inter-frame block is encountered or the backtracking distance reaches one CTU size. The extraction process then traces back the previously decoded block in the horizontal direction. The criterion for stopping the horizontal extraction process depends on whether the vertical non-adjacent candidate has been successfully extracted. If no vertical non-adjacent candidate has been extracted, the horizontal extraction process stops when an inter-frame block is encountered or the backtracking distance exceeds one CTU size threshold. If an extracted vertical non-adjacent candidate exists, the horizontal extraction process stops when an inter-frame block containing a different MV from the vertical non-adjacent candidate is encountered or the backtracking distance exceeds the CTU size threshold. In some examples, non-adjacent space merge candidates can be inserted before TMVP candidates in the merge candidate list.

[0133] In some examples, a non-adjacent space merge candidate can be inserted before a TMVP candidate in the same merge candidate list, which may include one or more HMVP candidates inserted after the TMVP candidate. See below for further details. Figure 5 This describes the identification and extraction of one or more non-adjacent space merge candidates that can be inserted into the merge candidate list.

[0134] As further described herein, in some examples, one or more HMVP candidates in HMVP table 400 may include affine motion vectors generated using affine motion vector prediction. For example, a block may have multiple affine motion vectors computed using affine motion vector prediction. However, instead of storing multiple motion vectors for the block in HMVP table 400, a single affine motion vector may be generated for the block and stored in HMVP table 400. In some cases, a single affine motion vector may be generated for a specific location of the block and / or a single affine motion vector corresponding to a specific location of the block may be generated. For example, in some cases, an affine motion vector may be generated for the center location of the block, and the affine motion vector generated for the center location of the block may be stored in HMVP table 400. Therefore, instead of storing multiple affine motion vectors for the block, a single affine motion vector is generated for the block and stored in HMVP table 400.

[0135] Figure 5This is a block diagram illustrating an image or slice 500, which includes the current block 502 to be encoded. In some examples, a merge candidate list can be constructed for encoding the current block 502. For example, a motion vector for the current block can be obtained from one or more merge candidates in the merge candidate list. The merge candidate list may include identifying non-adjacent spatial merge candidates. For example, non-adjacent spatial merge candidates may include new spatial candidates derived from two non-adjacent adjacent locations relative to the current block 502.

[0136] Several adjacent or neighboring blocks of the current block 502 are shown, including the upper-left block B2 510 (above and to the left of the current block 502), the upper block B1 512 (above the current block 502), the upper-right block B0 514 (above and to the right of the current block 502), the left-side block A1 516 (to the left of the current block 502), and the lower-left block A0 518 (to the left and below the current block 502). In some examples, non-adjacent space merge candidates can be obtained from one of the nearest non-adjacent blocks above and / or to the left of the current block.

[0137] In some examples, non-adjacent space merge candidates for the current block 502 may include tracing back previously decoded blocks in the vertical direction (above the current block 502) and / or the horizontal direction (to the left of the current block 502). The vertical tracing back distance 504 indicates the vertical distance relative to the current block 502 (e.g., the top boundary of the current block 502) and the vertical non-adjacent block V. N 520. Horizontal backtrack distance 508 indicates the horizontal distance relative to the current block 502 (e.g., the left boundary of the current block 502) and the horizontal non-adjacent block H. N 522. The vertical backtracking distance 504 and the horizontal backtracking distance 508 are limited to the maximum distance equal to one coding tree unit (CTU).

[0138] Vertically non-adjacent blocks, such as V, can be identified by tracing previously decoded blocks in both the vertical and horizontal directions. N 520 and horizontal non-adjacent block H N Candidates for merging non-adjacent spaces, such as 522. For example, extracting vertical non-adjacent blocks V. N 520 may include a vertical backtracking process to determine whether an inter-coded block (constrained to the maximum size of a CTU) exists within a vertical backtracking distance 504. If such a block exists, it is identified as a vertically non-adjacent block V. N520. In some examples, a horizontal backtracking process can be performed after the vertical backtracking process. The horizontal backtracking process may include determining whether an inter-coded block (constrained to the maximum size of a CTU) exists within a horizontal backtracking distance of 506, and if such a block is found, identifying it as a horizontally non-adjacent block H. N 522.

[0139] In some examples, vertical non-adjacent blocks V N 520 and horizontal non-adjacent block H N One or more of the 522 blocks can be extracted as candidates for non-adjacent space merging. If a vertical non-adjacent block V is identified during the vertical reverse tracing process... N 520, then the extraction process may include extracting vertical non-adjacent blocks V. N 520. Then, the extraction process can continue with the horizontal reverse tracing process. If a vertical non-adjacent block V is not identified during the vertical reverse tracing process... N 520, then when encountering an inter-frame coded block or when the horizontal backtracking distance exceeds the maximum distance by 508, the horizontal reverse tracing process can be terminated. If a vertical non-adjacent block V is identified and extracted... N 520, then when encountering a block V that contains and is contained within a vertically non-adjacent block. N In 520, the horizontal backtracking process terminates if the inter-frame coding blocks of different MVs are involved, or if the horizontal backtracking distance exceeds the maximum distance by 508. As mentioned earlier, the extracted non-adjacent spatial merging candidates (e.g., vertical non-adjacent blocks V) are added before the TMVP candidates in the merging candidate list. N 520 and horizontal non-adjacent block H N One or more of the ones in 522).

[0140] Return to reference Figure 4 In some cases, HMVP candidates can also be used when constructing the AMVP candidate list. During the AMVP candidate list construction process, some or all HMVP candidates from entries stored in the same HMVP table 400 (or a different HMVP table used for merging candidate list construction) can be inserted into the AMVP candidate list. In some examples, inserting HMVP candidates into the AMVP candidate list can include inserting a set of HMVP candidate entries (e.g., the number of k newest or oldest entries) after the TMVP candidates in the AMVP candidate list. In some examples, the aforementioned pruning process can be applied to HMVP candidates when constructing the AMVP candidate list. In some examples, only those HMVP candidates with the same reference image as the AMVP target reference image can be used to construct the AMVP candidate list.

[0141] Therefore, a history-based motion vector predictor (HMVP) prediction mode can involve using a history-based lookup table, such as an HMVP table 400 that includes one or more HMVP candidates. HMVP candidates can be used in inter-frame prediction modes, such as merge mode and AMVP mode. In some examples, different inter-frame prediction modes may use different methods to select HMVP candidates from the HMVP table 400.

[0142] In some cases, alternative motion vector prediction designs can be used. For example, alternative designs for spatial MVP (S-MVP) prediction and temporal MVP (T-MVP) prediction can be utilized. For instance, in some implementations of the merge mode (in some cases, the merge mode may be referred to as the skip mode or the direct mode), it can be performed according to... Figure 6A , Figure 6B and Figure 6C The given order shown is used to access (or search or select) the spatial and temporal MVP candidates shown in these figures to populate the MVP list. Figure 6A The locations of MVP candidates are shown. For example, the spatial and temporal locations used in MVP prediction are as follows: Figure 6A As shown. In Figure 6B The image shows an example of the access order (or search order or selection order) used for S-MVP. Figure 6C The diagram shows a spatial inversion pattern (with) Figure 6B (The order in the text is replaced by the order of the characters.)

[0143] In some examples, the spatial neighbors used as MVP candidates for the current block 620 may include blocks A(602), B(604), (C(606), A1(610)|B1(614)), A0(608), and B2(616), which can be in Figure 6B The access order, marked in the text and described below, is achieved using a two-phase process.

[0144] In some examples, the first group (e.g., group 1) may include block A (602) with access order 0, block B (604) with access order 1, and block C (606) with access order 2, where C (606) may be in the same position as B0 (612) in HEVC notation. Depending on the availability of the MVP in the central block C (606) and the type of block partitioning, the first group may also include block A1 (610) with access order 3 or block B1 (614) with access order 3. The second group (e.g., group 2) may include block A0 (608) with access order 5 and block B2 (616) with access order 4.

[0145] In addition, refer to Figure 6AThe temporally adjacent neighbors used as MVP candidates can be the block in the center of the current block 620 (center block C622) and the block H (624) located at the bottom right outside the current block 620. For example, a set could include blocks C 622 and H 624. If block H (624) is found outside the co-location picture, one or more backtracking H positions (H block 626) can be used instead.

[0146] In some cases, depending on the block splitting and encoding order used, a reverse S-MVP candidate order can be used, such as... Figure 6C As shown. For example, the first group may include block 614 with access order 0, block 614 with access order 1, block 616 with access order 2, block 604 with access order 3, or block 626 with access order 3. The second group may include block 612 with access order 4 and block 624 with access order 5.

[0147] In HEVC and earlier video coding standards, only translational motion models were applied to motion compensation prediction (MCP). For example, translational motion vectors could be determined for each block of an image (e.g., each CU or each PU). However, in various cases, other types of motion can exist besides translational motion, including scaling (e.g., zooming in and / or zooming out), rotation, perspective motion, and other irregular motions. Therefore, affine transformation motion compensation prediction can also be applied to improve coding efficiency.

[0148] For example, in some video coding standards such as HEVC, each block has a single motion vector (e.g., a translational motion vector). However, in affine coding modes, a block can have multiple affine motion vectors (e.g., each sample in the block can have an independent affine motion vector). Furthermore, the HMVP table in some video coding standards (e.g., HEVC) is not extended to include affine information. As further described herein, to enable the use of affine information, in some cases, a translational approximation for the affine motion vector can be derived and stored in the HMVP table. In some examples, for an affine block that can have multiple motion vectors, a single motion vector can be generated for a specific location of the block (e.g., the center location), and this single motion vector can be stored in the HMVP table to allow such an HMVP table to include affine information (even if not otherwise supported). In this way, affine transform motion compensation prediction can be applied to improve coding efficiency.

[0149] Figure 7 This shows two motion vectors passing through two control points 710 and 712. and A diagram describing the affine motion field of the current block 702. Motion vectors using control point 710. and the motion vector of control point 712 The motion vector field (MVF) of the current block 702 can be described by the following equation:

[0150]

[0151] In equation (1), v x and v y Form a motion vector for each pixel within the current block 702, where x and y provide the position of each pixel within the current block 702 (e.g., the top-left pixel in the block could have coordinates or indices (x,y) = (0,0)), (v 0x ,v 0y ) is the motion vector of the top-left control point 710, w is the width of the current block 702, and (v 1x ,v 1y ) is the motion vector of the upper right control point 712. 0x and v 1x The value is the horizontal value used for the corresponding motion vector, and v 0y and v 1y The value is the vertical value used for the corresponding motion vector. Additional control points can be defined by adding additional control point vectors (e.g., four control points, six control points, eight control points, or some other number of control points), such as at the lower corner of the current block 702, the center of the current block 702, or other locations within the current block 702.

[0152] Equation (1) above shows a 4-parameter motion model, where the four affine parameters a, b, c, and d are defined as: c = v 0x ; and d = v 0y Using equation (1), the motion vector (v) at the given upper left control point 710 is... 0x ,v 0y ) and the motion vector (v) of the upper right control point 712 1x ,v 1y In the case of ), the motion vector for each pixel in the current block can be calculated using the coordinates (x, y) of each pixel position. For example, for the top-left pixel position of the current block 702, the value of (x, y) can be equal to (0, 0). In this case, the motion vector for the top-left pixel becomes V. x =v 0x and V y =v 0y To further simplify MCP, block-based affine transformation prediction can be applied.

[0153] Figure 8This is a graph showing the block-based affine transformation prediction of the current block 802, which has been divided into sub-blocks. Figure 8 The example shown includes a 4x4 partition with 16 sub-blocks. Any suitable partition and a corresponding number of sub-blocks can be used. The motion vector for each sub-block can then be derived using equation (1). For example, to derive the motion vector for each 4x4 sub-block, the motion vector of the center sample of each sub-block is calculated according to equation (1) (e.g., Figure 8 (As shown). The obtained motion vectors can be rounded, for example, to 1 / 16 fractional precision or other appropriate precision (e.g., 1 / 4, 1 / 8, etc.). The derived sub-block motion vectors can then be used to apply motion compensation to generate predictions for each sub-block. For example, the decoding device can receive the motion vectors describing control point 810. and the motion vector of control point 812 The four affine parameters (a, b, c, d) are used, and the motion vector for each sub-block can be calculated based on the pixel coordinate index describing the position of the center sample of each sub-block. As mentioned above, after MCP, the high-precision motion vector of each sub-block can be rounded and can be saved with the same precision as the translation motion vector.

[0154] Figure 9 This diagram illustrates an example of motion vector prediction in affine inter-frame (AF_INTER) mode. In JEM, there are two affine motion modes: affine inter-frame (AF_INTER) mode and affine merge (AF_MERGE) mode. In some examples, the AF_INTER mode can be applied when the CU has a width and height greater than 8 pixels. An affine flag can be placed (or signaled) in the block-related bitstream (e.g., at the CU level) to indicate whether the AF_INTER mode is applied to that block. Figure 9As shown in the example, in AF_INTER mode, adjacent blocks can be used to construct a candidate list of motion vector pairs. For example, for child block 910 located at the top left corner of current block 902, motion vector v0 can be selected from adjacent block A 920 above and to the left of child block 910, adjacent block B 922 above child block 910, and adjacent block C 924 to the left of child block 910. As another example, for child block 912 located at the top right corner of current block 902, motion vector v1 can be selected from adjacent blocks D 926 above and adjacent block E 928 to the top right, respectively. Adjacent blocks can be used to construct a candidate list of motion vector pairs. For example, given that motion vectors vA, vB, vC, vD, and vE correspond to blocks A920, B922, C924, D926, and E928, respectively, the candidate list of motion vector pairs can be represented as {(v0, v1) | v0 = {vA, vB, vC}, v1 = {vD, vE}}.

[0155] As mentioned above and as Figure 9 As shown, in AF_INTER mode, motion vector v0 can be selected from the motion vectors of blocks A 9720, B 922, or C924. The motion vectors from neighboring blocks (blocks A, B, or C) can be scaled based on: a reference list, and the relationship between the POCs used as references for neighboring blocks, the POCs used as references for the current CU (e.g., current block 902), and the POCs of the current CU. In these examples, some or all POCs can be determined from the reference list. Selecting v1 from neighboring blocks D or E is similar to selecting v0.

[0156] In some cases, if the candidate list contains fewer than two candidates, the list can be populated using motion vector pairs by copying each of the AMVP candidates. When the candidate list contains more than two candidates, in some examples, the candidates in the list can be sorted first based on the consistency of adjacent motion vectors (e.g., consistency could be based on the similarity between two motion vectors in a motion vector pair candidate). In such examples, the first two candidates are retained, while the rest can be discarded.

[0157] In some examples, rate-distortion (RD) cost checks can be used to determine which motion vector pair candidate is selected as the Control Point Motion Vector Prediction (CPMVP) for the current CU (e.g., current block 902). In some cases, an index indicating the position of the CPMVP in the candidate list can be signaled (or otherwise indicated) in the bitstream. Once the CPMVP for the current affine CU is determined (based on the motion vector pair candidates), affine motion estimation can be applied, and the Control Point Motion Vector (CPMV) can be determined. In some cases, the difference between the CPMV and CPMVP can be signaled in the bitstream. Both the CPMV and CPMVP comprise two sets of translational motion vectors, in which case the signaling cost of the affine motion information is higher than the signaling cost of the translational motion.

[0158] Figure 10A and Figure 10B An example of motion vector prediction in AF_MERGE mode is shown. When encoding the current block 1002 (e.g., CU) using AF_MERGE mode, motion vectors can be obtained from valid neighboring reconstructed blocks. For example, the first block among valid neighboring reconstructed blocks encoded in affine mode can be selected as a candidate block. Figure 10A As shown, neighboring blocks can be selected from the set of adjacent blocks A 1020, B 1022, C 1024, D 1026, and E 1028. Neighboring blocks can be considered according to a specific selection order used to select them as candidate blocks. An example of the selection order is the left neighbor (block A 1020), followed by the top neighbor (block B 1022), then the top-right neighbor (block C 1024), then the bottom-left neighbor (block D 1026), and then the top-left neighbor (block E 1028).

[0159] As mentioned above, the selected adjacent blocks can be the first block already encoded in affine mode (e.g., in selection order). For example, block A 1020 may have already been encoded in affine mode. Figure 10B As shown, block A 1020 can be included in adjacent CU 1004. For adjacent CU 1004, the motion vectors for the upper left corner (v2 1030), upper right corner (v3 1032), and lower left corner (v4 1034) have been derived. In this example, the control point motion vector v01040 for the upper left corner of the current block 1002 is calculated based on v2 1030, v3 1032, and v4 1034. Then the control point motion vector v1 1042 for the upper right corner of the current block 1002 can be determined.

[0160] Once the control point motion vectors (CPMVs) (v0 1040 and v1 1042) of the current block 1002 are derived, equation (1) can be applied to determine the motion vector field for the current block 1002. To identify whether the current block 1002 is encoded in AF_MERGE mode, an affine flag can be included in the bitstream if at least one adjacent block is encoded in affine mode.

[0161] In many cases, the process of affine motion estimation involves determining the affine motion for a block on the encoder side by minimizing the distortion between the original block and the affine motion prediction block. Since affine motion has more parameters than translational motion, affine motion estimation can be more complex. In some cases, a fast affine motion estimation method based on the Taylor expansion of the signal can be performed to determine the affine motion parameters (e.g., affine motion parameters a, b, c, d in a 4-parameter model).

[0162] Fast affine motion estimation can include gradient-based affine motion search. For example, the pixel value I at a given time t. t (where t0 is the time of the reference image) In this case, the pixel value I can be... t The first-order Taylor expansion is determined as follows:

[0163]

[0164] in, and The pixel gradients G in the x and y directions are respectively. 0x G 0y ,and and Indicates the value I used for pixels t motion vector component V x and V y Used for pixel I in the current block t The motion vector points to pixel I in the reference image. to .

[0165] Equation (2) can be rewritten as equation (3) as follows:

[0166] I t =I to +G x0 ·V x +G y0 ·V y Equation (3)

[0167] Then, by making the prediction (I) to +G x0 ·V x +G y0·V y The solution for pixel value I is obtained by minimizing the distortion between the original signal and the pixel value I. t Affine motion V x and V y Taking a 4-parameter affine model as an example,

[0168] V x =a·xb·y+c Equation (4)

[0169] V y =b·x+a·y+d Equation (5)

[0170] Where x and y represent the positions of pixels or sub-blocks. Substituting equations (4) and (5) into equation (3), and then using equation (3) to minimize the distortion between the original signal and the prediction, the solution for the affine parameters a, b, c, and d can be determined:

[0171]

[0172] Once the affine motion parameters (which are defined as the affine motion vectors used for control points) are determined, the motion vectors for each pixel or sub-block can be determined using the affine motion parameters (e.g., using equations (4) and (5), which can also be expressed by equation (1)). Equation (3) can be performed for each pixel of the current block (e.g., CU). For example, if the current block is 16 pixels x 16 pixels, the least-squares solution in equation (6) can be used to derive the affine motion parameters (a, b, c, d) for the current block by minimizing the total value over 256 pixels.

[0173] Any number of parameters can be used in affine motion models used for video data. For example, 6-parameter affine motion or other affine motions can be solved in the same way as described above for a 4-parameter affine motion model. For example, a 6-parameter affine motion model can be described as follows:

[0174]

[0175] In equation (7), (v x ,v y Let f be the motion vector at coordinates (x, y), and a, b, c, d, e, and f are six affine parameters. The affine motion model for the block can also be described by three motion vectors (MV) at the three corners of the block: and

[0176] Figure 11This is a diagram showing the affine motion field of the current block 1102, described by three motion vectors at three control points 1110, 1112, and 1114. Motion vectors At control point 1110, located at the top left corner of the current block 1102, the motion vector... At control point 1112 located at the upper right corner of the current block 1102, and the motion vector At control point 1114, located at the lower left corner of the current block 1102, the motion vector field (MVF) of the current block 1102 can be described by the following equation:

[0177]

[0178] Equation (8) represents a 6-parameter affine motion model, where w and h are the width and height of the current block 1102.

[0179] Although the 4-parameter motion model is described with reference to equation (1) above, a simplified 4-parameter affine model using the width and height of the current block can be described by the following equation:

[0180]

[0181] The simplified 4-parameter affine model for the block based on equation (9) can be described by two motion vectors at two of the four corners of the block: and Therefore, a sports field can be described as:

[0182]

[0183] As mentioned earlier, motion vectors This is referred to as the Control Point Motion Vector (CPMV) in this paper. The CPMV used for a 4-parameter affine motion model is not necessarily the same as the CPMV used for a 6-parameter affine motion model. In some examples, different CPMVs can be chosen for the affine motion model.

[0184] Figure 12 This is a diagram illustrating the selection of control point vectors for the affine motion model used for the current block 1202. Four control points 1210, 1212, 1214, and 1216 are shown for the current block 1202. Motion vectors At control point 1210, located at the top left corner of the current block 1202, the motion vector... At control point 1212, located at the upper right corner of the current block 1202, the motion vector... At control point 1214 located at the lower left corner of the current block 1202, and the motion vector At control point 1216, located at the bottom right corner of the current block 1202.

[0185] In one example, for a 4-parameter affine motion model (according to equation (1) or equation (10)), it can be derived from four motion vectors. Control point pairs can be selected from any two vectors. In another example, for a 6-parameter affine motion model, control points can be selected from four motion vectors. A control point pair is selected from any three motion vectors. Based on the selected control point motion vectors, other motion vectors for the current block 1202 can be calculated, for example, using the derived affine motion model.

[0186] In some examples, alternative affine motion models can also be used. For instance, an affine motion model based on incremental MV can be represented by an anchor MV at coordinates (x0, y0). Horizontal increment and vertical increment MV To represent it. Typically, the MV at coordinates (x, y) can be expressed as... Calculated as

[0187] In some examples, the CPMV-based affine motion model representation can be converted into an alternative affine motion model representation with incremental MV. For example, in the incremental MV affine motion model representation... It is the same as the CPMV in the top left corner: It is important to note that for these vector operations, addition, division, and multiplication are applied element-wise.

[0188] In some examples, an affine motion predictor can be used to perform affine motion vector prediction. In some examples, an affine motion predictor for the current block can be derived from the affine motion vectors or normal motion vectors of adjacent coded blocks. As described above, an affine motion predictor can include an inherited affine motion vector predictor (e.g., inherited using the affine merge (AF_MERGE) mode) and a constructed affine motion vector predictor (e.g., constructed using the affine inter-frame (AF_INTER) mode).

[0189] The inherited affine motion vector predictor (MVP) uses one or more affine motion vectors from neighboring coded blocks to derive the predicted CPMV for the current block. With inherited affine, the current block can share the same affine motion model with neighboring coded blocks. Neighboring coded blocks are called neighboring blocks or candidate blocks. Neighboring blocks can be selected from different spatial or temporal adjacent locations.

[0190] Figure 13This is a diagram showing the inherited affine MVP of the current block 1302 from its neighboring block 1302 (block A). The following is based on the corresponding motion vectors at control points 1320, 1322, and 1324. To represent the affine motion vector of adjacent block 1302: In one example, the size of the adjacent block 1304 can be represented by parameters (w, h), where w is the width of the adjacent block 1304 and h is the height of the adjacent block 1304. The coordinates of the control points of the adjacent block 1304 are represented as (x0, y0), (x1, y1), and (x2, y2). Affine motion vectors at the corresponding control points 1310, 1312, and 1314 can be predicted for the current block 1302. For the affine motion vector predicted for the current block 1302 This can be derived by replacing (x,y) in equation (8) with the coordinate difference between the control point of the current block 1302 and the upper left control point of the adjacent block 1304, as described in the following equation:

[0191]

[0192]

[0193]

[0194] In equations (11)-(13), (x1',y1'), (x1',y1'), and (x2',y2') are the coordinates of the control points of the current block. In some examples, the predicted affine motion can also be represented using the increment MV: as well as

[0195] Similarly, if the affine motion model of an adjacent coded block (e.g., adjacent block 1304) is a 4-parameter affine motion model, then equation (10) can be applied to derive the affine motion vector at the control point of the current block 1302. In some examples, using equation (10) to obtain the 4-parameter affine motion model may include bypassing equation (13) above.

[0196] Figure 14 This is a diagram illustrating the possible locations of neighboring candidate blocks used in the inherited affine MVP model for the current block 1402. For example, the affine motion vectors at control points 1410, 1412, and 1414 of the current block. It can be derived from one of the adjacent blocks 1430 (block A0), 1426 (block B0), 1428 (block B1), 1432 (block A1), and / or 1420 (block B2). In some cases, adjacent blocks 1424 (block A2) and / or 1422 (block B3) can also be used. More specifically, the motion vector at control point 1410 located at the top left corner of the current block 1402. The motion vector can be inherited from the adjacent block 1420 (block B2) located above and to the left of control point 1410, the adjacent block 1422 (block B3) located above control point 1410, or the adjacent block 1424 (block A2) located to the left of control point 1410; the motion vector located at control point 1412 at the upper right corner of the current block 1402. It can inherit from the adjacent block 1426 (block B0) above control point 1410 or the adjacent block 1428 (block B1) above and to the right of control point 1410; and the motion vector at control point 1414 located at the lower left corner of the current block 1402. It can be inherited from the adjacent block 1430 (block A0) located to the left of control point 1410 or the adjacent block 1432 (block A1) located to the left and below control point 1410.

[0197] In some video coding standards, the HMVP buffer (e.g., the HMVP table) cannot be updated using motion information predicted by the CU using an affine motion model, but only using motion information from one or more regular inter-frame coded CUs. However, with the introduction of more sophisticated coding tools, the granularity of coded blocks can be as small as 4x4. Even though in some cases the CU size for affine mode (using an affine motion model) may be constrained to at least 8x8, considering that high-definition (HD) sequences with 2K resolution, ultra-high-definition (UHD) sequences with 4K resolution, or other high-resolution videos (e.g., which may be located in small affine coded blocks close to the translation-coded CUs) can help provide useful information for predicting motion information in the translation-coded CUs. Therefore, as further described herein, the method in this paper can allow motion information from one or more affine coded CUs to be included in the HMVP table used for regular inter-frame prediction modes.

[0198] In some examples, systems, methods (also referred to as processes), and computer-readable media are provided for improving history-based motion vector prediction. For example, in some cases, the HMVP table can be updated using motion information generated and / or utilized in CU coding using an affine motion model. In some implementations, the HMVP table can be updated using: motion information available in an affine coding block (e.g., CU, PU, ​​or other block), which may include motion information associated with control points of the affine block; or sub-block motion vector information derived from control points of the affine block; or motion information derived from the spatiotemporal and / or temporal neighborhood of the affine coding block; or motion information used as a predictor for the affine coding block (e.g., the output of the MVP generated from affine merging candidates).

[0199] In some cases, a history table containing motion vectors and reference indices of one or more previously decoded CUs can be defined as HMVPCandList. In the first illustrative example, the top-left CPMV of an affine coded block and its corresponding reference index (denoted as CPMV_top_left_info) can be inserted into the history table as follows: HMVPCandList = CPMV_top_left_info.

[0200] In another illustrative example, the top-right CPMV of the affine code block and its corresponding reference index (denoted as CPMV_top_right_info) can be inserted into the history table as follows: HMVPCandList = CPMV_top_right_info.

[0201] In some cases, an affine coded CU can be divided into sub-blocks, and motion vectors can be derived for each sub-block. When the affine CU size is large (e.g., larger than a threshold size, such as 8x8, 8x16, 16x8, 16x16, or other sizes), a specific CPMV from a given angle (e.g., top left, top right, or other angle) can be quite different from the motion of more distant sub-blocks. Therefore, in some cases, to generate representative motion vectors for a block, the calculated CPMV can be used as a representation of the motion information for the entire CU. Since the center position of the entire CU can provide an average of the overall motion information for each sub-block, in some examples, motion vectors derived using the center CU position can be inserted into a history table.

[0202] Figure 15 This is a diagram illustrating an exemplary motion vector estimated for an affine coded block and corresponding to the center position of the block. In this example, the current block 1500 is an affine coded block. Motion Vector and Corresponding to control points 1510 and 1512. Specifically, the motion vector... This represents the control point motion vector (CPMV) associated with the top-left control point 1510, and the motion vector... This represents the CPMV associated with the top-right control point 1512. Additionally, in some examples, the motion vector... and It can have corresponding horizontal and vertical values. For example, a motion vector. and v 0x and v 1x The value can be used for the corresponding motion vector (e.g., motion vector). and The level value of ) and v 0y and v 1y The value can be used for the corresponding motion vector (e.g., motion vector). and The vertical value of ). In some cases, additional control points (e.g., four control points, six control points, eight control points, or some other number of control points) can be defined by adding additional control point vectors (e.g., at the bottom corner of the current block 702, the center of the current block 1500, or other locations in the current block 1500).

[0203] Motion vectors corresponding to control points 1510 and 1512 and It can be used to generate motion vectors representing the current block 1500. 1504, which can be stored in an HMVP table (e.g., HMVP table 400) for use when encoding future blocks. As previously mentioned, the affine model uses multiple control points (e.g., control points 1510, 1512) to derive multiple motion vectors for a block. For example, the affine model can generate multiple local motion vectors for sub-blocks or pixels of a block. However, the HMVP table used in some video coding standards may only support and / or include a single translational motion vector per block in the HMVP list. Therefore, in some examples, in order to include motion information generated and / or utilized when encoding a block (e.g., current block 1500) using the affine motion model, additional motion vectors from control points (e.g., control points 1510 and 1512) within the block can be used (e.g., ...). and To generate a single motion vector representing the block (e.g., 1504). The HMVP table can then be updated to include motion vectors generated for affine-coded blocks.

[0204] In some examples, motion vectors 1504 can be based on the motion vectors corresponding to control points 1510 and 1512. and This generates a motion vector for a specific position in the current block 1500. Figure 15 In the example shown, the motion vector 1504 corresponds to sub-block 1502 located at the center of the current block 1500. The center position can reflect the overall motion information of the current block 1500 and / or the average and / or representation of each sub-block in the current block 1500. Therefore, sub-block 1502 at the center position can be used as a motion vector representing the overall motion information of the current block 1500.

[0205] In some cases, it can be based on motion vectors and The difference and / or motion vector between and The rate of change between them is used to generate motion vectors. 1504. For example, motion vectors can be calculated for both vertical and horizontal components. and The rate of change between them. In some cases, motion vectors can be calculated per sample for both vertical and horizontal dimensions. and The rate of change between them. For example, the motion vector can be calculated. and The difference between them is used to determine the total difference, and then the total difference is divided by the difference between the motion vectors. and The number of samples between associated control points to obtain the motion vector and The rate of change per sample between them.

[0206] The calculated rate of change can then be used to estimate the motion vector for a specific location within the current block 1500. For example, due to Figure 15 Sub-block 1502 corresponds to the center position within the current block 1500, so the motion vector for sub-block 1502 can be calculated as follows: divide the width of the current block 1500 by two to obtain the rate of change multiplier corresponding to the center position (e.g., the number of samples between the center position and control points 1510 or 1512, or in other words, the distance in samples from one end of the current block 1500 (e.g., from the left or right boundary) to the center position), and then... and The rate of change between them increases (e.g., by multiplying by the rate of change multiplier).

[0207] For the purpose of explanation, if the motion vector and If the rate of change per sample is x, and the width of the current block 1500 is 8, then the rate of change per sample x can be multiplied by the result obtained by dividing 8 (the width of the current block 1500) by 2. Here, the result of dividing 8 by 2 is 4, which represents the rate of change multiplier for the center position (e.g., the number of samples to the center position), and the motion vector for the sub-block 1502 is 4x (e.g., the rate of change x multiplied by 4). The motion vector for the sub-block 1502 can then be stored in the HMPV table as the translational motion vector corresponding to the current block 1500.

[0208] In some examples, the derivation of the motion vector (denoted as center_subblock_mv) for the affine-coded CU, using the center position of the CU, can be as follows. Given the top-left and top-right CPMVs (denoted as CPMV_top_left and CPMV_top_right, respectively) and the CU width and height denoted as w and h, respectively:

[0209] mvScaleHor = CPMV_top_left_hor << 7 Equation (14)

[0210] mvScaleVer = CPMV_top_left_ver << 7 Equation (15)

[0211] dHorX=(CPMV_top_right_hor-CPMV_top_left_hor)<<(7–log2(w))

[0212] Equation (16)

[0213] dVerX=(CPMV_top_right_ver-CPMV_top_left_ver)<<(7–log2(w))

[0214] Equation (17)

[0215] In the case where a CPMV exists at the bottom left and is represented as CPMV_bottom_left:

[0216] dHorY=(CPMV_bottom_left_hor-CPMV_top_left_hor)<<(7–log2(h))

[0217] Equation (18)

[0218] dVerY=(CPMV_bottom_left_ver-CPMV_top_left_ver)<<(7–log2(h))

[0219] Equation (19)

[0220] otherwise:

[0221] Equation (20) is dHorY = -dVerX

[0222] Equation (21) is: dVerY = dHorX

[0223] Where CPMV_top_left_hor and CPMV_top_left_ver are the horizontal and vertical components of the top-left CPMV (CPMV_top_left); CPMV_top_right_hor and CPMV_top_right_ver are the horizontal and vertical components of the top-right CPMV (CPMV_top_right). The motion vector for the center sub-block can then be derived as follows:

[0224] center_subblock_mv_hor=(mvScaleHor+dHorX*(w>>1)+dHorY*(h>>1))

[0225] Equation (22)

[0226] center_subblock_mv_ver=(mvScaleVer+dVerX*(w>>1)+dVerY*(h>>1))

[0227] Equation (23)

[0228] center_subblock_mv=(center_subblock_mv_hor,center_subblock_mv_ver)

[0229] Equation (24)

[0230] Furthermore, `center_subblock_mv_info` can include the motion vector `center_subblock_mv` and the corresponding reference index (or in some cases, consist of them). The HMVP candidate list can then be updated as follows: `HMVPCandList = center_subblock_mv_inf`. In some examples, Figure 15 The motion vector 1504 in the equation can be the center-sub-block_mv calculated in equations (14)-(24) as described above.

[0231] In some implementations, another sub-block position can be used for HMVP table updates, such as a sub-block position at the top left, top right, bottom left, bottom right, or near the center (e.g., offset from the center). In some implementations, the techniques described herein can insert the motion vector of the bottom left CPMV or a sub-block other than the CU center position sub-block MV. In some cases, a prescriptive procedure such as the prescriptive procedure provided below can be performed.

[0232] In an example of the process for updating a candidate list of history-based motion vector predictors, the inputs may include luminance motion vectors mvL0 and mvL1 with 1 / 16 fractional sample accuracy, reference indices refIdxL0 and refIdxL1, variables cbWidth and cbHeight specifying the width and height of the luminance coding block, the luminance position (xCb, yCb) of the top-left sample of the current luminance coding block relative to the top-left luminance sample of the current image, and a history-based motion information table HmvpCandList. The output of the process may be a modified history-based motion information table HmvpCandList.

[0233] If affine_flag[xCb][yCb] equals 1, numCpMv is set to the number of control point motion vectors, and cpMvLX[cpIdx] is set to the control point motion vector of the current block, where cpIdx = 0, ..., numCpMv – 1 and X is 0 or 1. Then, the horizontal variation dX, the vertical variation dY, and the base motion vector mvBaseScaled are derived as follows: The derivation process for the affine motion model parameters from the control point motion vectors is invoked, where the luminance block width cbWidth, the luminance block height cbHeight, the number of control point motion vectors numCpMv, and the control point motion vector cpMvLX[cpIdx] (where cpIdx = 0, ..., numCpMv – 1) are used as input. The motion vector MvLX can be calculated as follows:

[0234] xPosSb=cbWidth>>1 Equation (25)

[0235] yPosSb=cbHeight>>1 Equation (26)

[0236] mvLX[0]=(mvBaseScaled[0]+dX[0]*xPosSb+dY[0]*yPosSb)

[0237] Equation (27)

[0238] mvLX[1]=(mvBaseScaled[1]+dX[1]*xPosSb+dY[1]*yPosSb)

[0239] Equation (28)

[0240] As further described below, a rounding process for the motion vector can be invoked, where mvX (set to equal mvLX), rightShift (set to equal 5), and leftShift (set to equal 0) are taken as inputs, and the rounded mvLX is taken as the output. Furthermore, the motion vector mvLX can be clipped as follows:

[0241] mvLX[0]=Clip3(-2 17 ,2 17 -1,mvLX[0]) Equation (29)

[0242] mvLX[1]=Clip3(-2 17 ,2 17 -1,mvLX[1]) Equation (30)

[0243] The MVP candidate hMvpCand can include the luminance motion vectors mvL0 and mvL1, and the reference indices refIdxL0 and refIdxL1 (or in some cases, they can be combined). NumHmvpCand can be set to the number of motion entries in HmvpCandList.

[0244] If slice_type equals P and refIdxL0 is valid, or if slice_type equals B and refIdxL0 or refIdxL1 is valid, then the candidate list HmvpCandList is modified using the candidate mvCands through the following steps (in some cases, these steps can be performed in any other order). First, the variable curIdx is set to equal NumHmvpCand. If NumHmvpCand equals 23, then for each index hMvpIdx = 1, ..., NumHmvpCand - 1, HMVPCandList[hMvpIdx] is copied to HMVPCandList[hMvpIdx - 1], and then hMvpCand is copied to HMVPCandList[hMvpIdx]. If NumHmvpCand is less than 23, then NumHmvpCand is incremented by 1.

[0245] The derivation process for the affine motion model parameters from the control point motion vectors mentioned above can be invoked as follows. First, the inputs to the derivation process can include: variables cbWidth and cbHeight specifying the width and height of the luma encoding block, the number of control point motion vectors numCpMv, and the control point motion vector cpMvLX[cpIdx], where cpIdx = 0, ..., numCpMv–1 and X is 0 or 1. The output of this process can include the horizontal variation dX of the motion vectors, the vertical variation dY of the motion vectors, and the motion vector mvBaseScaled corresponding to the top-left corner of the luma encoding block.

[0246] The variables log2CbW and log2CbH can be derived as follows:

[0247] log2CbW=Log2(cbWidth) Equation (31)

[0248] log2CbH=Log2(cbHeight) Equation (32)

[0249] The horizontal change dX of the motion vector can be derived as follows:

[0250] dX[0]=(cpMvLX[1][0]-cpMvLX[0][0])<<(7-log2CbW)

[0251] Equation (33)

[0252] dX[1]=(cpMvLX[1][1]-cpMvLX[0][1])<<(7-log2CbW)

[0253] Equation (34)

[0254] The vertical change of the motion vector dY can be derived as follows. If numCpMv equals 3, then dY can be derived as follows:

[0255] dY[0]=(cpMvLX[2][0]-cpMvLX[0][0])<<(7-log2CbH)

[0256] Equation (35)

[0257] dY[1]=(cpMvLX[2][1]-cpMvLX[0][1])<<(7-log2CbH)

[0258] Equation (36)

[0259] Otherwise (numCpMv equals 2), dY can be derived as follows:

[0260] dY[0]=-dX[1] Equation (37)

[0261] dY[1]=dX[0] Equation (38)

[0262] The motion vector mvBaseScaled corresponding to the top left corner of the luminance coding block can be derived as follows:

[0263] mvBaseScaled[0]=cpMvLX[0][0]<<7 Equation (40)

[0264] mvBaseScaled[1]=cpMvLX[0][1]<<7 Equation (41)

[0265] The rounding procedure for motion vectors mentioned above can be invoked as follows. The inputs to this procedure can include the motion vector mvX, a right-shift parameter (rightShift) for rounding, and a left-shift parameter (leftShift) for resolution improvement. The output of this procedure is the rounded motion vector mvX. For rounding mvX, the following applies:

[0266] Offset = (rightShift == 0)? 0: (1 << (rightShift - 1)) Equation (42)

[0267] mvX[0]=((mvX[0]+offset-(mvX[0]>=0))>>rightShift)< <leftShift

[0268] Equation (43)

[0269] mvX[1]=((mvX[1]+offset-(mvX[1]>=0))>>rightShift)< <leftShift

[0270] Equation (44)

[0271] Figure 16 This is a flowchart illustrating an exemplary method 1600 (also called a process) for updating a history-based motion prediction table using motion information generated from affine coding blocks.

[0272] At box 1602, method 1600 may include: obtaining one or more blocks of video data. For example, the one or more blocks may include blocks encoded using affine motion patterns, as previously described. In some cases, the one or more blocks may include the current block. Furthermore, in some examples, the video data may include a current image and a reference image.

[0273] At box 1604, method 1600 may include: determining a first motion vector (e.g., control point 1510) derived from a first control point (e.g., control point 1510) of one or more blocks (e.g., current block 1500). This block may include blocks encoded using affine motion patterns. At block 1606, method 1600 may include: determining a second motion vector (e.g., ...) derived from a second control point (e.g., control point 1512) of the block. ).

[0274] At box 1608, method 1600 may include estimating a third motion vector (e.g., motion vector 1504) for a predetermined position within the block based on a first motion vector and a second motion vector. In some examples, the third motion vector may be calculated based on the above equations (14)-(24). In some cases, the third motion vector may be a representation of motion information for the block (e.g., for the entire block).

[0275] In some cases, the predetermined location can be the center of the block or a location within the block. In some examples, the first control point can be the top-left control point, and the second control point can be the top-right control point. In some cases, a third motion vector can be further estimated based on the motion vectors of the control points associated with bottom control points such as the bottom-left control point.

[0276] In some examples, estimating the third motion vector may include: determining the rate of change between the first and second motion vectors based on the difference between them, and multiplying that rate of change by a multiplication factor corresponding to a predetermined position. In some cases, the predetermined position may include the center of the block, and the multiplication factor may include the width of the block and / or half the height of the block.

[0277] In some cases, the rate of change can be a rate of change per unit. Furthermore, each unit of the rate of change per unit can include a sample, a sub-block, and / or a pixel. In some examples, the multiplication factor can include the number of samples between a predetermined location and the boundary of the block. For example, if the predetermined location is the center of the block, and the block comprises n samples along its dimensions, the multiplication factor could be n divided by 2, which in this example corresponds to the number of samples to the center of the block.

[0278] In some implementations, the third motion vector may be based on a first motion change between the horizontal component of the first motion vector and the horizontal component of the second motion vector, and a second motion change between the vertical component of the first motion vector and the vertical component of the second motion vector.

[0279] In some examples, the third motion vector may be a translational motion vector generated based on affine motion information associated with the block. Furthermore, the third motion vector may include motion information associated with one or more sub-blocks of the block. In some examples, at least one of the one or more sub-blocks may correspond to a predetermined position.

[0280] At box 1610, method 1600 may include: populating a history-based motion vector predictor (HMVP) table (e.g., HMVP table 400) with a third motion vector. In some cases, the HMVP table and / or the third motion vector may be used in motion prediction for additional blocks (such as future blocks).

[0281] In some aspects, method 1600 may include adding one or more HMVP candidates from the HMVP table to the advanced motion vector prediction (AMVP) candidate list and / or merging pattern candidate list.

[0282] In some examples, the processes described herein may be performed by a computing device or apparatus (e.g., encoding device 104, decoding device 112, and / or any other computing device). In some cases, the computing device or apparatus may include a processor, microprocessor, or other components of a device configured to perform the steps of the processes described herein. In some examples, the computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) comprising video frames. For example, the computing device may include a camera device, which may or may not include a video codec. As another example, the computing device may include a mobile device with a camera (e.g., a camera device such as a digital camera, IP camera, etc., a mobile phone or tablet device including a camera, or another type of device with a camera). In some cases, the computing device may include a display for displaying images. In some examples, the camera or other capturing device that captures video data is separate from the computing device, in which case the computing device receives the captured video data. The computing device may also include a network interface, transceiver, and / or transmitter configured to transmit video data. The network interface, transceiver, and / or transmitter may be configured to transmit Internet Protocol (IP) based data or other network data.

[0283] The processes described herein can be implemented using hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement these processes.

[0284] Furthermore, the processes described herein can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or a combination thereof. As mentioned above, the code can be stored, for example, in the form of a computer program comprising multiple instructions executable by one or more processors on a computer-readable or machine-readable storage medium. The computer-readable or machine-readable storage medium can be non-transitory.

[0285] The encoding techniques discussed herein can be implemented in an example video encoding and decoding system (e.g., System 100). In some examples, the system includes a source device that provides encoded video data to be decoded later by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source and destination devices can include any of a variety of devices, including desktop computers, laptop computers, tablet computers, set-top boxes, mobile phones (e.g., so-called "smartphones"), so-called "smartboards," televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source and destination devices can be equipped for wireless communication.

[0286] A destination device can receive encoded video data to be decoded via a computer-readable medium. The computer-readable medium can be any type of medium or device capable of moving encoded video data from a source device to a destination device. In one example, the computer-readable medium may include a communication medium enabling the source device to transmit encoded video data directly to the destination device in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network (LAN), a wide area network (WAN), or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other means that can be used to facilitate communication from the source device to the destination device.

[0287] In some examples, encoded data can be output from an output interface to a storage device. Similarly, encoded data can be accessed from a storage device via an input interface. The storage device can include any of a variety of distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In other examples, the storage device can correspond to a file server or another intermediate storage device that can store encoded video generated by the source device. The destination device can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing and sending encoded video data to the destination device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device can access the encoded video data via any standard data connection, including an internet connection. This can include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., DSL, cable modems, etc.), or combinations thereof, suitable for accessing encoded video data stored on a file server. Transfer of the encoded video data from the storage device can be streaming, downloading, or a combination thereof.

[0288] The technology disclosed herein is not necessarily limited to wireless applications or setups. The technology can be applied to video encoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video (e.g., HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding digital video stored on data storage media, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0289] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device may include an input interface, a video decoder, and a display device. The video encoder of the source device may be configured to apply the techniques disclosed herein. In other examples, the source and destination devices may include other components or arrangements. For example, the source device may receive video data from an external video source such as an external camera. Similarly, the destination device may interface with an external display device, rather than including an integrated display device.

[0290] The example system described above is merely an example. Techniques for processing video data in parallel can be implemented by any digital video encoding and / or decoding device. While the techniques of this disclosure are generally implemented by video encoding devices, they can also be implemented by a video encoder / decoder, commonly referred to as a "CODEC". Furthermore, the techniques of this disclosure can also be implemented by a video preprocessor. The source and destination devices are merely examples of such encoding devices, where the source device generates encoded video data for transmission to the destination device. In some examples, the source and destination devices can operate in a substantially symmetrical manner, such that each of these devices includes video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0291] Video sources may include video capture devices, such as cameras, video archiving units containing previously captured video, and / or video feed interfaces for receiving video from video content providers. Alternatively, a video source may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In some cases, if the video source is a camera, the source and destination devices may form a so-called camera phone or video phone. However, as described above, the techniques described in this disclosure are generally applicable to video encoding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by a video encoder. The encoded video information can then be output to a computer-readable medium via an output interface.

[0292] As mentioned, computer-readable media can include temporary media such as wireless broadcasting or wired network transmissions, or storage media such as hard disks, flash drives, compressed optical discs, digital versatile optical discs, Blu-ray discs (i.e., non-temporary storage media), or other computer-readable media. In some examples, a network server (not shown) can, for example, receive encoded video data from a source device via a network transmission and provide the encoded video data to a destination device. Similarly, a computing device in a media production facility, such as an optical disc stamping facility, can receive encoded video data from a source device and manufacture an optical disc containing the encoded video data. Therefore, in the various examples, computer-readable media can be understood to include one or more computer-readable media of various forms.

[0293] The input interface of the destination device receives information from a computer-readable medium. The information in the computer-readable medium may include syntax information defined by the video encoder (which is also used by the video decoder), including characteristics of description blocks and other coding units (e.g., groups of pictures (GOPs)) and / or processed syntax elements. The display device displays the decoded video data to the user and may include any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device. Various embodiments of this application have been described.

[0294] exist Figure 17 and Figure 18 The specific details of the encoding device 104 and the decoding device 112 are shown in the figure. Figure 17This is a block diagram illustrating an exemplary encoding device 104 that can implement one or more of the techniques described in this disclosure. Encoding device 104 can, for example, generate the syntax structures described herein (e.g., syntax structures of VPS, SPS, PPS, or other syntax elements). Encoding device 104 can perform intra-frame prediction and inter-frame prediction coding of video blocks within a video slice. As previously described, intra-frame coding relies at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter-frame coding relies at least in part on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames of a video sequence. Intra-frame mode (I-mode) can refer to any of several spatial-based compression modes. Inter-frame modes such as one-way prediction (P-mode) or two-way prediction (B-mode) can refer to any of several temporal-based compression modes.

[0295] Encoding device 104 includes a segmentation unit 35, a prediction processing unit 41, a filter unit 63, an image memory 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-frame prediction processing unit 46. For video block reconstruction, encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62. Filter unit 63 is intended to represent one or more loop filters, such as deblocking filters, adaptive loop filters (ALF), and sample adaptive offset (SAO) filters. Although in Figure 17 The filter unit 63 is shown as an in-loop filter, but in other configurations, it can be implemented as a post-loop filter. The post-processing device 57 can perform additional processing on the encoded video data generated by the encoding device 104. In some cases, the techniques of this disclosure can be implemented by the encoding device 104. However, in other cases, one or more of the techniques of this disclosure can be implemented by the post-processing device 57.

[0296] like Figure 17As shown, encoding device 104 receives video data, and segmentation unit 35 segments the data into video blocks. This segmentation may also include, for example, segmentation into slices, segments, tiles, or other larger units based on the quadtree structure of LCUs and CUs, as well as video block segmentation. Encoding device 104 generally illustrates components for encoding video blocks within a video slice to be encoded. The slice can be divided into multiple video blocks (and may be divided into a set of video blocks referred to as tiles). Prediction processing unit 41 can select one of several possible coding modes for the current video block based on error results (e.g., coding rate and distortion level, etc.), such as one of several intra-frame prediction coding modes or one of several inter-frame prediction coding modes. Prediction processing unit 41 can provide the obtained intra-frame or inter-frame coded block to summer 50 to generate residual block data, and to summer 62 to reconstruct the coded block for use as a reference picture.

[0297] The intra-prediction processing unit 46 within the prediction processing unit 41 can perform intra-prediction coding of the current video block relative to one or more neighboring blocks in the same frame or slice as the current video block to be encoded, to provide spatial compression. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-prediction coding of the current video block relative to one or more prediction blocks in one or more reference pictures, to provide temporal compression.

[0298] Motion estimation unit 42 can be configured to determine an inter-frame prediction mode for video slices based on a predetermined pattern for the video sequence. The predetermined pattern can designate video slices in the sequence as P-slices, B-slices, or GPB-slices. Motion estimation unit 42 and motion compensation unit 44 can be highly integrated, but are shown separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. Motion vectors can, for example, indicate the displacement of a prediction unit (PU) of a video block within the current video frame or picture relative to a prediction block within a reference picture.

[0299] A predicted block is a block found to closely match the PU of the video block to be encoded in terms of pixel difference, which can be determined by the sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some examples, encoding device 104 can calculate values ​​for pixel positions less than integers of a reference image stored in image memory 64. For example, encoding device 104 can interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, motion estimation unit 42 can perform motion search relative to full-pixel positions and fractional-pixel positions and output motion vectors with fractional-pixel precision.

[0300] The motion estimation unit 42 calculates the motion vector for the PU by comparing the position of the PU of the video block in the inter-frame coded slice with the position of the predicted block of the reference picture. Reference pictures can be selected from either a first list of reference pictures (list 0) or a second list of reference pictures (list 1), each of which identifies one or more reference pictures stored in the picture memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.

[0301] Motion compensation performed by motion compensation unit 44 may involve acquiring or generating prediction blocks based on motion vectors determined by motion estimation, possibly performing interpolation with sub-pixel precision. Upon receiving the motion vector of the PU for the current video block, motion compensation unit 44 can locate the prediction block pointed to by the motion vector in a list of reference images. Encoding device 104 forms a residual video block by subtracting the pixel values ​​of the prediction block from the pixel values ​​of the current video block being encoded. The pixel difference forms residual data for that block and may include both luminance difference components and chrominance difference components. Summer 50 represents one or more components performing this subtraction operation. Motion compensation unit 44 may also generate syntax elements associated with video blocks and video slices for use by decoding device 112 when decoding video blocks of video slices.

[0302] As described above, the intra-prediction processing unit 46 can perform intra-prediction on the current block as an alternative to the inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44. Specifically, the intra-prediction processing unit 46 can determine the intra-prediction mode to be used for encoding the current block. In some examples, the intra-prediction processing unit 46 can use various intra-prediction modes to encode the current block, for example, during a separate encoding process, and the intra-prediction processing unit 46 can select a suitable intra-prediction mode from the tested modes. For example, the intra-prediction processing unit 46 can use rate-distortion analysis for various tested intra-prediction modes to calculate rate-distortion values, and can select the intra-prediction mode with the best rate-distortion characteristics from the tested modes. Rate-distortion analysis typically determines the amount of distortion (or error) between the coded block and the original uncoded block encoded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra-prediction processing unit 46 can calculate the ratio based on the distortion and rate for various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for that block.

[0303] In any case, after selecting an intra-prediction mode for a block, the intra-prediction processing unit 46 can provide information indicating the selected intra-prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 can encode the information indicating the selected intra-prediction mode. The encoding device 104 can include in the transmitted bitstream configuration data definitions for various blocks, as well as indications of the most probable intra-prediction modes to be used for each of these contexts, an intra-prediction mode index table, and a modified intra-prediction mode index table. The bitstream configuration data may include multiple intra-prediction mode index tables and multiple modified intra-prediction mode index tables (also referred to as codeword maps).

[0304] After prediction processing unit 41 generates a prediction block for the current video block via inter-frame prediction or intra-frame prediction, encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to transform processing unit 52. Transform processing unit 52 uses a transform (such as Discrete Cosine Transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients. Transform processing unit 52 can transform the residual video data from the pixel domain to the transform domain (such as the frequency domain).

[0305] The transform processing unit 52 can send the obtained transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of these coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can then perform a scan of a matrix including the quantized transform coefficients. Alternatively or additionally, the entropy coding unit 56 can perform this scan.

[0306] After quantization, entropy coding unit 56 performs entropy coding on the quantized transform coefficients. For example, entropy coding unit 56 can perform context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probabilistic interval partitioned entropy (PIPE) coding, or another entropy coding technique. After entropy coding by entropy coding unit 56, the encoded bitstream can be sent to decoding device 112, or archived for later transmission or retrieved by decoding device 112. Entropy coding unit 56 can also perform entropy coding on motion vectors and other syntax elements used for the current video slice being encoded.

[0307] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual block in the pixel domain for use as a reference block later as a reference image. Motion compensation unit 44 can calculate the reference block by adding the residual block to a predicted block of one of the reference images in the reference image list. Motion compensation unit 44 can also apply one or more interpolation filters to the reconstructed residual block to calculate pixel values ​​below the integer level for motion estimation. Summer 62 adds the reconstructed residual block to the motion-compensated predicted block produced by motion compensation unit 44 to produce a reference block for storage in image memory 64. The reference block can be used by motion estimation unit 42 and motion compensation unit 44 as a reference block for inter-frame prediction of blocks in subsequent video frames or images.

[0308] Encoding device 104 can perform any of the techniques described herein. Some techniques of this disclosure have been generally described with respect to encoding device 104, but as mentioned above, some of the techniques of this disclosure can also be implemented by post-processing device 57.

[0309] Figure 17 Encoding device 104 represents an example of a video encoder configured to perform one or more of the transform coding techniques described herein. Encoding device 104 can perform any of the techniques described herein, including those mentioned above. Figure 16 The process described.

[0310] Figure 18 This is a block diagram illustrating an exemplary decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, a filter unit 91, and an image memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-frame prediction processing unit 84. In some examples, the decoding device 112 may perform operations generally related to... Figure 17 The encoding stage described by the encoding device 104 is the opposite of the decoding stage.

[0311] During the decoding process, decoding device 112 receives an encoded video bitstream sent by encoding device 104, which represents video blocks of encoded video slices and associated syntax elements. In some embodiments, decoding device 112 may receive the encoded video bitstream from encoding device 104. In some embodiments, decoding device 112 may receive the encoded video bitstream from network entity 79 (such as a server, a media-aware network element (MANE), a video editor / stitcher, or other such device configured to implement one or more of the technologies described above). Network entity 79 may or may not include encoding device 104. Before sending the encoded video bitstream to decoding device 112, network entity 79 may implement some of the technologies described in this disclosure. In some video decoding systems, network entity 79 and decoding device 112 may be part of a separate device, while in other cases, the functionality described with respect to network entity 79 may be performed by the same device including decoding device 112.

[0312] The entropy decoding unit 80 of the decoding device 112 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 can receive syntax elements at the video slice level and / or video block level. The entropy decoding unit 80 can process and parse both fixed-length and variable-length syntax elements in more parameter sets such as VPS, SPS, and PPS.

[0313] When a video slice is encoded as an intra-coded (I) slice, the intra-prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video slice based on the intra-prediction mode signaled by the signal and data from previously decoded blocks from the current frame or picture. When a video frame is encoded as an inter-coded (i.e., B, P, or GPB) slice, the motion compensation unit 82 of the prediction processing unit 81 generates a prediction block for the video block of the current video slice based on motion vectors received from the entropy decoding unit 80 and other syntax elements. A prediction block can be generated from one of the reference pictures in the reference picture list. The decoding device 112 can construct the reference frame list, i.e., list 0 and list 1, based on the reference pictures stored in the picture memory 92 using the default construction technique.

[0314] Motion compensation unit 82 determines prediction information for video blocks in the current video slice by parsing motion vectors and other syntax elements, and uses this prediction information to generate prediction blocks for the current video slice being decoded. For example, motion compensation unit 82 may use one or more syntax elements in the parameter set to determine the prediction mode (e.g., intra-frame or inter-frame prediction) for encoding video blocks in the video slice, the inter-frame prediction slice type (e.g., B-slice, P-slice, or GPB-slice), construction information for one or more reference picture lists for the slice, motion vectors for each inter-frame coded video block in the slice, inter-frame prediction state for each inter-frame coded video block in the slice, and other information for decoding video blocks in the current video slice.

[0315] The motion compensation unit 82 can also perform interpolation based on an interpolation filter. The motion compensation unit 82 can use the interpolation filter used by the encoding device 104 during the encoding of the video block to calculate the interpolated values ​​for pixels less than an integer value for the reference block. In this case, the motion compensation unit 82 can determine the interpolation filter used by the encoding device 104 based on the received syntax elements, and can use the interpolation filter to generate a prediction block.

[0316] The inverse quantization unit 86 inverse-quantizes or dequantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include determining the degree of quantization using quantization parameters calculated by the encoding device 104 for each video block in the video slice, and similarly determining the degree of inverse quantization to be applied. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT or other suitable inverse transform), an inverse integer transform, or a conceptually similar inverse transform process to the transform coefficients to produce residual blocks in the pixel domain.

[0317] After the motion compensation unit 82 generates a prediction block for the current video block based on motion vectors and other syntax elements, the decoding device 112 forms a decoded video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82. The summer 90 represents one or more components performing this summation operation. Loop filters (in or after the coding loop) can also be used, if needed, to smooth pixel transitions or otherwise improve video quality. The filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although in Figure 18The filter unit 91 is shown as a filter in the loop, but in other configurations, filter unit 91 can be implemented as a post-loop filter. The decoded video block from a given frame or image is then stored in image memory 92, which stores reference images for subsequent motion compensation. Image memory 92 also stores the decoded video for later display on a display device (such as a monitor). Figure 1 The video is displayed on the destination device 122 shown in the figure.

[0318] Figure 18 Decoding device 112 represents an example of a video decoder configured to perform one or more of the transform coding techniques described herein. Decoding device 112 can perform any of the techniques described herein, including those mentioned above. Figure 16 Method 1600 is described.

[0319] In the foregoing description, various aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that the subject matter of this application is not limited thereto. Therefore, although illustrative embodiments of this application have been described in detail herein, it should be understood that the inventive concept may be embodied and employed in other ways, and the appended claims are intended to be construed as including such variations, in addition to those limited by the prior art. Various features and aspects of the foregoing subject matter may be used individually or collectively. Furthermore, embodiments may be used in any number of environments and applications other than those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings are to be considered illustrative rather than restrictive. For illustrative purposes, the methods have been described in a particular order. It should be understood that, in alternative embodiments, the methods may be performed in a different order than that described.

[0320] It will be understood by those skilled in the art that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.

[0321] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing a circuit or other hardware to perform the operation, programming a programmable circuit (e.g., a microprocessor or other suitable circuit) to perform the operation, or any combination thereof.

[0322] The language of a claim that states "at least one" and / or "one or more" in a set indicates that one or more members of that set (in any combination) satisfy the claim. For example, the language of a claim stating "at least one of A and B" means A, B, or A and B. In another example, the language of a claim stating "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The use of "at least one" and / or "one or more" in a language set does not limit the set to items listed in that set. For example, the language of a claim stating "at least one of A and B" could mean A, B, or A and B, and could additionally include items not listed in the set of A and B.

[0323] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been generally described above regarding their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0324] The techniques described herein can also be implemented using electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, mobile phones with wireless communication devices, or integrated circuit devices with multiple uses (including applications in mobile phones with wireless communication devices and other devices). Any feature described as a module or component can be implemented together in an integrated logic device, or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code that includes instructions for performing one or more of the methods described above when executed. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Alternatively or concurrently, the technology may be implemented, at least in part, by a computer-readable communication medium (such as a propagating signal or wave) that carries or transmits program code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer.

[0325] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated software or hardware modules configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).

Claims

1. An apparatus for processing video data, the apparatus comprising: Memory; as well as One or more processors coupled to the memory, the one or more processors being configured to: Obtain one or more blocks of video data; Determine a first motion vector derived from a first control point of one or more blocks, the blocks being encoded using an affine motion pattern; Determine the second motion vector derived from the second control point of the block; A third motion vector for the center of the block is estimated based on the first motion vector and the second motion vector, wherein the third motion vector for estimating the center of the block is an estimate of the translational motion of the block; and The estimated third motion vector for the center of the block is used to populate the history-based motion vector predictor HMVP table.

2. The apparatus according to claim 1, wherein, The first control point includes a control point at the upper left, and the second control point includes a control point at the upper right.

3. The apparatus according to claim 2, wherein, The third motion vector is further estimated based on the control point motion vector associated with the bottom control point.

4. The apparatus according to claim 1, wherein, Estimating the third motion vector includes: The rate of change between the first motion vector and the second motion vector is determined based on the difference between them; and Multiply the rate of change by a multiplication factor corresponding to the center of the block.

5. The apparatus according to claim 4, wherein, The multiplication factor includes half of at least one of the width and height of the block.

6. The apparatus according to claim 4, wherein, The rate of change includes a rate of change per unit, wherein each unit of the rate of change per unit includes at least one of a sample, a sub-block, and a pixel.

7. The apparatus according to claim 4, wherein, The multiplication factor includes the number of samples between the center of the block and the boundary of the block.

8. The apparatus according to claim 1, wherein, The third motion vector is based on a first motion change between the horizontal components of the first motion vector and the horizontal components of the second motion vector, and a second motion change between the vertical components of the first motion vector and the vertical components of the second motion vector.

9. The apparatus according to claim 1, wherein, The third motion vector includes motion information associated with one or more sub-blocks of the block, wherein at least one of the one or more sub-blocks corresponds to the center of the block.

10. The apparatus according to claim 1, wherein, At least one of the HMVP table and the third motion vector is used for motion prediction of the additional block.

11. The apparatus according to claim 1, wherein, The one or more processors are configured to: One or more HMVP candidates from the HMVP table are added to at least one of the Advanced Motion Vector Prediction (AMVP) candidate list, the merged pattern candidate list, and the motion vector prediction predictor for encoding using the affine motion pattern.

12. The apparatus according to claim 1, wherein, Processing video data includes encoding the video data, wherein the one or more processors are configured to: An encoded video bitstream is generated, the encoded video bitstream comprising the one or more blocks of video data.

13. The apparatus according to claim 12, wherein, The one or more processors are configured to: The encoded video bitstream, the first control point, and the second control point are transmitted.

14. The apparatus according to claim 1, wherein, The one or more processors are configured to: The affine motion pattern is used to decode one or more blocks of video data.

15. The apparatus according to claim 1, wherein, The device is a mobile device.

16. A method for processing video data, the method comprising: Obtain one or more blocks of video data; Determine a first motion vector derived from a first control point of one or more blocks, the blocks being encoded using an affine motion pattern; Determine the second motion vector derived from the second control point of the block; A third motion vector for the center of the block is estimated based on the first motion vector and the second motion vector, wherein the third motion vector for estimating the center of the block is an estimate of the translational motion of the block; and The estimated third motion vector for the center of the block is used to populate the history-based motion vector predictor HMVP table.

17. The method according to claim 16, wherein, The first control point includes a control point at the upper left, and the second control point includes a control point at the upper right.

18. The method according to claim 17, wherein, The third motion vector is further estimated based on the control point motion vector associated with the bottom control point.

19. The method of claim 16, wherein, Estimating the third motion vector includes: The rate of change between the first motion vector and the second motion vector is determined based on the difference between them; and Multiply the rate of change by a multiplication factor corresponding to the center of the block.

20. The method according to claim 19, wherein, The multiplication factor includes half of at least one of the width and height of the block.

21. The method according to claim 19, wherein, The rate of change includes a rate of change per unit, wherein each unit of the rate of change per unit includes at least one of a sample, a sub-block, and a pixel.

22. The method according to claim 19, wherein, The multiplication factor includes the number of samples between the center of the block and the boundary of the block.

23. The method according to claim 16, wherein, The third motion vector is based on a first motion change between the horizontal components of the first motion vector and the horizontal components of the second motion vector, and a second motion change between the vertical components of the first motion vector and the vertical components of the second motion vector.

24. The method of claim 16, wherein, The third motion vector includes motion information associated with one or more sub-blocks of the block, wherein at least one of the one or more sub-blocks corresponds to the center of the block.

25. The method according to claim 16, wherein, At least one of the HMVP table and the third motion vector is used for motion prediction of the additional block.

26. The method of claim 16, further comprising: One or more HMVP candidates from the HMVP table are added to at least one of the Advanced Motion Vector Prediction (AMVP) candidate list, the merged pattern candidate list, and the motion vector prediction predictor for encoding using the affine motion pattern.

27. The method according to claim 16, wherein, The third motion vector includes a representation of the motion information of the block.

28. The method according to claim 16, wherein, Processing video data includes encoding the video data, and the method further includes: An encoded video bitstream is generated, the encoded video bitstream comprising the one or more blocks of video data.

29. The method of claim 28, further comprising: The encoded video bitstream, the first control point, and the second control point are transmitted.

30. The method of claim 16, further comprising: The affine motion pattern is used to decode one or more blocks of video data.

31. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing the one or more processors to perform the following operations when executed by one or more processors: Obtain one or more blocks of video data; Determine a first motion vector derived from a first control point of one or more blocks, the blocks being encoded using an affine motion pattern; Determine the second motion vector derived from the second control point of the block; A third motion vector for the center of the block is estimated based on the first motion vector and the second motion vector, wherein the third motion vector for estimating the center of the block is an estimate of the translational motion of the block; and The estimated third motion vector for the center of the block is used to populate the history-based motion vector predictor HMVP table.