Affine coding with vector clipping
Patent Information
- Application Number
- CN202080065276.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-25
- Filing Date
- 2020-09-28
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2040-09-28
AI Technical Summary
因此,满足这些需求所需要的大量视频数据为处理和存储视频数据的通信网络和设备带来了负担
Smart Images

Figure CN114402617B_ABST
Abstract
Description
Technical Field
[0001] This application relates to video decoding and compression. More specifically, this application relates to affine decoding patterns for video decoding and compression. Background Technology
[0002] Many devices and systems allow video data to be processed and output for consumption. Digital video data typically comprises large amounts of data to meet the needs of video consumers and providers. For example, video data consumers expect high-quality, high-fidelity, high-resolution, high-frame-rate video. Therefore, the large amounts of video data required to meet these needs place a burden on the communication networks and equipment that process and store video data.
[0003] Various video decoding techniques can be used to compress video data. Video decoding techniques can be performed according to one or more video decoding standards. For example, video decoding standards include High Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), Moving Picture Experts Group (MPEG) Part 2 Decoding, VP9, Open Media Consortium (AOMedia) Video 1 (AV1), and Basic Video Decoding (EVC). Video decoding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that take advantage of redundancy present in video images or sequences. A key goal of video decoding technology is to compress video data to a form using a lower bit rate while avoiding or minimizing video quality degradation. As more and more video services become available, there is a need for coding techniques with improved decoding accuracy and efficiency. Summary of the Invention
[0004] This paper describes systems and methods for improving video processing. In some examples, video decoding techniques using affine decoding patterns are described to efficiently encode and decode video data.
[0005] In one illustrative example, a method for decoding video data is described. The method includes: obtaining a current decoding block from the video data; determining control data for the current decoding block; determining one or more affine motion vector clipping parameters based on the control data; selecting a sample of the current decoding block; determining an affine motion vector for the sample of the current decoding block; and clipping the affine motion vector using the one or more affine motion vector clipping parameters to generate a clipped affine motion vector.
[0006] In another illustrative example, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to: obtain a current decoding block from video data; determine control data for the current decoding block; determine one or more affine motion vector clipping parameters based on the control data; select a sample of the current decoding block; determine an affine motion vector for the sample of the current decoding block; and clip the affine motion vector using the one or more affine motion vector clipping parameters to generate a clipped affine motion vector.
[0007] In another illustrative example, a different apparatus for decoding video data is described. The apparatus includes: a unit for obtaining a current decoding block from the video data; a unit for determining control data for the current decoding block; a unit for determining one or more affine motion vector clipping parameters based on the control data; a unit for selecting a sample of the current decoding block; a unit for determining an affine motion vector for the sample of the current decoding block; and a unit for clipping the affine motion vector using the one or more affine motion vector clipping parameters to generate a clipped affine motion vector.
[0008] In another illustrative example, an apparatus for decoding video data is described. The apparatus includes: a memory; and one or more processors coupled to the memory, the processors being configured to: obtain a current decoding block from the video data; determine control data for the current decoding block; determine one or more affine motion vector clipping parameters based on the control data; select a sample of the current decoding block; determine an affine motion vector for the sample of the current decoding block; and clip the affine motion vector using the one or more affine motion vector clipping parameters to generate a clipped affine motion vector.
[0009] In some aspects, the control data includes: a position having associated horizontal and associated vertical coordinates within the full sample cell; a width variable specifying the width of the current decoded block; a height variable specifying the height of the current decoded block; horizontal variations of the motion vector; vertical variations of the motion vector; and a base scaling motion vector. In some examples, the control data may also include the height of the image in the sample associated with the current decoded block and the width of the image in the sample.
[0010] In some aspects, the one or more affine motion vector clipping parameters include: a horizontal maximum variable; a horizontal minimum variable; a vertical maximum variable; and a vertical minimum variable. In some aspects, the horizontal minimum variable is defined by the maximum value selected from the horizontal minimum image value and the horizontal minimum motion vector value.
[0011] In some aspects, the minimum horizontal image value is determined based on the associated horizontal coordinates. In some aspects, the minimum horizontal motion vector value is determined based on a center motion vector value, an array of values based on resolution values or block size (e.g., current decoded block width x height) associated with the video data, and a width variable specifying the width of the current decoded block. In some aspects, the center motion vector value is determined based on the base scaling motion vector, the horizontal variation of the motion vector, the width variable, and the height variable. In some aspects, the base scaling motion vector corresponds to the top-left corner of the current decoded block and is determined based on control point motion vector values. In some aspects, the maximum vertical variable is defined by the minimum value selected from the maximum vertical image value and the maximum vertical motion vector value.
[0012] In some aspects, the maximum vertical image value is determined based on the image's height, the associated vertical coordinates, and the height variable. In some aspects, the maximum vertical motion vector value is determined based on a center motion vector value, an array of values based on resolution values or block size (e.g., current decoded block width x height) associated with the video data, and the height variable specifying the width of the current decoded block.
[0013] In some aspects, the example involves sequentially obtaining multiple current decoded blocks from the video data; determining a set of affine motion vector clipping parameters for each block within the multiple current decoded blocks; and retrieving a portion of a corresponding reference image for each block using the set of affine motion vector clipping parameters.
[0014] In some aspects, the example identifies a reference image associated with the current decoding block; and stores a portion of the reference image defined by the one or more affine motion vector clipping parameters. In some aspects, the example uses reference image data derived from the reference image indicated by the clipped affine motion vectors to process the current decoding block.
[0015] In some aspects, the affine motion vector for the sample of the current decoded block is determined based on: a first base scaling motion vector value, a first horizontal change in the motion vector value, a first vertical change in the motion vector value, a second base scaling motion vector value, a second horizontal change in the motion vector value, a second vertical change in the motion vector value, the horizontal coordinate of the sample, and the vertical coordinate of the sample. In some such aspects, the control data includes values from a derivation table.
[0016] In some aspects, the aforementioned apparatus may include a mobile device having a camera for capturing one or more images. In some aspects, the aforementioned apparatus may include a display for displaying one or more images. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to define the scope of the claimed subject matter. The subject matter should be understood by referring to appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0017] The foregoing, as well as other features and embodiments, will become more apparent upon reference to the following description, claims, and drawings. Attached Figure Description
[0018] Examples of various implementations are described in detail below with reference to the following figures:
[0019] Figure 1 This is a block diagram illustrating encoding and decoding devices according to some examples;
[0020] Figure 2A This is a conceptual diagram illustrating spatially adjacent motion vector candidates for merging patterns, based on some examples;
[0021] Figure 2B This is a conceptual diagram illustrating spatially adjacent motion vector candidates for an Advanced Motion Vector Prediction (AMVP) model, based on some examples.
[0022] Figure 3A This is a conceptual diagram illustrating some examples of Time Motion Vector Predictor (TMVP) candidates;
[0023] Figure 3B This is a conceptual diagram illustrating motion vector scaling based on some examples;
[0024] Figure 4 This is a diagram showing a table of historical motion vector predictors (HMVP) based on some examples;
[0025] Figure 5 This is a graph illustrating the retrieval of non-adjacent space merge candidates based on some examples;
[0026] Figure 6A This is a diagram showing the spatial and temporal locations utilized in MVP prediction based on some examples;
[0027] Figure 6B This is a graph illustrating various aspects of spatial and temporal location utilized in MVP prediction based on some examples;
[0028] Figure 6C This is a diagram illustrating the access order for a spatial MVP (S-MVP) based on some examples;
[0029] Figure 6D This is a diagram illustrating spatial inverse pattern substitutions based on some examples;
[0030] Figure 7 This is a diagram illustrating a simplified affine motion model for the current block, based on some examples;
[0031] Figure 8 This is a diagram showing the motion vector field of sub-blocks of a block according to some examples;
[0032] Figure 9 This is a diagram illustrating motion vector predictions based on some examples of affine inter-frame (AF_INTER) mode;
[0033] Figure 10A and Figure 10B This is a graph showing motion vector predictions under the affine merging (AF_MERGE) mode based on some examples;
[0034] Figure 11 This is a diagram illustrating an affine motion model for the current block based on some examples;
[0035] Figure 12 This is a diagram illustrating another affine motion model for the current block, based on some examples;
[0036] Figure 13 This is a diagram showing the current block and candidate blocks based on some examples;
[0037] Figure 14 This is a diagram showing the current block, the control point of the current block, and candidate blocks based on some examples;
[0038] Figure 15 This is a diagram illustrating the affine model and spatial neighborhood in MPEG5 EVC based on some examples;
[0039] Figure 16 It is a diagram showing various aspects of the affine model and spatial neighborhood based on some examples;
[0040] Figure 17 It is a diagram showing various aspects of the affine model and spatial neighborhood based on some examples;
[0041] Figure 18A This is a diagram illustrating various aspects of clipping using thresholds based on some examples;
[0042] Figure 18B This is a diagram illustrating various aspects of clipping using thresholds based on some examples;
[0043] Figure 18C This is a diagram illustrating various aspects of clipping using thresholds based on some examples;
[0044] Figure 19 This is a flowchart illustrating the decoding process using affine patterns according to the example described herein;
[0045] Figure 20 This is a block diagram illustrating a video encoding device according to some examples; and
[0046] Figure 21 This is a block diagram illustrating a video decoding device based on some examples. Detailed Implementation
[0047] Certain aspects and embodiments of this disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments can be applied independently, and some can be applied in combination. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments may be practiced without these specific details. The accompanying drawings and description are not intended to be limiting.
[0048] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the subsequent description of exemplary embodiments will provide those skilled in the art with a feasible description for implementing the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.
[0049] As described above, this paper describes examples for improving video processing. In some examples, video decoding techniques are described that use affine decoding patterns to efficiently encode and decode video data. An affine model is a model that can be used to approximate a streaming pattern associated with a specific type of image motion in a video, particularly a streaming pattern associated with camera motion (e.g., viewpoint motion or capture position for the video stream). Video processing systems may include affine decoding patterns configured to decode video using affine motion models. Additional details about affine patterns used for video decoding are described below. Examples described herein include the operation and structure of video decoding devices that improve operation by improving memory bandwidth usage in affine decoding patterns. In some examples, memory bandwidth improvements are achieved by cropping the motion vectors used by the affine decoding pattern (which can reduce the data used in the local buffer by limiting the possible reference area (e.g., and associated data) used for affine decoding).
[0050] Some systems use per-sample motion vector generation, which can significantly increase the number of memory access operations required to retrieve filter samples for affine decoding. While a system can handle a large number of retrieval operations if the local buffer can hold the reference data, memory bandwidth usage can degrade system performance if the reference data for each retrieval is large (e.g., exceeding the size of the local buffer, such as the size of the decoded image buffer). By limiting the memory bandwidth usage associated with reference image accesses, a large number of retrieval operations can be used without degrading memory bandwidth performance, thereby improving device operation. The examples described in this paper can provide such benefits in the context of larger video decoding systems and as part of a video decoding device.
[0051] Video decoding devices implement video compression techniques to efficiently encode and decode video data. Video compression techniques may include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques to reduce or remove redundancy inherent in the video sequence. A video encoder can segment each frame of the original video sequence into rectangular regions, which are called video blocks or decoding units (described in more detail below). These video blocks may be encoded using specific prediction modes.
[0052] Video blocks can be divided into one or more smaller blocks in one or more ways. Blocks may include decode tree blocks, prediction blocks, transform blocks, and / or other suitable blocks. Unless otherwise specified, the reference to “block” generally refers to such a video block (e.g., decode tree block, decode block, prediction block, transform block, or other suitable block or sub-block, as will be understood by one of ordinary skill in the art). Furthermore, each of these blocks may also be interchangeably referred to herein as a “unit” (e.g., decode tree unit (CTU), decode unit, prediction unit (PU), transform unit (TU), etc.). In some cases, a unit may refer to a decoded logic unit encoded in the bitstream, while a block may refer to a portion of the process targeted in the video frame buffer.
[0053] For inter-frame prediction mode, the video encoder searches for blocks similar to the block being encoded in a frame (or picture) located at another time position (called a reference frame or reference picture). The video encoder can limit the search to a certain spatial displacement from the block to be encoded. The best match can be located using two-dimensional (2D) motion vectors that include horizontal and vertical displacement components. For intra-frame prediction mode, the video encoder can use spatial prediction techniques to form a prediction block based on data from previously encoded adjacent blocks within the same picture.
[0054] A video encoder can determine prediction error. For example, prediction can be determined as the difference between the pixel values in the block being encoded and the predicted block. Prediction error can also be referred to as residual. The video encoder can also apply a transform to the prediction error using transform decoding (e.g., using the form of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), or other suitable transforms) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and motion vectors can be represented using syntax elements and, together with control information, form a decoded representation of the video sequence. In some cases, the video encoder can perform entropy decoding on the syntax elements, further reducing the number of bits required for its representation.
[0055] The video decoder can use the syntax elements and control information discussed above to construct prediction data (e.g., prediction blocks) for decoding the current frame. For example, the video decoder can add the prediction blocks to the compressed prediction error. The video decoder can determine the compressed prediction error by weighting the transform basis function with quantized coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.
[0056] As described in more detail below, this document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively, the “Systems and Techniques”) for providing improved predictions of history-based motion vectors. The techniques described herein can be applied to one or more of a variety of block-based video decoding techniques, where video is reconstructed on a block-by-block basis. For example, the systems and techniques described herein can be applied to any of existing video codecs (e.g., High Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), or other suitable existing video codecs), and / or can be efficient decoding tools for any video decoding standards under development and / or future video decoding standards, such as Universal Video Decoding (VVC), Joint Exploration Model (JEM), VP9, AV1, Basic Video Decoding (EVC), and / or other video decoding standards under development or to be developed.
[0057] This paper will discuss various aspects of the systems and techniques described in this paper, with reference to the figures. Figure 1 This is a block diagram illustrating an example of a system 100 including an encoding device 104 and a decoding device 112, which can operate in affine decoding mode, according to the examples described herein. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device (also referred to as a client device). The source device and / or receiving device may include electronic devices such as mobile or landline handsets (e.g., smartphones, cellular phones, etc.), desktop computers, laptop or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, Internet Protocol (IP) cameras, server devices in server systems (e.g., video streaming server systems, or other suitable server systems) including one or more server devices, head-mounted displays (HMDs), head-up displays (HUDs), smart glasses (e.g., virtual reality (VR) glasses, augmented reality (AR) glasses, or other smart glasses), or any other suitable electronic devices.
[0058] The components of system 100 may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs) and / or other suitable circuits), and / or may include and / or be implemented using computer software, firmware or any combination thereof to perform the various operations described herein.
[0059] Although system 100 is shown to include certain components, those skilled in the art will understand that system 100 may include more than [other components]. Figure 1The components shown may include more or fewer components. For example, in some cases, system 100 may also include one or more memory devices (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices) in addition to storage devices 108 and 118; one or more processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices) communicating with and / or electrically connected to one or more memory devices; one or more wireless interfaces for performing wireless communication (e.g., including one or more transceivers and baseband processors for each wireless interface); one or more wired interfaces for performing communication over one or more hardwired connections (e.g., serial interfaces such as Universal Serial Bus (USB) inputs, Lightning connectors, and / or other wired interfaces); and / or Figure 1 Other components not shown.
[0060] The decoding techniques described herein are applicable to video decoding in a variety of multimedia applications, including streaming video (e.g., via the Internet), television broadcasting or transmission, encoding digital video for storage on data storage media, decoding digital video stored on data storage media, or other applications. In some examples, system 100 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.
[0061] Encoding device 104 (or encoder) can be used to encode video data using video decoding standards or protocols to generate an encoded video bitstream. Examples of video decoding standards include ITU-T H.261, ISO / IEC MPEG-1 Video, ITU-T H.262, or ISO / IEC MPEG-2 Video, ITU-T H.263, ISO / IEC MPEG-4 Video, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its Scalable Video Decoding (SVC) and Multi-View Video Decoding (MVC) extensions), and High Efficiency Video Decoding (HEVC) or ITU-T H.265. Various extensions to HEVC for multi-layer video decoding exist, including range and screen content decoding extensions, 3D video decoding (3D-HEVC), multi-view extensions (MV-HEVC), and scalable extensions (SHVC). HEVC and its extensions have been developed by the ITU-T Video Decoding Experts Group (VCEG) and the Joint Collaborative Team for Video Decoding (JCT-VC) of the ISO / IEC Moving Picture Experts Group (MPEG), as well as the Joint Collaborative Team for the Development of 3D Video Decoding Extensions (JCT-3V).
[0062] MPEG and ITU-T VCEG have also formed the Joint Video Exploration Group (JVET) to explore and develop new video decoding tools for next-generation video decoding standards, named Universal Video Decoding (VVC). The reference software is called the VVC Test Model (VTM). The goal of VVC is to provide significant improvements in compression performance compared to the existing HEVC standard, thereby facilitating the deployment of higher-quality video services and emerging applications (e.g., 360° omnidirectional immersive multimedia, high dynamic range (HDR) video, etc.). VP9, Open Media Consortium (AOMedia) Video 1 (AV1), and Basic Video Decoding (EVC) are other video decoding standards to which the technologies described herein can be applied.
[0063] Many of the embodiments described herein can be performed using video codecs such as VTM, VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein are also applicable to other decoding standards, such as MPEG, JPEG (or other decoding standards for still images), VP9, AV1, their extensions, or other suitable decoding standards that are already available or not yet available or under development. Therefore, although the techniques and systems described herein may be described with reference to a specific video decoding standard, it will be understood by those skilled in the art that the description should not be construed as applicable only to that particular standard.
[0064] refer to Figure 1 Video source 102 can provide video data to encoding device 104. Video source 102 can be part of a source device, or it can be part of a device other than a source device. Video source 102 can include video capture devices (e.g., cameras, camera phones, video phones, etc.), video archiving devices containing stored video, video servers or content providers that provide video data, video feed interfaces that receive video from video servers or content providers, computer graphics systems for generating computer graphics video data, combinations of such sources, or any other suitable video source.
[0065] Video data from video source 102 may include one or more input images. Images may also be referred to as "frames." An image or frame is a still image that, in some cases, is part of the video. In some examples, data from video source 102 may be still images that are not part of the video. In HEVC, VVC, and other video decoding specifications, a video sequence may include a series of images. An image may include a three-sample array, represented as S. L S Cb and S Cr S L It is a two-dimensional array of brightness samples, SCb It is a two-dimensional array of Cb chromaticity samples, and S Cr It is a two-dimensional array of chrominance samples. Chrominance samples may also be referred to as "chroma" samples in this paper. In other cases, the image may be monochrome and may consist only of an array of luminance samples.
[0066] Encoding device 104's encoder engine 106 (or encoder) encodes the video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more decoded video sequences. The decoded video sequence (CVS) comprises a series of access units (AUs) that begin with an AU that has a random access point picture in the base layer and possesses certain attributes, and continue until, but do not include, the next AU that has a random access point picture in the base layer and possesses certain attributes. For example, certain attributes of the random access point picture that begins the CVS may include a RASL flag equal to 1 (e.g., NoRaslOutputFlag). Otherwise, a random access point picture (where the RASL flag is equal to 0) does not begin the CVS. An access unit (AU) comprises one or more decoded pictures and corresponding control information for decoded pictures sharing the same output time. Decoded slices of pictures are encapsulated as data units at the bitstream level, referred to as Network Abstraction Layer (NAL) units. For example, an HEVC video bitstream may include one or more CVSs, which include NAL units. Each NAL unit has a NAL unit header. In one example, the header is one byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. The syntax elements in the NAL unit header use specified bits and are therefore visible to all kinds of systems and transport layers, such as transport streams, Real-Time Transport (RTP) protocols, file formats, and others.
[0067] The HEVC standard contains two types of NAL units: Video Decoding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain decoded picture data that forms the decoded video bitstream. For example, a VCL NAL unit contains a sequence of bits that form the decoded video bitstream. A VCL NAL unit may include a slice or fragment of the decoded picture data (described below), while non-VCL NAL units contain control information related to one or more decoded pictures. In some cases, NAL units may be referred to as packets. A HEVC AU includes: VCL NAL units containing decoded picture data, and non-VCL NAL units (if any) corresponding to the decoded picture data. Among other information, non-VCL NAL units may also contain a set of parameters with high-level information related to the encoded video bitstream. For example, the parameter set may include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). In some cases, each slice or other portion of the bitstream may reference a single valid PPS, SPS, and / or VPS to allow the decoding device 112 to access information that can be used to decode the slice or other portion of the bitstream.
[0068] NAL units can contain bit sequences that form a decoded representation of video data (e.g., an encoded video bitstream, a CVS of a bitstream, etc.), such as a decoded representation of a picture in a video. Encoder engine 106 generates a decoded representation of a picture by dividing each picture into multiple slices. Each slice is independent of other slices, allowing information in that slice to be decoded without relying on data from other slices within the same picture. A slice includes one or more segments, which include independent segments and (if present) one or more dependent segments that depend on previous segments.
[0069] In HEVC, slices are divided into decoder tree blocks (CTBs) for luma and chroma samples. One or more CTBs for luma samples and one CTB for chroma samples, along with the syntax used for these samples, are called decoder tree units (CTUs). CTUs can also be called "tree blocks" or "maximum decoder units" (LCUs). A CTU is the basic processing unit used for HEVC encoding. A CTU can be subdivided into multiple decoder units (CUs) of different sizes. A CU contains an array of luma and chroma samples called a decoder block (CB).
[0070] Luminance and chrominance CBs can be further subdivided into prediction blocks (PBs). A PB is a sample block of the luminance or chrominance component that uses the same motion parameters for inter-frame prediction or intra-block copy (IBC) prediction (when available or enabled for use). A luminance PB and one or more chrominance PBs, along with their associated syntax, form a prediction unit (PU). For inter-frame prediction, the set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU and used for inter-frame prediction of the luminance PB and one or more chrominance PBs. Motion parameters can also be referred to as motion information. CBs can also be subdivided into one or more transform blocks (TBs). A TB represents a square block of samples of the chrominance component, to which a residual transform (e.g., in some cases, the same two-dimensional transform) is applied to decode the prediction residual signal. A transform unit (TU) represents a TB of luminance and chrominance samples along with the corresponding syntax elements. Transform decoding is described in more detail below.
[0071] The size of a CU corresponds to the size of the decoding mode and can be square. For example, the size of a CU can be 8 x 8 samples, 16 x 16 samples, 32 x 32 samples, 64 x 64 samples, or any other suitable size up to the corresponding CTU size. The phrase “N x N” is used herein to refer to the pixel size of a video block in both the vertical and horizontal dimensions (e.g., 8 pixels x 8 pixels). Pixels in a block can be arranged in rows and columns. In some embodiments, a block may not have the same number of pixels in the horizontal direction as it does in the vertical direction. The syntax data associated with a CU can describe, for example, the segmentation of the CU into one or more PUs. The segmentation mode can differ between CUs encoded using intra-frame prediction mode and inter-frame prediction mode. PUs can be segmented into non-square shapes. The syntax data associated with a CU can also describe, for example, the segmentation of the CU into one or more TUs according to the CTU. TUs can be square or non-square shapes.
[0072] According to the HEVC standard, transform units (TUs) can be used to perform transforms. The TU can be different for different CUs. The size of the TU can be set based on the size of the PU within a given CU. The TU can have the same size as or smaller than the PU. In some examples, a quadtree structure called a residual quadtree (RQT) can be used to subdivide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can correspond to TUs. The pixel differences associated with the TU can be transformed to produce transform coefficients. The transform coefficients can be quantized by the encoder engine 206.
[0073] Once the video data is segmented into Units (CUs), encoder engine 106 uses a prediction mode to predict each Processing Unit (PU). Prediction units or blocks are subtracted from the original video data to obtain a residual (described below). For each CU, the prediction mode can be signaled within the bitstream using syntax data. Prediction modes can include intra-frame prediction (or intra-picture prediction) or inter-frame prediction (or inter-picture prediction). Intra-frame prediction utilizes the correlation between spatially adjacent samples within a picture. For example, using intra-frame prediction, each PU is predicted from neighboring image data in the same picture using, for example, DC prediction to find an average value for the PU, planar prediction to adapt a planar surface to the PU, orientation prediction to infer from neighboring data, or any other suitable prediction type. Inter-frame prediction uses temporal correlations between pictures to derive motion-compensated predictions for blocks of image samples. For example, using inter-frame prediction, each PU is predicted from image data in one or more reference pictures (in the output order before or after the current picture) using motion-compensated prediction. For example, a decision can be made at the CU level whether to use inter-picture prediction or intra-picture prediction to decode a picture region.
[0074] Encoder engine 106 and decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, the video decoder (such as encoder engine 106 and / or decoder engine 116) segments the image into multiple decoder tree units (CTUs) (where one or more CTBs of luminance samples and chrominance samples, together with the syntax used for these samples, are referred to as CTUs). The video decoder can segment CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple segmentation types (such as the distinction between CU, PU, and TU in HEVC). The QTBT structure includes two levels: a first level segmented according to quadtree segmentation and a second level segmented according to binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to decoder units (CUs).
[0075] In the MTT partitioning structure, blocks can be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. A ternary tree partitioning is a partition where a block is split into three sub-blocks. In some examples, a ternary tree partitioning divides a block into three sub-blocks without partitioning the original block by a center. The partitioning types in MTT (e.g., quadtree, binary tree, and ternary tree) can be symmetric or asymmetric.
[0076] In some examples, the video decoder may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video decoder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luminance component and another QTBT or MTT structure for the two chrominance components (or two QTBT and / or MTT structures for the respective chrominance components).
[0077] Video decoders can be configured to use quadtree segmentation based on HEVC, QTBT segmentation, MTT segmentation, or other segmentation structures. For illustrative purposes, the description herein may refer to QTBT segmentation. However, it should be understood that the techniques of this disclosure can also be applied to video decoders configured to use quadtree segmentation or other types of segmentation.
[0078] In some examples, a slice type is assigned to one or more slices of an image. Slice types include intra-frame decoded slices (I-slices), inter-frame decoded P-slices, and inter-frame decoded B-slices. An I-slice (intra-frame decoded frame, independently decodeable) is a slice of an image that is decoded solely by intra-frame prediction and is therefore independently decodeable because an I-slice only requires intra-frame data to predict any prediction unit or prediction block of the slice. A P-slice (one-way prediction frame) is a slice of an image that can be decoded using both intra-frame prediction and one-way inter-frame prediction. Each prediction unit or prediction block within a P-slice is decoded using either intra-frame prediction or inter-frame prediction. When inter-frame prediction is applied, the prediction unit or prediction block is predicted using only one reference image, and therefore the reference sample comes from only one reference region within a frame. A B-slice (two-way prediction frame) is a slice of an image that can be decoded using both intra-frame prediction and inter-frame prediction (e.g., two-way or one-way prediction). Bidirectional prediction of a B-slice's prediction unit or block can be performed from two reference images, where each image contributes a reference region, and the sample sets of the two reference regions are weighted (e.g., using equal weights or different weights) to generate the prediction signal for the bidirectional prediction block. As explained above, a slice of an image is decoded independently. In some cases, an image may be decoded as a single slice.
[0079] As mentioned above, intra-image prediction leverages the correlation between spatially adjacent samples within an image. Several intra-frame prediction modes exist (also referred to as "intra-frame modes"). In some examples, intra-frame prediction for luma patches includes 35 modes, comprising planar modes, DC modes, and 33 angular modes (e.g., diagonal intra-frame prediction modes and angular modes adjacent to the diagonal intra-frame prediction modes). The 35 intra-frame prediction modes are indexed as shown in Table 1 below. In other examples, more intra-frame modes can be defined, including prediction angles that may not yet be represented by the 33 angular modes. In other examples, the prediction angles associated with angular modes may differ from those used in HEVC. Table 1 - Specification of Intra-Frame Prediction Modes and Associated Names
[0080] Inter-image prediction utilizes the temporal correlation between images to derive motion-compensated predictions for image sample blocks. Using a translational motion model, the position of a block in a previously decoded image (reference image) is determined by a motion vector (…). It means that among them Specifies the horizontal displacement of the reference block relative to the position of the current block, and Specifies the vertical displacement of the reference block relative to the position of the current block. In some cases, the motion vector ( The motion vector can be integer sample precision (also known as integer precision), in which case the motion vector points to an integer pixel grid (or integer pixel sampling grid) of the reference frame. In some cases, the motion vector ( Motion vectors can have fractional sample precision (also known as fractional pixel precision or non-integer precision) to capture the motion of the underlying object more accurately, without being limited to the integer pixel grid of the reference frame. The precision of a motion vector can be expressed by its quantization level. For example, the quantization level can be integer precision (e.g., 1 pixel) or fractional pixel precision (e.g., ¼ pixel, ½ pixel, or other sub-pixel values). When the corresponding motion vector has fractional sample precision, interpolation is applied to the reference image to derive the predicted signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate values at fractional positions. The previously decoded reference image is indicated by a reference index (refIdx) for a list of reference images. The motion vector and reference index can be referred to as motion parameters. Two types of inter-image prediction can be performed, including one-way prediction and two-way prediction.
[0081] When using bidirectional prediction for inter-frame prediction, two sets of motion parameters are used ( and Two motion-compensated predictions (from the same reference image or possibly from different reference images) are generated. For example, in the case of bidirectional prediction, each prediction block uses two motion-compensated prediction signals and generates B prediction units. The two motion-compensated predictions are combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which different weights can be applied to each motion-compensated prediction. The reference images that can be used in bidirectional prediction are stored in two separate lists, denoted as List 0 and List 1, respectively. Motion parameters can be derived at the encoder using a motion estimation process.
[0082] When using unidirectional prediction for inter-frame prediction, a set of motion parameters is used ( This is used to generate motion-compensated predictions from a reference image. For example, in the case of unidirectional prediction, each prediction block uses at most one motion-compensated prediction signal, and P prediction units are generated.
[0083] The prediction unit (PU) may include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when the PU is encoded using intra-frame prediction, the PU may include data describing the intra-frame prediction mode used for the PU. As another example, when the PU is encoded using inter-frame prediction, the PU may include data defining the motion vectors used for the PU. The data defining the motion vectors used for the PU may describe, for example, the horizontal component of the motion vector (…). ), the vertical component of the motion vector ( The resolution used for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference image to which the motion vector points, the reference index, the list of reference images used for the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.
[0084] After performing prediction using intra-frame prediction and / or inter-frame prediction, encoding device 104 can perform transform and quantization. For example, after prediction, encoder engine 106 can compute a residual value corresponding to the PU. The residual value can include the pixel difference between the current pixel block (PU) being decoded and the prediction block used to predict the current block (e.g., a predicted version of the current block). For example, after generating a prediction block (e.g., using inter-frame prediction or intra-frame prediction), encoder engine 106 can generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of pixel differences that quantize the differences between the pixel values of the current block and the pixel values of the prediction block. In some examples, the residual block can be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of the pixel values.
[0085] Block transforms are used to transform any residual data that may remain after prediction is performed. These block transforms can be based on Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), integer transforms, wavelet transforms, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., kernels of size 32 x 32, 16 x 16, 8 x 8, 4 x 4, or other suitable sizes) can be applied to the residual data in each CU. In some examples, TUs can be used for the transform and quantization processes implemented by encoder engine 106. A given CU with one or more PUs can also include one or more TUs. As described further below, residual values can be transformed into transform coefficients using block transforms, and quantization and scanning can be performed using TUs to produce serialized transform coefficients for entropy decoding.
[0086] In some embodiments, after intra-frame prediction or inter-frame prediction decoding using the PU of the CU, the encoder engine 106 can compute residual data for the TU of the CU. The PU may include pixel data in the spatial domain (or pixel domain). As previously described, the residual data may correspond to the pixel difference between a pixel in the uncoded image and the predicted value corresponding to the PU. The encoder engine 106 may form one or more TUs including residual data for the CU (which includes the PU), and may transform the TUs to produce transform coefficients for the CU. The TUs may include coefficients in the transform domain after applying the block transform.
[0087] The encoder engine 106 can perform quantization on the transform coefficients. Quantization provides further compression by reducing the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient with an n-bit value can be rounded down to an m-bit value during quantization, where n is greater than m.
[0088] Once quantization is performed, the decoded video bitstream includes quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), segmentation information, and any other suitable data (such as other syntax data). The different elements of the decoded video bitstream can be entropy-coded by encoder engine 106. In some examples, encoder engine 106 can scan the quantized transform coefficients using a predefined scan order to produce a serialized vector that can be entropy-coded. In some examples, encoder engine 106 can perform adaptive scanning. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), encoder engine 106 can entropy-code that vector. For example, encoder engine 106 can use context-adaptive variable-length decoding, context-adaptive binary arithmetic decoding, syntax-based context-adaptive binary arithmetic decoding, probabilistic interval segmentation entropy decoding, or another suitable entropy coding technique.
[0089] The output 110 of encoding device 104 can transmit NAL units constituting the encoded video bitstream data to decoding device 112 of the receiving device on communication link 120. The input 114 of decoding device 112 can receive the NAL units. Communication link 120 may include a channel provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network may include any wireless interface or combination of wireless interfaces and may include any suitable wireless network (e.g., the Internet or other wide area networks, packet-based networks, WiFi). TM Radio frequency (RF), UWB, WiFi Direct, Cellular, Long Term Evolution (LTE), WiMax TM Wired networks can include any wired interface (e.g., fiber optic, Ethernet, powerline Ethernet, coaxial cable Ethernet, digital signal line (DSL), etc.). Various devices can be used to implement wired and / or wireless networks, such as base stations, routers, access points, bridges, gateways, switches, etc. Encoded video bitstream data can be modulated according to communication standards such as wireless communication protocols and transmitted to receiving devices.
[0090] In some examples, encoding device 104 may store encoded video bitstream data in storage device 108. Output 110 may retrieve the encoded video bitstream data from encoder engine 106 or from storage device 108. Storage device 108 may include any of a variety of distributed or locally accessed data storage media. For example, storage device 108 may include hard disk drives, storage disks, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. Storage device 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. In further examples, storage device 108 may correspond to a file server or another intermediate storage device that may store encoded video generated by a source device. In such a case, a receiving device including decoding device 112 may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing encoded video data and sending such encoded video data to the receiving device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device can access the encoded video data via any standard data connection, including an internet connection. Access may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on a file server. Transmission of the encoded video data from storage device 108 may be streaming, downloading, or a combination thereof.
[0091] Input 114 of decoding device 112 receives encoded video bitstream data and can provide the video bitstream data to decoder engine 116 or to storage device 118 for later use by decoder engine 116. For example, storage device 118 may include a DPB for storing reference pictures used in inter-frame prediction. A receiving device including decoding device 112 can receive the encoded video data to be decoded via storage device 108. The encoded video data can be modulated and transmitted to the receiving device according to a communication standard such as a wireless communication protocol. The communication medium used to transmit the encoded video data can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, switch, base station, or any other means that can be used to facilitate communication from the source device to the receiving device.
[0092] Decoder engine 116 decodes encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more decoded video sequences that constitute the encoded video data. Decoder engine 116 can rescale the encoded video bitstream data and perform an inverse transform on it. Residual data is passed to the prediction stage of decoder engine 116. Decoder engine 116 predicts pixel blocks (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (residual data).
[0093] Video decoding device 112 can output decoded video to video destination device 122, which may include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, video destination device 122 may be part of a receiving device that includes decoding device 112. In some aspects, video destination device 122 may be part of a separate device, distinct from the receiving device.
[0094] In some embodiments, the video encoding device 104 and / or the video decoding device 112 may be integrated with the audio encoding device and the audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 may also include other hardware or software necessary for implementing the decoding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. The video encoding device 104 and the video decoding device 112 may be integrated as part of a combined encoder / decoder (codec) in the respective device.
[0095] exist Figure 1 The example system shown is an illustrative example that can be used herein. The techniques used to process video data using the techniques described herein can be implemented by any digital video encoding and / or decoding device. Although, in general, the techniques of this disclosure are implemented by video encoding or video decoding devices, the techniques can also be implemented by a combination of video encoder-decoder, commonly referred to as a “CODEC.” Furthermore, the techniques of this disclosure can also be implemented by a video preprocessor. The source device and receiving device are merely examples of such decoding devices, wherein the source device generates decoded video data for transmission to the receiving device. In some examples, the source device and receiving device can operate in a substantially symmetrical manner, such that each of these devices includes video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0096] Extensions to the HEVC standard include the Multi-View Video Decoding Extension (MV-HEVC) and the Scalable Video Decoding Extension (SHVC). MV-HEVC and SHVC extensions share the concept of layered decoding, where different layers are included in the encoded video bitstream. Each layer in the decoded video sequence is addressed by a unique layer identifier (ID). The layer ID can be present in the header of a NAL unit to identify the layer associated with that NAL unit. In MV-HEVC, different layers typically represent different views of the same scene in the video bitstream. In SHVC, different scalable layers are provided to represent the video bitstream with different spatial resolutions (or picture resolutions) or different reconstruction fidelities. A scalable layer can include a base layer (where layer ID = 0) and one or more enhancement layers (where layer IDs = 1, 2, … n). The base layer may conform to the first version of the HEVC profile and represents the lowest available layer in the bitstream. Compared to the base layer, enhancement layers have increased spatial resolution, temporal resolution, or frame rate and / or reconstruction fidelity (or quality). Enhancement layers are organized hierarchically and may depend on (or not depend on) lower layers. In some examples, a single-standard codec can be used to decode different layers (e.g., using HEVC, SHVC, or other decoding standards to encode all layers). In other examples, multi-standard codecs can be used to decode different layers. For example, AVC can be used to decode the base layer, while SHVC and / or MV-HEVC extensions of the HEVC standard can be used to decode one or more enhancement layers.
[0097] As described above, for each block, a set of motion information (also referred to herein as motion parameters) may be available. This set of motion information may contain motion information for both the forward and backward prediction directions. Here, the forward and backward prediction directions are two prediction directions in a bidirectional prediction mode, and the terms "forward" and "backward" do not necessarily have geometric meaning. Instead, forward and backward may correspond to a reference image list 0 (RefPicList0) and a reference image list 1 (RefPicList1) for the current image, slice, or block. In some examples, when only one reference image list is available for an image, slice, or block, only RefPicList0 is available, and the motion information for each block of the slice is always forward. In some examples, RefPicList0 includes reference images that are temporally preceding the current image, and RefPicList1 includes reference images that are temporally following the current image. In some cases, motion vectors along with associated reference indices may be used during decoding. Such motion vectors with associated reference indices are represented as a unidirectional prediction set of motion information.
[0098] For each prediction direction, motion information may include a reference index and a motion vector. In some cases, for simplicity, the motion vector may have associated information, based on which it can be assumed that the motion vector has an associated reference index. The reference index can be used to identify reference images in the current list of reference images (RefPicList0 or RefPicList1). The motion vector may have horizontal and vertical components, providing the offset from the coordinate position in the current image to the coordinate position in the reference image identified by the reference index. For example, the reference index may indicate a specific reference image that should be used for a block in the current image, and the motion vector may indicate where the best-matching block in the reference image (the block that best matches the current block) is located in the reference image.
[0099] Picture order counts (POCs) can be used in video decoding standards to identify the display order of pictures. Although it is possible for two pictures within a decoded video sequence to have the same POC value, it is generally not the case that two pictures with the same POC value will have the same POC value within a decoded video sequence. When multiple decoded video sequences exist in a bitstream, pictures with the same POC value are likely to be closer to each other in terms of decoding order. The POC value of a picture can be used for constructing a reference picture list (such as deriving a reference picture set in HEVC) and / or motion vector scaling, etc.
[0100] In H.264 / AVC, each inter-frame macroblock (MB) can be partitioned in four different ways, including: one 16x16 macroblock partition; two 16x8 macroblock partitions; two 8x16 macroblock partitions; and four 8x8 macroblock partitions; and so on. Different macroblock partitions within a macroblock can have different reference index values for each prediction direction (different reference index values for RefPicList0 and RefPicList1).
[0101] In some cases, when a macroblock is not divided into four 8x8 macroblock partitions, it may have only one motion vector per macroblock partition in each prediction direction. In other cases, when a macroblock is divided into four 8x8 macroblock partitions, each 8x8 macroblock partition can be further subdivided into sub-blocks, where each sub-block can have a different motion vector in each prediction direction. An 8x8 macroblock partition can be divided into sub-blocks in different ways, including: one 8x8 sub-block; two 8x4 sub-blocks; two 4x8 sub-blocks; and four 4x4 sub-blocks; and so on. Each sub-block can have a different motion vector in each prediction direction. Therefore, motion vectors can exist at a level equal to or higher than that of the sub-block.
[0102] In HEVC, the largest decoding unit in a slice is called a decoding tree block (CTB) or decoding tree unit (CTU). A CTB contains a quadtree, where the nodes are decoding units. In the HEVC master profile, the size of a CTB can range from 16x16 pixels to 64x64 pixels. In some cases, an 8x8 pixel CTB size is supported. A CTB can be recursively divided into decoding units (CUs) in a quadtree manner. The size of a CU can be the same as the CTB and can be as small as 8x8 pixels. In some cases, each decoding unit uses a mode for decoding, such as intra-prediction mode or inter-prediction mode. When using inter-prediction mode for inter-decoding of a CU, the CU can be further divided into two or four prediction units (PUs), or, when further division is not applicable, the CU can be considered as a single PU. When two PUs exist within a CU, these two PUs can be rectangles of half the size, or two rectangles of 1 / 4 or 3 / 4 the size of the CU.
[0103] When performing inter-frame decoding on a CU, a set of motion information can exist for each PU, which can be derived using a unique inter-frame prediction mode. For example, an inter-frame prediction mode can be used to decode each PU to derive the set of motion information. In some cases, when using intra-frame prediction modes to intra-decode the CU, the PU shape can be 2Nx2N or NxN. Within each PU, a single intra-frame prediction mode is decoded (while the chroma prediction mode is signaled at the CU level). In some cases, an NxN intra-frame PU shape is permitted when the current CU size is equal to the minimum CU size defined in the SPS.
[0104] For motion prediction in HEVC, there are two inter-frame prediction modes for prediction units (PUs): merge mode and Advanced Motion Vector Prediction (AMVP) mode. The special case considered as merge is skipped. In AMVP or merge mode, a candidate list of motion vectors (MVs) for multiple motion vector predictors can be maintained. In merge mode, the motion vector of the current PU and its reference index are generated by extracting a candidate from the MV candidate list.
[0105] In some examples, the MV candidate list contains up to five candidates for the merge mode and two candidates for the AMVP mode. In other examples, different numbers of candidates can be included in the MV candidate list for the merge mode and / or AMVP mode. The merge candidate can contain a set of motion information. For example, the set of motion information can include motion vectors corresponding to two lists of reference images (list 0 and list 1) and reference indices. If the merge candidate is identified by a merge index, the reference image is used for the prediction of the current block and to determine the associated motion vector. However, in AMVP mode, for each potential prediction direction from list 0 or list 1, the reference index needs to be explicitly signaled along with the MV predictor (MVP) index of the MV candidate list, because AMVP candidates only contain motion vectors. In AMVP mode, the predicted motion vectors can be further refined.
[0106] Merged candidates can correspond to a complete set of motion information, while AMVP candidates can contain a motion vector and a reference index for a specific prediction direction. Candidates for both modes can be derived similarly from the same spatially and temporally adjacent blocks.
[0107] In some examples, the merge mode allows an inter-frame predicted PU to inherit one or more motion vectors, prediction directions, and one or more reference image indices from an inter-frame predicted PU that includes motion data locations selected from a set of spatially adjacent motion data locations and one of two temporally co-located motion data locations. In AMVP mode, one or more motion vectors of the PU can be predicted and decoded relative to one or more motion vector predictors (MVPs) from an AMVP candidate list constructed by the encoder. In some cases, for unidirectional inter-frame prediction of the PU, the encoder can generate a single AMVP candidate list. In some cases, for bidirectional prediction of the PU, the encoder can generate two AMVP candidate lists, one using motion data from spatially and temporally adjacent PUs in the forward prediction direction, and the other using motion data from spatially and temporally adjacent PUs in the backward prediction direction.
[0108] Candidates for both modes can be derived from spatially and / or temporally adjacent blocks. For example, Figure 2A and Figure 2B Includes a conceptual diagram illustrating spatial adjacency candidates in HEVC. Figure 2A Candidate spatially adjacent motion vectors (MVs) for merging patterns are shown. Figure 2BSpatial neighbor motion vector (MV) candidates for AMVP mode are shown. Spatial MV candidates are derived from neighboring blocks for a specific PU (PU0), but the method of generating candidates based on blocks differs for merge and AMVP modes.
[0109] In merging mode, the encoder can form a list of merging candidates by considering merging candidates from various motion data locations. For example, ... Figure 2A As shown, regarding in Figure 2A Using the spatially adjacent motion data positions indicated by numbers 0-4, up to four spatial MV candidates can be derived. The MV candidates can be sorted in the merged candidate list in the order shown by numbers 0-4. For example, the positions and order can include: left position (0), top position (1), upper right position (2), lower left position (3), and upper left position (4).
[0110] exist Figure 2B In the AVMP mode shown, adjacent blocks are divided into two groups: the left group comprising blocks 0 and 1, and the upper group comprising blocks 2, 3, and 4. For each group, potential candidates in adjacent blocks that reference the same reference image as indicated by the signaled reference index have the highest priority to be selected to form the final candidates for that group. It is possible that none of the adjacent blocks contain motion vectors pointing to the same reference image. Therefore, if no such candidate can be found, the first available candidate is scaled to form the final candidate; thus, temporal distance differences can be compensated for.
[0111] Figure 3A and Figure 3B This includes a conceptual diagram illustrating temporal motion vector prediction in HEVC. Temporal Motion Vector Predictor (TMVP) candidates (if enabled and available) are added to the MV candidate list after spatial motion vector candidates. The process for deriving motion vectors for TMVP candidates is the same as in both merge mode and AMVP mode. However, in some cases, the target reference index used for TMVP candidates is always set to zero in merge mode.
[0112] The main block position used for TMVP candidate derivation is the lower right block outside the corresponding PU (e.g., in...). Figure 3A The block is designated as "T" to compensate for the offset of the upper and left blocks used to generate spatially adjacent candidates. However, if the block is outside the current CTB (or LCU) row or motion information is not available, the center block of the PU is used to replace the block. Motion vectors for TMVP candidates are derived from the co-located PUs of the co-located images indicated at the slice level. Similar to the temporal direct mode in AVC, the motion vectors of TMVP candidates can undergo motion vector scaling, which is performed to compensate for distance differences.
[0113] Other aspects of motion prediction are also covered in HEVC, VVC, and other video decoding specifications. One aspect, for example, includes motion vector scaling. In motion vector scaling, it is assumed that the value of a motion vector is proportional to the distance between images according to their presentation time. In some examples, a first motion vector may be associated with two images: a first reference image and a first containing image that includes the first motion vector. The first motion vector can be used to predict a second motion vector. To predict the second motion vector, a first distance between the first containing image and the first reference image can be calculated based on the Picture Order Count (POC) values associated with the first reference image and the first containing image.
[0114] A second reference image and a second contained image can be associated with the second motion vector to be predicted, wherein the second reference image may be different from the first reference image, and the second contained image may be different from the first contained image. A second distance between the second reference image and the second contained image can be calculated based on the POC values associated with the second reference image and the second contained image, wherein the second distance may be different from the first distance. To predict the second motion vector, the first motion vector can be scaled based on the first distance and the second distance. For spatially adjacent candidates, the first contained image and the second contained image of the first motion vector and the second motion vector can be the same, while the first reference image and the second reference image can be different. In some examples, motion vector scaling can be applied to TMVP and AMVP modes for spatially and temporally adjacent candidates.
[0115] Another aspect of motion prediction involves the generation of artificial motion vector candidates. For example, if the list of motion vector candidates is incomplete, artificial motion vector candidates are generated and inserted at the end of the list until all candidates are obtained. In the merging mode, there are two types of artificial motion vector candidates: a first type, which includes combined candidates derived only for B-slices; and a second type, which includes zero candidates for AMVP only (if the first type does not provide enough artificial candidates). For each pair of candidates that is already in the motion vector candidate list and has relevant motion information, a bidirectional combined motion vector candidate can be derived by combining the motion vector of the first candidate referencing the image in list 0 and the motion vector of the second candidate referencing the image in list 1.
[0116] Another aspect of the merge and AMVP patterns includes a pruning process for candidate insertion. For example, candidates from different blocks might be identical, which reduces the efficiency of merging and / or AMVP candidate lists. A pruning process can be applied to address this problem. The pruning process involves comparing candidates with those already existing in the current candidate list to avoid inserting identical or duplicate candidates. To reduce the complexity of comparisons, the pruning process can be performed on fewer potential candidates than are to be inserted into the candidate list.
[0117] In some examples, enhanced motion vector prediction can be achieved. For instance, video decoding standards such as VVC specify inter-frame decoding tools that, based on VCC, can derive or refine a candidate list for motion vector prediction or merge prediction for the current block. Examples of such methods are described below.
[0118] History-based motion vector prediction (HMVP) is a motion vector prediction method that allows each block to find its MV predictor from a list of previously decoded motion vectors (MVs) excluding those in the immediate causal neighboring motion fields. For example, in addition to MVs in the immediate causal neighboring motion fields, HMVP can also obtain or predict one or more MV predictors for the current block from a list of previously decoded MVs. The MV predictors in the list of previously decoded MVs are called HMVP candidates. HMVP candidates can include motion information associated with the inter-frame decoded block. An HMVP table with multiple HMVP candidates can be maintained during the encoding and / or decoding process for a slice. In some examples, the HMVP table can be updated dynamically. For example, after decoding an inter-frame decoded block, the HMVP table can be updated by adding the associated motion information of the decoded inter-frame decoded block as a new HMVP candidate. In some examples, the HMVP table can be cleared when a new slice is encountered.
[0119] In some cases, whenever an inter-frame decoding block exists, the associated motion information can be inserted into the table as a new HMVP candidate in a first-in-first-out (FIFO) manner. Constrained FIFO rules can be applied. When inserting an HMVP into the table, a redundancy check can first be applied to find if the same HMVP already exists in the table. If found, that specific HMVP can be removed from the table, and then all candidate HMVPs can be moved.
[0120] In some examples, HMVP candidates can be used during the merge candidate list construction process. In some cases, all HMVP candidates from the last entry to the first entry in the table are inserted after the TMVP candidates. Pruning can be applied to HMVP candidates. The merge candidate list construction process can be terminated once the total number of available merge candidates reaches the maximum allowed merge candidates signaled.
[0121] In some examples, HMVP candidates can be used during the AMVP candidate list construction process. In some cases, the last candidate in the table will be... K The motion vectors of HMVP candidates are inserted after the TMVP candidates. In some implementations, only HMVP candidates with the same reference image as the AMVP target reference image are used to construct the AMVP candidate list. Pruning can be applied to HMVP candidates.
[0122] Figure 4 This is a block diagram illustrating an example of an HMVP table 400. The HMVP table 400 can be implemented as a storage device managed using a first-in, first-out (FIFO) rule. For example, HMVP candidates, including MV predictors, can be stored in the HMVP table 400. The HMVP candidates can be stored in the order they are encoded or decoded. In one example, the order in which HMVP candidates are stored in the HMVP table 400 can correspond to the time in which the HMVP candidates are constructed. For example, when implemented in a decoder such as decoding device 112, HMVP candidates can be constructed to include motion information of decoded inter-frame decoded blocks. In some examples, one or more HMVP candidates from the HMVP table 400 can include motion vector predictors, which can be used to predict motion vectors for the current block to be decoded. In some examples, one or more HMVP candidates can include one or more previously decoded blocks, which can be stored in one or more entries of the HMVP table 400 in a FIFO manner according to the time order in which they are decoded.
[0123] HMVP candidate index 402 is shown as being associated with HMVP table 400. HMVP candidate index 402 may identify one or more entries in HMVP table 400. According to the illustrative example, HMVP candidate index 402 is shown as including index values 0 to 4, where each index value of HMVP candidate index 402 is associated with a corresponding entry. HMVP table 400 may include entries referenced in other examples. Figure 4The entries shown and described may have more or fewer entries. When constructing HMVP candidates, they are filled into HMVP table 400 in a FIFO manner. For example, as HMVP candidates are decoded, they are inserted into HMVP table 400 at one end, and entries are moved sequentially through HMVP table 400 until they exit HMVP table 400 from the other end. Therefore, in some examples, memory structures such as shift registers can be used to implement HMVP table 400. In one example, index 0 can point to the first entry of HMVP table 400, where the first entry may correspond to the first end of HMVP table 400 from which HMVP candidates are inserted. Correspondingly, index 4 can point to the second entry of HMVP table 400, where the second entry may correspond to the second end of HMVP table 400 from which HMVP candidates exit or are cleared from HMVP table 400. Therefore, the HMVP candidate inserted at the first entry at index 0 can traverse HMVP table 400 to make room for newer or most recently decoded HMVP candidates until the HMVP candidate reaches the second entry at index 4. Therefore, among the HVMP candidates appearing in HMVP table 400 at any given time, the HMVP candidate in the second entry at index 4 can be the oldest or most distant, while the HMVP candidate in the first entry at index 0 can be the most recent or most recent. Typically, the HMVP candidate in the second entry can be an older or more recently constructed HMVP candidate than the HMVP candidate in the first entry.
[0124] exist Figure 4 In the diagram, reference numerals 400A, 400B, and 400C are used to identify different states of HMVP table 400. Referring to state 400A, it is shown that HMVP candidates HMVP0 to HMVP4 exist in entries of HMVP table 400 with their respective index values 4 to 0. For example, HMVP0 could be the oldest or most distant HMVP candidate inserted into HMVP table 400 at the first entry at index value 0. HMVP0 can be shifted sequentially to make room for the most recently inserted and newer HMVP candidates HMVP1 to HMVP4 until HMVP0 reaches the second entry at index value 4 as shown in state 400A. Correspondingly, HMVP4 could be the most recent HMVP candidate inserted into the first entry at index value 0. Therefore, HMVP0 is an older or more distant HMVP candidate in HMVP table 400 relative to HMVP4.
[0125] In some examples, one or more of the HMVP candidates HMVP0 through HMVP4 may include potentially redundant motion vector information. For example, a redundant HMVP candidate may include the same motion vector information as that in one or more other HMVP candidates stored in HMVP table 400. Since the motion vector information of the redundant HMVP candidate can be obtained from one or more other HMVP candidates, storing the redundant HMVP candidate in HMVP table 400 can be avoided. By avoiding storing redundant HMVP candidates in HMVP table 400, the resources of HMVP table 400 can be utilized more efficiently. In some examples, a redundancy check can be performed before storing an HMVP candidate in HMVP table 400 to determine whether the HMVP candidate would be redundant (e.g., the motion vector information of the HMVP candidate can be compared with the motion vector information of other already stored HMVP candidates to determine if a match exists).
[0126] In some examples, the state of reference numeral 400B in HMVP table 400 is a conceptual illustration of the redundancy check described above. In some examples, as HMVP candidates are decoded, they can be populated in HMVP table 400, and redundancy checks can be performed periodically instead of as a threshold test before storing HMVP candidates. For example, as shown in the state of reference numeral 400B, HMVP candidates HMVP1 and HMVP3 can be identified as redundant candidates (i.e., their motion information is the same as the motion information of one of the other HMVP candidates in HMVP table 400). Redundant HMVP candidates HMVP1 and HMVP3 can be removed, and the remaining HMVP candidates can be shifted accordingly.
[0127] For example, as shown by reference numeral 400C in the attached figure, HMVP candidates HMVP2 and HMVP4 are shifted toward higher index values corresponding to older entries, while HMVP0, already in the second entry at the end of HMVP table 400, is not shown as being shifted further. In some examples, shifting HMVP candidates HMVP2 and HMVP4 can free up space in HMVP table 400 for newer HMVP candidates. Thus, new HMVP candidates HMVP5 and HMVP6 are shown as being shifted into HMVP table 400, where HMVP6 is the most recent or includes recently decoded motion vector information and is stored in the first entry at index 0.
[0128] In some examples, one or more HMVP candidates from HMVP table 400 can be used to construct an additional candidate list that can be used for motion prediction of the current block. For example, one or more HMVP candidates from HMVP table 400 can be added to the merged candidate list, for example, as additional merged candidates. In some examples, one or more HMVP candidates from the same HMVP table 400 or another such HMVP table can be added to the Advanced Motion Vector Prediction (AMVP) candidate list, for example, as additional AMVP predictors.
[0129] For example, during the construction of the merged candidate list, some or all of the HMVP candidates stored in the entries of HMVP table 400 can be inserted into the merged candidate list. In some examples, inserting HMVP candidates into the merged candidate list may include inserting HMVP candidates after the Time Motion Vector Predictor (TMVP) candidates in the merged candidate list. (See previous reference...) Figure 3A and Figure 3B As discussed, if a TMVP candidate is enabled and available, it can be added to the MV candidate list after the spatial motion vector candidate.
[0130] In some examples, the pruning process described above can be applied to HMVP candidates when constructing the merge candidate list. For instance, once the total number of merge candidates in the merge candidate list reaches the maximum allowed number of merge candidates, the merge candidate list construction process can be terminated, and no more HMVP candidates will be inserted into the merge candidate list. The maximum allowed number of merge candidates in the merge candidate list can be a predetermined number, or it can be, for example, the number notified by the encoder to the decoder via a signal, at which the merge candidate list can be constructed.
[0131] In some examples of constructing a merge candidate list, one or more other candidates can be inserted into the merge candidate list. In some examples, motion information from previously decoded blocks that are not adjacent to the current block can be used for more efficient motion vector prediction. For example, non-adjacent spatial merge candidates can be used when constructing the merge candidate list. In some cases, the construction of non-adjacent spatial merge candidates (e.g., as described in JVET-K0228, which is incorporated herein by reference in its entirety and used for all purposes) involves deriving a new spatial candidate from two non-adjacent adjacent positions (e.g., from the nearest non-adjacent block to the left / above, such as in...). Figure 5(As shown in the diagram and discussed below). These blocks can be limited to a maximum distance of 1 CTU from the current block. The retrieval process for non-adjacent candidates begins by tracing previously decoded blocks in the vertical direction. Vertical reverse tracing stops when an intra-frame block is encountered or the tracing distance reaches 1 CTU size. The retrieval process traces previously decoded blocks in the horizontal direction. The criterion for stopping the horizontal retrieval process depends on whether a successfully retrieved vertical non-adjacent candidate exists. If no vertical non-adjacent candidate is retrieved, the horizontal retrieval process stops when an intra-frame block is encountered or the tracing distance exceeds a threshold of 1 CTU size. If a retrieved vertical non-adjacent candidate exists, the horizontal retrieval process stops when an inter-frame block containing a different MV than the vertical non-adjacent candidate is encountered or the tracing distance exceeds a threshold of 1 CTU size. In some examples, non-adjacent spatial merge candidates can be inserted before TMVP candidates in the merge candidate list. In some examples, non-adjacent spatial merge candidates can be inserted before TMVP candidates in the same merge candidate list, which may include one or more HMVP candidates inserted after the TMVP candidates. Refer to the following. Figure 5 The description identifies and retrieves one or more non-adjacent space merge candidates that can be inserted into the merge candidate list.
[0132] Figure 5 This is a block diagram showing an image or slice 500 including the current block 502 to be decoded. In some examples, a merge candidate list can be constructed for decoding the current block 502. For example, a motion vector for the current block can be obtained from one or more merge candidates in the merge candidate list. The merge candidate list may include determining non-adjacent spatial merge candidates. For example, non-adjacent spatial merge candidates may include new spatial candidates derived from two non-adjacent adjacent positions relative to the current block 502.
[0133] The diagram shows several adjacent or neighboring blocks of the current block 502, including the top-left block B2510 (above-left of the current block 502), the top block B1512 (above the current block 502), the top-right block B0514 (above-right of the current block 502), the left-side block A1516 (to the left of the current block 502), and the bottom-left block A0518 (below-left of the current block 502). In some examples, non-adjacent space merge candidates can be obtained from a nearest non-adjacent block above and / or to the left of the current block.
[0134] In some examples, non-adjacent space merge candidates for the current block 502 may include tracing previously decoded blocks in the vertical direction (above the current block 502) and / or in the horizontal direction (to the left of the current block 502). The vertical tracing distance 504 indicates the distance relative to the current block 502 (e.g., the upper boundary of the current block 502) and the vertical non-adjacent block V. NThe vertical distance is 520. The horizontal tracing distance 506 indicates the distance relative to the current block 502 (e.g., the left boundary of the current block 502) and the horizontal non-adjacent block H. N The horizontal distance is 522. The vertical tracing distance 504 and the horizontal tracing distance 506 are limited to a maximum distance equal to the size of a decoding tree unit (CTU).
[0135] Non-adjacent spatial merge candidates, such as vertical non-adjacent blocks V, can be identified by tracing previously decoded blocks in both the vertical and horizontal directions. N 520 and horizontal non-adjacent block H N 522. For example, retrieve the vertical non-adjacent block V. N 520 may include a vertical reverse tracing process to determine whether an inter-frame decoded block exists within a vertical tracing distance 504 (constrained to the maximum size of a CTU). If such a block exists, it is identified as a vertically non-adjacent block V. N 520. In some examples, a horizontal reverse tracing process can be performed after the vertical reverse tracing process. The horizontal reverse tracing process may include determining whether an inter-frame decoded block exists within a horizontal tracing distance of 506 (constrained to the maximum size of a CTU), and if such a block is found, identifying it as a horizontally non-adjacent block H. N 522.
[0136] In some examples, the vertical non-adjacent block V can be retrieved. N 520 and horizontal non-adjacent block H N One or more of 522 can be used as candidates for non-adjacent space merging. The retrieval process may include: if a vertical non-adjacent block V is identified during the vertical reverse tracing process. N 520, then retrieve the vertical non-adjacent block V. N 520. The retrieval process can continue with the horizontal reverse tracing process. If the vertical non-adjacent block V is not identified during the vertical reverse tracing process... N If 520 is encountered, the horizontal reverse tracing process can be terminated when an inter-frame decoded block is encountered or the horizontal tracing distance exceeds the maximum distance (506). If a vertical non-adjacent block V is identified and retrieved... N 520, then when encountering a block V that is not adjacent to the vertical block... N If the horizontal tracing distance 506 exceeds the maximum distance, the horizontal reverse tracing process is terminated when the MV included in 520 is different from the inter-frame decoded block. As mentioned earlier, the retrieved non-adjacent space candidates (such as vertical non-adjacent blocks V) can be merged. N 520 and horizontal non-adjacent block H N One or more of the candidates in 522) are added to the merged candidate list before the TMVP candidate.
[0137] Return to reference Figure 4 In some cases, HMVP candidates can also be used to construct the AMVP candidate list. During the AMVP candidate list construction process, some or all HMVP candidates from entries stored in the same HMVP table 400 (or a different HMVP table than the one used to construct the merged candidate list) can be inserted into the AMVP candidate list. In some examples, inserting HMVP candidates into the AMVP candidate list may include adding a set of HMVP candidate entries (e.g., ...). k The most recent or oldest entry is inserted into the AMVP candidate list after the TMVP candidate. In some examples, the above pruning process can be applied to HMVP candidates when constructing the AMVP candidate list. In some examples, only those HMVP candidates with the same reference image as the AMVP target reference image can be used to construct the AMVP candidate list.
[0138] Therefore, history-based motion vector predictor (HMVP) prediction modes can involve using history-based lookup tables, such as HMVP table 400 which includes one or more HMVP candidates. HMVP candidates can be used in inter-frame prediction modes such as merge mode and AMVP mode. In some examples, different inter-frame prediction modes may use different methods to select HMVP candidates from HMVP table 400.
[0139] In some cases, alternative motion vector prediction designs can be used. For example, alternative designs for spatial MVP (S-MVP) prediction and temporal MVP (T-MVP) prediction can be utilized. For instance, in some implementations of the merge mode (in some cases, the merge mode may be referred to as the skip mode or the direct mode), in... Figure 6A , Figure 6B and Figure 6C The spatial and temporal MVP candidates shown in these figures can be accessed (or searched or selected) in the given order shown in these figures to populate the MVP list.
[0140] Figure 6A The positions of MVP candidates A, B, (C, A1 | B1), A0, and B2 for the current block 600 are shown. Figure 6B The diagram shows the temporally co-located neighbor at center position 610, which has a backoff candidate H for the current block 600. The spatial and temporal locations used in MVP prediction are as follows... Figure 6A As shown. In Figure 6C The example of an access order (e.g., search order or selection order) for an S-MVP is shown using search order blocks 0, 1, 2, 3, 4, and 5. Figure 6D The diagram shows the spatial reverse pattern used for searching sequence blocks 0-5 (and...). Figure 6C(The order in the text is replaced by the order of the characters.)
[0141] The spatial neighbors used as MVP candidates are A, B, (C, A1 | B1), A0, and B2. This is achieved using a two-phase process, where... Figure 6C The access order is marked in the middle: Group 1: a. A, B, C (in HEVC notation, they are in the same position as B0) b.A1 or B1, depending on the availability of the MVP at position C and the type of block splitting. Group 2: a. A0 and B2
[0142] The temporally adjacent neighbors used as MVP candidates are the block at the center position (610) of the current block and the block at the bottom right outside the current block: Group 3: a. C, H b. If the H position is found to be outside the corresponding image, the backtracking H position can be used as an alternative.
[0143] In some implementations, depending on the block splitting and decoding order used, a reverse S-MVP candidate order can be used, such as... Figure 6D As shown.
[0144] In HEVC and earlier video decoding standards, only translational motion models were applied to motion-compensated prediction (MCP). For example, translational motion vectors could be determined for each block of an image (e.g., each CU or each PU). However, in the real world, many more types of motion exist besides translational motion, including scaling (e.g., zooming in and / or zooming out), rotation, perspective motion, and other irregular motions. In the Joint Exploration Model (JEM) of ITU-T VCEG and MPEG, affine transformation motion-compensated prediction can be applied to improve decoding efficiency by using affine decoding modes.
[0145] Figure 7 This is a diagram showing the affine motion field of the current block 702 described by two motion vectors, which are shown as vector 720 of two corresponding control points 510 and 512. ) and vector 722 ( Motion vector 720 using control point 710. and the motion vector 722 of control point 712 The motion vector field (MVF) of the current block 702 can be described by the following equation: Equation (1)
[0146] In equation (1), and Form a motion vector for each pixel within the current block 702. x and y Provide the position of each pixel within the current block 702 (for example, the top-left pixel in the block can have coordinates or an index). x , y ) = (0,0)), and ( , ) is the motion vector of the upper left control point 710. w It is the width of the current block 702, and ( , ) is the motion vector 722 of the upper right control point 712. and The value is the horizontal value of the corresponding motion vector, and and The value is the vertical value of the corresponding motion vector. Additional control points (e.g., four control points, six control points, eight control points, or some other number of control points) can be defined by adding additional control point vectors, for example, at the lower corner of the current block 702, the center of the current block 702, or other locations within the current block 702.
[0147] Equation (1) above shows a 4-parameter motion model, where the four affine parameters a, b, c, and d are defined as: ; ; ;as well as Using equation (1), the motion vector of the upper left control point 710 is given ( , ) and the motion vector of the upper right control point 712 ( , ), can use the coordinates of each pixel position ( x , y () is used to calculate the motion vector for each pixel in the current block. For example, for the top-left pixel position of the current block 702, () x , y The value of ) can be equal to (0, 0), in which case the motion vector for the top left pixel becomes = and = To further simplify MCP, block-based affine transformation prediction can be applied.
[0148] Figure 8This is a graph showing the block-based affine transformation prediction of the current block 802 (e.g., which may resemble the current block 600 or the current block 702) that has been divided into sub-blocks (including the shown sub-blocks 804, 806, and 808). Figure 8 The example shown includes a 4 x 4 partition with a total of 16 sub-blocks. In other examples, any suitable partition and corresponding number of sub-blocks can be used. The motion vector can be derived for each sub-block using equation (1). In some examples, the motion vector for each 4 x 4 sub-block is derived by calculating the motion vector of the center sample of each sub-block according to equation (1) (e.g., Figure 8 As shown, motion vector 805 derived for sub-block 804, motion vector 807 derived for sub-block 806, and motion vector 809 derived for sub-block 808, each motion vector is derived from the center sample of the corresponding sub-block. In other examples, other samples may be used. In some examples, each obtained motion vector may be rounded to, for example, 1 / 16 fractional precision or other suitable precision (e.g., 1 / 4, 1 / 8, etc.). Motion compensation can be applied using the derived motion vectors of the sub-blocks to generate predictions for each sub-block. For example, the decoding device may receive motion vectors describing control point 810. Motion vectors of 820 and control point 812 The four affine parameters (a, b, c, d) of the 822 can be used to calculate the motion vector for each sub-block based on the pixel coordinate index describing the position of the center sample of each sub-block. After MCP, the high-precision motion vector of each sub-block can be rounded as described above and can be preserved with the same precision as the translation motion vector. Furthermore, in some examples, the motion vector in the affine mode can be restricted to limit the reference data to be used during the affine decoding operation, which uses the motion vector as the affine motion vector in the affine decoding mode. In some such examples, clipping can be applied to such vectors, as described in more detail below, particularly regarding... Figure 18A , 18B And 18C as described.
[0149] Figure 9 This is a diagram illustrating an example of motion vector prediction in affine inter-frame (AF_INTER) mode. In JEM, there are two affine motion modes: affine inter-frame (AF_INTER) mode and affine merge (AF_MERGE) mode. In some examples, the AF_INTER mode can be applied when the CU has a width and height greater than 8 pixels. Affine flags about blocks (e.g., at the CU level) can be placed in the bitstream (or signaled) to indicate whether the AF_INTER mode has been applied to the block. Figure 9In the example, under AF_INTER mode, adjacent blocks can be used to construct a candidate list of motion vector pairs. For instance, for a child block 910 located at the top left corner of the current block 902, motion vectors can be selected from the adjacent block 920 to the top left of child block 910, the adjacent block B 922 above child block 910, and the adjacent block C 924 to the left of child block 910. v 0 As another example, for the child block 912 located at the upper right corner of the current block 902, the motion vector can be selected from the adjacent blocks D926 and E928 in the upper and upper right directions, respectively. v 1 A candidate list of motion vector pairs can be constructed using adjacent blocks. For example, given motion vectors corresponding to blocks A 920, B 922, C 924, D 926, and E 928 respectively. v A , v B , v C , v D and v E The candidate list of motion vector pairs can be expressed as {( v 0 , v 1 ) | v 0 = { v A , v B , v C}, v 1 = { v D , v E}}.
[0150] As described above and as Figure 9 As shown, in AF_INTER mode, a motion vector can be selected from the motion vectors of blocks A 920, B 922, or C 924. v 0 Motion vectors from neighboring blocks (e.g., blocks A, B, or C) can be scaled based on a reference list and the relationship between the POCs used for references to neighboring blocks, the POCs used for references to the current CU (e.g., current block 902), and the POCs of the current CU. In these examples, some or all of the POCs can be determined from the reference list. Selected from neighboring blocks D926 or E928. v 1 Similar to a choice .
[0151] In some cases, if the candidate list contains fewer than two candidates, the candidate list can be populated using motion vector pairs by copying each AMVP candidate from the AMVP candidates. When the candidate list contains more than two candidates, in some examples, the candidates in the candidate list can be sorted first based on the consistency of adjacent motion vectors (e.g., consistency could be based on the similarity between two motion vectors in a motion vector pair candidate). In such examples, the first two candidates are retained, and the remaining candidates can be discarded.
[0152] In some examples, rate-distortion (RD) cost checks can be used to determine which motion vector pair candidate to select as the Control Point Motion Vector Prediction (CPMVP) for the current CU (e.g., current block 902). In some cases, the index of the CPMVP's position in the candidate list can be signaled (or otherwise indicated) in the bitstream. Once the CPMVP of the current affine CU (based on the motion vector pair candidates) is determined, affine motion estimation can be applied, and the Control Point Motion Vector (CPMV) can be determined. In some cases, the difference between the CPMV and CPMVP can be signaled in the bitstream. Both the CPMV and CPMVP consist of two sets of translational motion vectors, in which case the signaling cost of the affine motion information is higher than the signaling cost of the translational motion.
[0153] Figure 10A and Figure 10B An example of motion vector prediction in AF_MERGE mode is shown. When decoding the current block 802 (e.g., CU) using AF_MERGE mode, motion vectors can be obtained from valid neighboring reconstructed blocks. For example, the first valid neighboring reconstructed block from those decoded using affine mode can be selected as a candidate block. Figure 10A As shown, neighboring blocks can be selected from the set of adjacent blocks A 1020, B 1022, C 1024, D 1026, and E 1028. A specific selection order can be used to consider neighboring blocks as candidate blocks. An example of this selection order is the left neighbor (e.g., block A 1020), followed by the top neighbor (block B 1022), the top right neighbor (block C 1024), the bottom left neighbor (block D 1026), and the top left neighbor (block E 1028).
[0154] As mentioned above, the selected adjacent block can be the first block that has already been decoded using affine mode (e.g., in the order of selection). For example, block A 820 may have already been decoded in affine mode. Figure 10BAs shown, block A 1020 can be included in adjacent CU 1004. For adjacent CU 1004, the motion vector for the upper left corner of adjacent CU 1004 may have been derived. v 2 1030), used for the motion vector in the upper right corner ( v 3 1032) and used for the bottom left corner ( v 4 The motion vector of 1034). In the example above, according to v 2 1030 v 3 1032 and v 4 1034 is used to calculate the motion vector of the control point at the top left corner of the current block 1002. v 0 1040. The motion vector of the control point at the upper right corner of the current block 1002 can be determined. v 1 1042.
[0155] Once the control point motion vector (CPMV) of the current block 1002 has been derived ( v 0 1040 and v 1 1042), then equation (1) can be applied to determine the motion vector field for the current block 1002. In order to identify whether the current block 1002 is decoded using AF_MERGE mode, an affine flag can be included in the bit stream when there is at least one adjacent block decoded in affine mode.
[0156] In many cases, the process of affine motion estimation involves determining the affine motion for a block on the encoder side by minimizing the distortion between the original block and the affine motion prediction block. Because affine motion has more parameters than translational motion, affine motion estimation can be more complex. In some cases, a fast affine motion estimation method based on the Taylor expansion of the signal can be performed to determine the affine motion parameters (e.g., affine motion parameters a, b, c, d in a 4-parameter model).
[0157] Fast affine motion estimation can include gradient-based affine motion search. For example, given a pixel value at time t... (Where t0 is the time of the reference image), for pixel values The first-order Taylor expansion can be determined as: Equation (2)
[0158] in and These are the pixel gradients in the x and y directions, respectively. , ,and and Instructions for pixel values Motion vector components and Used for pixels in the current block. The motion vector points to the pixels in the reference image. .
[0159] Equation (2) can be rewritten as equation (3) as follows: Equation (3)
[0160] By making predictions ( Minimize the distortion between the original signal and the pixel value to solve for the pixel value. Affine motion and Taking the 4-parameter affine model as an example, Equation (4) Equation (5)
[0161] Where x and y indicate the position of a pixel or sub-block. Substituting equations (4) and (5) into equation (3), and using equation (3) to minimize the distortion between the original signal and the prediction, the solution for the affine parameters a, b, c, and d can be determined: Equation (6)
[0162] Once the affine motion parameters that define the affine motion vectors used for the control points are determined, the affine motion parameters (e.g., using equations (4) and (5), which are also expressed in equation (1)) can be used to determine the motion vectors per pixel or per sub-block. Equation (3) can be performed for each pixel of the current block (e.g., CU). For example, if the current block is 16 pixels × 16 pixels, the least-squares solution in equation (6) can be used to derive the affine motion parameters (a, b, c, d) for the current block by minimizing the overall value over 256 pixels.
[0163] Any number of parameters can be used in affine motion models for video data. For example, 6-parameter affine motion or other affine motions can be solved in the same way as described above for the 4-parameter affine motion model. For example, a 6-parameter affine motion model can be described as follows:
[0164] In equation (7), ( ) is the coordinate ( The motion vector at point () is given, and a, b, c, d, e, and f are six affine parameters. The affine motion model for a block can also be derived from three motion vectors (MV) at the three corners of the block. , and To describe.
[0165] Figure 11 This is a diagram showing the affine motion field of the current block 1102, described by three motion vectors 1120, 1122, and 1124 at three corresponding control points 1110, 1112, and 1114. Motion vector 1120 (e.g., At control point 1110, located at the top left corner of the current block 1102, motion vector 1122 (e.g., At control point 1112 located at the upper right corner of the current block 1102, and motion vector 1124 (e.g., The control point 1114 is located at the lower left corner of the current block 1102. The motion vector field (MVF) of the current block 1102 can be described by the following equation:
[0166] Equation (8) represents a 6-parameter affine motion model, where, and It represents the width and height of the current block 1102.
[0167] Although the above reference equation (1) describes the 4-parameter motion model, a simplified 4-parameter affine model using the width and height of the current block can be described by the following equation:
[0168] The simplified 4-parameter affine model for the block based on equation (9) can be derived from two motion vectors at two of the four corners of the block. and To describe. A sports field can be described as:
[0169] As mentioned earlier, motion vectors This is referred to as the Control Point Motion Vector (CPMV) in this paper. The CPMV used for a 4-parameter affine motion model is not necessarily the same as the CPMV used for a 6-parameter affine motion model. In some examples, different CPMVs can be chosen for the affine motion model.
[0170] Figure 12This is a diagram illustrating the selection of control point vectors for the affine motion model of the current block 1202. Four control points 1210, 1212, 1214, and 1216 are shown for the current block 1202. Motion vector 1220 (e.g., At control point 1210 located at the top left corner of the current block 1202, motion vector 1222 (e.g., At control point 1212, located at the upper right corner of the current block 1202, motion vector 1224 (e.g., At control point 1214 located at the lower left corner of the current block 1202, and motion vector 1226 (e.g., At control point 1216, located at the lower right corner of the current block 1202.
[0171] In one example, for a 4-parameter affine motion model (according to equation (1) or equation (10)), it can be derived from four motion vectors. Control point pairs can be selected from any two motion vectors. In another example, for a 6-parameter affine motion model, control points can be selected from four motion vectors. A control point pair is selected from any three motion vectors. Based on the selected control point motion vectors, other motion vectors for the current block 1002 can be calculated, for example, using the derived affine motion model.
[0172] In some examples, alternative affine motion models can also be used. For instance, an affine motion model based on incremental motion vectors can be represented by coordinates ( Anchor motion vector at point ) Horizontal incremental motion vector and vertical incremental motion vector Representation. Typically, coordinates ( The motion vector at point ) It can be calculated as .
[0173] In some examples, the CPMV-based affine motion model representation can be converted into an alternative affine motion model representation with incremental motion vectors. For example, in the incremental motion vector affine motion model representation... Same as CPMV in the upper left corner. It is important to note that for these vector operations, addition, division, and multiplication are applied element-wise.
[0174] In some examples, an affine motion predictor can be used to perform affine motion vector prediction. In some examples, the affine motion predictor for the current block can be derived from the affine motion vectors or normal motion vectors of adjacent decoded blocks. As described above, the affine motion predictor can include inherited affine motion vector predictors (e.g., those inherited using the affine merge (AF_MERGE) mode) and constructed affine motion vector predictors (e.g., those constructed using the affine inter-frame (AF_INTER) mode).
[0175] The inherited affine motion vector predictor (MVP) uses one or more affine motion vectors from neighboring decoded blocks to derive the predicted CPMV for the current block. The inherited affine MVP is based on the assumption that the current block shares the same affine motion model with neighboring decoded blocks. Neighboring decoded blocks are referred to as neighboring blocks or candidate blocks. Neighboring blocks can be selected from different spatial or temporal proximity locations.
[0176] Figure 13 This is a diagram showing the affine MVP of the current block 1302, inherited from the neighboring block 1304 (block A). The affine motion vector of the neighboring block 1304 is based on the corresponding motion vectors 1330, 1332, and 1334 at control points 1320, 1322, and 1324. It is expressed as follows: In one example, the size of the adjacent block 1304 can be represented by parameters (w, h), where w is the width of the adjacent block 1304 and h is the height of the adjacent block 1304. The coordinates of the control points of the adjacent block 1304 are represented as (x0, y0), (x1, y1), and (x2, y2). Affine motion vectors 1340, 1342, and 1344 can be predicted for the current block 1302 at the corresponding control points 1310, 1312, and 1314, and are represented as follows: By replacing (x, y) in equation (8) with the coordinate difference between the control point of the current block 1302 and the upper left control point of the adjacent block 1304, the predicted affine motion vector for the current block 1302 can be derived. As described in the following equation:
[0177] In equations (11)-(13), (x0', y0'), (x1', y1'), and (x2', y2') are the coordinates of the control points of the current block 1102. If represented as incremental MV, then ,and , .
[0178] Similarly, if the affine motion model of an adjacent decoded block (e.g., adjacent block 1304) is a 4-parameter affine motion model, equation (10) can be applied when deriving the affine motion vector at the control point for the current block 1102. In some examples, using equation (10) to obtain the 4-parameter affine motion model may include circumventing the above equation (13).
[0179] Figure 14 This is a diagram showing the possible locations of neighboring candidate blocks for use in the affine MVP model inherited for the current block 1402. For example, the affine motion vectors 1440, 1442, and 1444 at control points 1410, 1412, and 1414 of the current block, or... It can be derived from one of the adjacent blocks 1430 (block A0), 1426 (block B0), 1428 (block B1), 1432 (block A1), and / or 1420 (block B2). In some cases, adjacent blocks 1424 (block A2) and / or 1422 (block B3) can also be used. More specifically, the motion vector 1440 at the control point 1410 located at the top left corner of the current block 1402 (e.g., The motion vector 1442 can be inherited from the adjacent block 1420 (block B2) located to the upper left of control point 1410, the adjacent block 1422 (block B3) located above control point 1410, or the adjacent block 1424 (block A2) located to the left of control point 1410; the motion vector 1442 located at control point 1412 at the upper right corner of the current block 1402 can also be inherited from the adjacent block 1420 (block B2). It can inherit from the adjacent block 1426 (block B0) above control point 1410 or the adjacent block 1428 (block B1) to the upper right of control point 1410; and the motion vector 1444 at control point 1414 located at the lower left corner of the current block 1402. It can be inherited from the adjacent block 1430 (block A0) located to the left of control point 1410 or the adjacent block 1432 (block A1) located to the lower left of control point 1410.
[0180] Currently, in some designs (e.g., in MPEG5 Elementary Video Decoding (EVC)), when affine inheritance is from the affine decoding adjacent block in the upper CTU row, the lower left and lower right sub-block MVs are used as CPMVs, and a 4-parameter affine model is always used to derive the CPMV of the current CU.
[0181] Figure 15 This is a diagram illustrating the affine model and spatial neighborhood in MPEG5 EVC. Figure 15 The current CTU 1500 is shown, which has an adjacent candidate block 1540 (with sub-blocks 1542 and 1544) and a left neighbor sub-block 1552 and a lower left neighbor sub-block 1554. Although Figure 15The use of a CTU provides an illustrative example, but in other examples, the current CTU could be another block, such as CU, PU, TU, etc. Current CTU 1500 includes current block 1502 with control points 1510, 1512, and 1514, and associated CPMVs 1520, 1522, and 1524. The upper left CPMV 1520 is referred to as... And CPMV 1522 in the upper right corner is shown as In the example shown, CPMVs 1520 and 1522 are designated as the CPMVs of the current CU (e.g., current block 1502), which will be used by adjacent affine decoding CUs (including candidate block 1540 located above the current CTU 1500) with associated MVs. (Not shown) The lower left sub-block 1542 and having motion vectors The motion vector of the lower right sub-block 1544 (not shown) is used for derivation. CPMV 1520 and 1522 are in... Figure 15 The middle is shown as and And it can be derived through the following equation: Equation (14) Where neiW is the width of the adjacent block, curW is the width of the current block, posNeiX is the x-coordinate of the top-left pixel of the adjacent block (or, in some examples, the sample), and posCurX is the x-coordinate of the top-left pixel of the current block (or, in some examples, the sample).
[0182] In some cases, three affine prediction motion modes exist: AF_4_INTER mode, AF_6_INTER mode, and AF_MERGE mode. When the merge / skip flag is true (e.g., equal to a value of 1) and both the width and height of the CU are greater than or equal to 8 samples (or other sample numbers), the affine flags at the CU level (or other block level) are signaled in the bitstream to indicate whether an affine merge mode is used. Furthermore, when the CU is decoded as AF_MERGE, a merge candidate index with a maximum value of 4 (or, in some cases, other values) is signaled to specify which motion information candidate from the affine merge candidate list is used for the CU.
[0183] The affine merging candidate list can be constructed by following these steps: 1) Insert model-based affine candidates, where the model-based candidates are derived from the affine motion models of their effective spatially adjacent affine decoded blocks. The scan order used for candidate positions can be... Figure 6A , Figure 6B and / or Figure 6CThe merged lists in the first two parts have the same order and include positions from 0 to 5. 2) Insert affine candidates based on control points. If the size limit for the affine merged list is not met, insert affine candidates based on control points. Affine candidates based on control points mean that candidates are constructed by combining the adjacent motion information of each control point to form an affine merged candidate.
[0184] Use a total of 4 control points or CPs (denoted as CP1-CP4), which have coordinates (0, 0), (W, 0), (H, 0) and (W, H) respectively, where W and H are the width and height of the current block.
[0185] To simplify the construction of the affine merge list, scaling is not performed when deriving control point-based affine merge candidates. Candidates are considered unavailable if the control point motion vectors point to different reference indices or if the reference indices are invalid.
[0186] When the merge / skip flag is false (e.g., equal to a value of 0) and both the width and height of the CU are greater than or equal to 16 samples (or, in some cases, other sample numbers), the affine flag at the CU level is signaled in the bitstream to indicate whether an affine inter-frame mode (e.g., AF_4_INTER mode or AF_6_INTER mode) is used. When the CU is decoded in affine inter-frame mode, the model flag is signaled to specify whether a 4-parameter affine model or a 6-parameter affine model is used for the CU. If the model flag is true (e.g., equal to a value of 1), the AF_6_INTER mode (6-parameter affine model) is applied, and 3 MVDs are parsed; otherwise, if the model flag is false (e.g., equal to a value of 0), the AF_4_INTER mode (4-parameter affine model) is applied, and 2 MVDs are parsed.
[0187] The affine AMVP candidate list can be constructed by following these steps: 1) inserting model-based affine candidates; 2) inserting control point-based affine candidates; 3) inserting translation-based affine AMVP candidates; and 4) filling with zero motion vectors.
[0188] If the number of candidates in the affine merge candidate list is less than 2 (or, in some cases, other values), a zero motion vector with a zero reference index is inserted until the list is full. To reduce the complexity of list construction, pruning is not applied.
[0189] Sample-derived affine patterns can be performed for small block sizes (e.g., 4x8 and 8x4). In MPEG5 EVC, the minimum block size for affine decoding is set to 8x8. However, the encoder can choose to implement affine prediction in sub-block sizes of 4x8 or 8x4. MPEG5 EVC specifies affine prediction for such sub-block sizes using an Enhanced Interpolation Filter (EIF). The EIF implementation utilizes per-sample prediction, which computes motion vectors independently for each sample. To prevent the motion vector (MV) from pointing outside the reference image, the MV obtained for each sample is cropped to the image size. The following excerpt from MPEG EVC illustrates an implementation of affine prediction using the EIF, which is used in… <highlight>"and" <highlighted>"The underlined text between the symbols is used to mark (e.g., " <highlight> Highlighted text <highlighted>(): If affine_flag equals 1 and either sbWidth or sbHeight is less than 8, then the following applies: – The horizontal change dX, the vertical change dY, and the base motion vector mvBaseScaled of the motion vector are derived by invoking the procedure specified in Clause 8.5.3.9, with the luminance decoder block width nCbW, the luminance decoder block height nCbH, the number of control point motion vectors numCpMv, and the control point motion vector cpMvLX[cpIdx] (where cpIdx = 0...numCpMv-1) as inputs. – The array predSamplesLXL is derived by invoking the interpolation procedure for the enhanced interpolation filter as specified in Clause 8.5.4.3, where the luminance position (xSb, ySb), luminance decoder block width nCbW, luminance decoder block height nCbH, horizontal variation of the motion vector dX, vertical variation of the motion vector dY, base motion vector mvBaseScaled, reference array refPicLXL, sample bit depth bitDepthY, image width pic_width_in_luma_samples, and height pic_height_in_luma_samples are taken as inputs.
[0190] 1.1.1.1 Derivation of the parameters of the affine motion model from the motion vectors of the control points The input to this process is: – Two variables, cbWidth and cbHeight, specify the width and height of the luminance decoding block. – Number of control point motion vectors numCpMv – Control point motion vector cpMvLX[ cpIdx ], where cpIdx = 0..numCpMv - 1, and X equals 0 or 1.
[0191] The output of this process is: – Horizontal change of the motion vector dX, – The vertical change of the motion vector dY, – Motion vector mvBaseScaled, which corresponds to the top left corner of the luminance decoding block.
[0192] The variables log2CbW and log2CbH are derived as follows: log2CbW = Log2( cbWidth ) (8-688) log2CbH = Log2( cbHeight ) (8-689)
[0193] The horizontal change dX of the motion vector is derived as follows: dX[ 0 ] = ( cpMvLX[ 1 ][ 0 ] - cpMvLX[ 0 ][ 0 ] ) << ( 7 - log2CbW )(8-690) dX[ 1 ] = ( cpMvLX[ 1 ][ 1 ] - cpMvLX[ 0 ][ 1 ] ) << ( 7 - log2CbW )(8-691)
[0194] The vertical change dY of the motion vector is derived as follows: – If numCpMv equals 3, then dY is derived as follows: dY[ 0 ] = ( cpMvLX[ 2 ][ 0 ] - cpMvLX[ 0 ][ 0 ] ) << ( 7 - log2CbH )(8-692) dY[ 1 ] = ( cpMvLX[ 2 ][ 1 ] - cpMvLX[ 0 ][ 1 ] ) << ( 7 - log2CbH )(8-693) – Otherwise (numCpMv equals 2), dY is derived as follows: dY[0] = -dX[1] (8-694) dY[1] = dX[0] (8-695)
[0195] The following derivation corresponds to the motion vector mvBaseScaled at the top left corner of the luminance decoding block: mvBaseScaled[ 0 ] = cpMvLX[ 0 ][ 0 ] << 7 (8-696) mvBaseScaled[ 1 ] = cpMvLX[ 0 ][ 1 ] << 7 (8-697)
[0196] 1.1.1.2 Interpolation process for enhanced interpolation filters The input to this process is: – The position (xCb, yCb) within the entire sample cell. – Two variables, cbWidth and cbHeight, specify the width and height of the current coded block. – Horizontal change of the motion vector dX, – The vertical change of the motion vector dY, – Motion vector mvBaseScaled – The selected reference image sample array refPicLX, – Sample bit depth – The width of the image in the sample, pic_width. – The height of the image in the sample, pic_height.
[0197] The output of this process is: – The array predSamplesLX, representing the predicted sample values (cbWidth) x (cbHeight).
[0198] The variables shift1, shift2, shift3, offset1, offset2, and offset3 are derived as follows: shift0 is set to equal bitDepth – 6, and offset0 is set to 2. shift1-1 , shift1 is set to 11, and offset1 is equal to 1024. <highlight>for For x = -1.. cbWidth and y = -1.. cbHeight, the following applies: –The motion vector mvX is derived as follows: mvX[ 0 ] = ( mvBaseScaled[ 0 ] + dX[ 0 ] x + dY[0] y ) (8-728) mvX[ 1 ] = ( mvBaseScaled[ 1 ] + dX[ 1 ] x + dY[ 1 ] y ) (8-729) <highlightend> – The variables xInt, yInt, xFrac, and yFrac are derived as follows: xInt = xCb + ( mvX[ 0 ]>>9 ) + x (8-730) yInt = yCb + ( mvX[ 1 ]>>9 ) + y (8-731) xFrac = mvX[ 0 ] & 511 (8-732) yFrac = mvX[ 1 ] & 511 (8-733)
[0199] The following derivation shows the position (xInt, yInt) within a given array refPicLX: <highlight> xInt = Clip3( 0, pic_width - 1, xInt ) (8-734) yInt = Clip3( 0, pic_height -1, yInt ) (8-735) <highlightend>
[0200] The following is a derivation of variable a x,y a x+1,y a x,y+1 a x+1,y+1 : a x,y = ( ( refPicLX[ xInt ][ yInt ] ( 512 – xFrac ) + offset0 ) >>shift0 ) (512 – yFrac) (8-736) a x+1,y = ( ( refPicLX[ xInt + 1 ][ yInt ] xFrac + offset0) >> shift0) (512 – yFrac) (8-737) a x,y+1 = ( ( refPicLX[ xInt ][ yInt + 1 ] ( 512 – xFrac ) + offset0 )>> shift0 ) yFrac (8-738) a x+1,y+1 = ( ( ( refPicLX[ xInt ][ yInt ] xFrac + offset0) >> shift0) yFrac (8-739)
[0201] The following derivation gives the sample value b corresponding to position (x, y). x,y : b x,y = (a x,y + a x+1,y + a x,y+1 + a x+1,y+1 + offset1 ) >> shift1 (8-740)
[0202] The enhancement interpolation filter coefficients eF[] are specified as {-1, 10, -1}.
[0203] The variables shift2, shift3, offset2, and offset3 are derived as follows: shift2 is set to 4, and offset2 is set to 8. shift3 is set to 15 – bitDepth, and offset3 equals 2. shift3-1 ,
[0204] For x = 0...cbWidth–1 and y = –1...cbHeight, the following applies: -h x,y = (eF[0]) b x-1,y + eF[ 1 ] b x,y + eF[ 2 ] b x+1,y + offset2 )>>shift2 (8-741)
[0205] For x = 0.. cbWidth – 1 and y = 0.. cbHeight – 1, the following applies: –predSamplesLX L [x][y]= Clip3( 0, ( 1 << bitDepth ) – 1, ( eF[ 0 ] h x,y-1 + eF[ 1 ] h x,y + eF[ 2 ] b x,y+1 + offset3 )>>shift3)
[0206] The per-sample MV generation introduced in Enhanced Interpolation Filters (EIFs) can potentially increase the number of memory accesses required to retrieve filter samples, thereby increasing memory bandwidth. This increase in memory accesses can be significantly greater than the one memory retrieval typically used in unidirectional prediction for block sizes of 4x8 or 8x4 (or other block sizes) or the two memory retrievals for bidirectional prediction blocks.
[0207] As mentioned above, a large number of retrieval operations may not be a problem, for example, if the required reference region is available in the local buffer. Current EIF designs introduce MV cropping to image boundaries, which requires the entire image to be available in the local buffer. As described above, this paper describes techniques and systems for improving affine pattern decoding. Each technique described herein can be performed individually or in any combination. In some examples, the systems and techniques described herein limit (using restrictions or constraints) the reference image region accessed from the generation of affine samples (e.g., via EIF) to a specific constraint, which in some cases can be set as a function of the block size. In some examples, the systems and techniques apply restrictions or constraints to certain block sizes, such as less than 8x8, or less than 4x8, or less than 8x4, or other block sizes. In some cases, restrictions or constraints can be specified as a function of the block size.
[0208] Limitations or constraints can be imposed in various ways. An illustrative and non-limiting example of such constraints can be implemented as a modification of the MPEG EVC description used for affine motion constraints. For example, an encoding and / or decoding device can constrain and / or trim one or more affine motion vectors or their outputs (e.g., the coordinates of a reference sample to which the affine motion vector points) such that the constraint / trimming ensures that no higher-granularity affine vector (e.g., a sub-block or sample) will exceed the allowed area. Two examples of forms in which such constraints can be introduced include bitstream requirements (consistency) and canonical decoding procedures. An illustrative example of a canonical decoding procedure that can be implemented by trimming one or more affine motion vectors (MVs) is as follows (via "... <insert>"and" <insertend>Modify the highlighted section above by adding underlined text between the symbols (e.g., " <insert>Added text <insertend>(): The motion vector mvX is derived as follows: mvX[ 0 ] = ( mvBaseScaled[ 0 ] + dX[ 0 ] x + dY[0] y ) (8-728) mvX[ 1 ] = ( mvBaseScaled[ 1 ] + dX[ 1 ] x + dY[ 1 ] y ) (8-729) <insert> mvX[ 0 ] = Clip3( MinX, MaxX - 1, mvX[ 0 ] ) mvX[ 1 ] = Clip3( MinY, MaxX - 1, mvX[ 1 ] ) <insertend> The clipping parameters are derived as a function of the block size, the current block / sample coordinates, and the MV.
[0209] In the above case, spatial coordinate clipping is not required, and spatial coordinate clipping can be removed from one or more implementations, as shown below. <delete>and <deleteend>Crossed-out text between symbols <delete>Deleted text <deleteend>(Refer to the corresponding sections 8-734 and 8-735 shown above:) <delete>xInt = Clip3( MinX, MaxX - 1, xInt ) (8-734) yInt = Clip3( MinY, MaxY -1, yInt ) (8-735) <deleteend> iii. Standardize the decoding process, which can be achieved by retrieving data by cropping the actual coordinates, as follows: xInt = Clip3(MinX, MaxX - 1, xInt) (8-734) yInt = Clip3( MinY, MaxY -1, yInt ) (8-735) The clipping parameters are derived as a function of the block size, the current block / sample coordinates, and the MV.
[0210] In some examples, clipping parameters can be derived by considering one or more MVs derived for different spatial locations of the affine block, such as the X / Y coordinates pointed to by MV v0 (top left CP) or by other MVs (e.g., v1 or v2) provided by the affine model. An example of such an implementation using thresholds is as follows: {minX, minY,maxX,maxY} = function(Threshold, {v0||v1||v2}, {x0,y0}) {minX, minY,maxX,maxY} = function(Threshold, {v0||v1||v2}, {x0,y0}) xInt = Clip3(MinX, MaxX - 1, xInt) yInt = Clip3(MinY, MaxY -1, yInt)
[0211] Figure 16 It is a diagram showing various aspects of the affine model and spatial neighborhood based on some examples. Figure 16 Showing from Figure 15 The current CTU 1500, as well as the adjacent blocks and sub-blocks above it, and the associated control points and motion vectors. Although Figure 16 The use of CTU provides an illustrative example, but in other examples, the current CTU could be another block, such as CU, PU, TU, etc.
[0212] As described in detail above, affine decoding of the current block 1502 can use reference data. Such reference data can be derived from... Figure 16 Reference 1670 is shown in the diagram. In some cases, reference 1670 may be a portion of a picture identified as a reference image for the current block 1502. In some cases, the affine motion vectors may not be consistent with the indications of significantly different portions of the reference image (e.g., they may vary considerably from the current block). For example, as mentioned above, since affine motion may be associated with motion due to changes in viewpoint (e.g., movement of the camera position), it can be expected that the affine motion vectors will be fairly consistent across the block. However, in some cases, the affine motion vectors for one sample of the block may be significantly different from those for another sample of the same block (e.g., pointing in significantly different directions). When this occurs, the memory bandwidth used to access the reference data indicated by the affine motion vectors (e.g., reference 1670) may degrade performance.
[0213] The examples described herein may include devices (e.g., encoding device 104 or decoding device 112) that perform clipping of affine motion vectors to limit (e.g., to boundary region 1660) the data in reference 1670 that may be indicated by affine motion vectors. In some examples, such clipping may be accomplished using a threshold (e.g., from "Threshold" above). In some examples, the threshold may be a user- and / or system-specified block size ratio, which serves as a criterion for defining boundary region 1660 (e.g., it may be considered a memory access region, which is the region of reference 1670 stored in memory or DCB for decoding the current block 1502), such as... Figure 16 As shown. Reference 1670 (e.g., a reference image or a portion of a reference image, such as a reference block) includes a boundary region 1660 (e.g., a portion of reference 1670), which may be pointed to by one or more affine motion vectors based on constraints (e.g., clipping parameters) applied to the affine motion vectors. Arrow 1690 indicates the relationship between samples or points in the current block 1502 and the boundary region 1660, such that the affine motion vectors are constrained (e.g., by clipping parameters) to the boundary region 1660. Depending on different affine motion parameters, the relationship between samples in the current block 1502 and the data in the boundary region 1660 indicated by the affine vectors can be changed to match a specific affine motion being decoded in affine decoding mode. The following is about Figure 18A , Figure 18B and Figure 18C Additional details are provided regarding the relationship between samples of the current block (e.g., current block 1502) and data referenced from a reference image (e.g., data from boundary region 1660 in reference 1670). In many cases, the affine motion vector will have a consistent value across the current block (e.g., current block 1502) (e.g., due to the nature of affine motion, such as viewpoint movement as described above), in which case the performance degradation caused by limiting the affine motion vector will generally be limited.
[0214] Limiting possible reference data to boundary region 1660 prevents performance degradation associated with memory bandwidth and limits the possible data to be referenced to a manageable size that can be buffered in memory and used for affine decoding of the current block 1502. The clipping parameters described herein (e.g., a (cbWidth)x(cbHeight) array; variables such as a horizontal maximum, horizontal minimum, vertical maximum, and vertical minimum; or any other such parameters used to limit reference picture data indicated by an affine motion vector for the current block (such as current block 1502)) can be used in various examples to define boundary region 1660 in the context of current block 1502 and reference 1670, and can also be used to store reference data associated with boundary region 1660 for use during decoding of the current block 1502.
[0215] Figure 17 This is a diagram illustrating various aspects of the affine model and spatial neighborhood based on some examples. Figure 16 similar, Figure 17 Showing from Figure 15 The current CTU 1500, as well as adjacent blocks and sub-blocks, and associated control points and motion vectors. Although Figure 17 The use of CTU provides an illustrative example, but in other examples, the current CTU can be another block, such as CU, PU, TU, etc. In some examples, such as Figure 17 As shown and as described above, clipping parameters can be derived by considering one or more motion vectors derived for different spatial locations of the affine block (e.g., the actual affine MV generated for the affine sub-block) or affine samples within the current block. Example implementations may include clipping parameters as follows: {minX, minY,maxX,maxY} = function(Threshold, {mv(x,y)}, {x,y}) {minX, minY,maxX,maxY} = function(Threshold, { mv(x,y)}, {x,y})
[0216] The clipped motion vectors are as follows: xInt = Clip3(MinX, MaxX - 1, xInt) yInt = Clip3(MinY, MaxY -1, yInt).
[0217] Other examples may include other implementations of such motion vectors and clipping parameters.
[0218] As described above, the threshold (e.g., a threshold indicating the boundary region 1660 of reference data that can be indicated by an affine motion vector) can be a user- and / or system-specified block size ratio, which serves as a criterion for defining memory access regions (e.g., boundary region 1760), such as... Figure 17 As shown. In Figure 17 In the middle, refer to the boundary region 1760 (e.g., similar to...). Figure 16 Boundary region 1660) specifies the boundary region for data used in the reference image, which can be pointed to by an affine motion vector from a sample of the current block 1502 (e.g., given a clipping constraint indicating to prevent data from the reference image outside boundary region 1760 from being included). Current block reference region 1750 shows an example of the block size of the current block 1502 to which the motion vector for the center position is pointed (e.g., under the assumption of using translational motion associated with arrow 1790). Due to acceptable variations in the affine vectors, the area of the reference image accessible to process the current block 1502 (e.g., boundary region 1760) is larger than the current block reference region 1750, as described in more detail below. Depending on the constraint or clipping parameters implemented as part of affine decoding, the affine motion vector generated for the sample position of the current block 1502 or for a vector from a related sub-block (such as sub-blocks 1542, 1544, 1552, or 1554) will point to a position within boundary region 1660. The following discusses… Figure 18A , Figure 18B and Figure 18C Additional details are described regarding the affine motion vectors and associated reference data that are subject to clipping or thresholding.
[0219] In some examples, scaling and / or pruning of the CP motion vector of the CU can be performed, or the resulting motion vector variation parameters (dXmv, dYmv) can be used to verify that higher-granularity (e.g., sub-blocks or samples) affine vectors do not exceed allowed regions (e.g., boundary regions 1660 or 1760). Two examples of forms in which such constraints can be introduced include bitstream requirements (e.g., for consistency) and canonical decoding procedures. Canonical decoding procedures can ensure that constraints imposed on affine motion vectors are achieved by pruning the CP MV or readjusting / scaling the CP MV or changing the parameters.
[0220] In some examples, the MV{v0, v1, v2} of the CP position can be clipped before being used in the affine MV derivation. For example, one of the CP MVs (e.g., v0) can be used as the base, and other CP MVs (e.g., v1 and v2) can be checked to determine if the other CP MVs point outside the boundary region (this can be referred to as checking for boundary block violations). If such a boundary block violation is identified, the identified vector can be scaled and pointed within the boundary region, one side (corner) of which is specified by the base MV (e.g., v0). Similar techniques can be applied to affine motion models with fewer than three CP motion vectors (e.g., fewer than v0, v1, v2) and / or more than three CP motion vectors (e.g., more than v0, v1, v2).
[0221] In another example, the motion information of v0, v1, and v2 can remain unchanged; however, the affine parameters dX and dY will be scaled accordingly to prevent the affine MV from pointing outside the boundary block. The following utilizes... <insert2>"and" <insertend2>Text marked with underlined text between symbols (e.g., " <insert2> Added text <insertend2>The image shows an example of such an implementation. Similar techniques can be applied to affine motion models with fewer than three CP motion vectors (e.g., fewer than v0, v1, v2) and / or more than three CP motion vectors (e.g., more than v0, v1, v2). 1.1.1.3 Derivation of the parameters of the affine motion model from the motion vectors of the control points The input to this process is: – Two variables, cbWidth and cbHeight, specify the width and height of the luminance decoding block. – Number of control point motion vectors numCpMv – Control point motion vector cpMvLX[ cpIdx ], where cpIdx = 0..numCpMv - 1, and X equals 0 or 1. The output of this process is: – Horizontal change of the motion vector dX, – The vertical change of the motion vector dY, – Motion vector mvBaseScaled, which corresponds to the top left corner of the luminance decoding block. The variables log2CbW and log2CbH are derived as follows: log2CbW = Log2( cbWidth ) (8-688) log2CbH = Log2( cbHeight ) (8-689) <insert1> The call clips the cpMvLX motion vectors to the bounding block size (wBB, hBB, cbWidth, ... cbHeight, xCb, yCb, ratio); <insertend1> The horizontal change dX of the motion vector is derived as follows: dX[ 0 ] = ( cpMvLX[ 1 ][ 0 ] - cpMvLX[ 0 ][ 0 ] ) << ( 7 - log2CbW )(8-690) dX[ 1 ] = ( cpMvLX[ 1 ][ 1 ] - cpMvLX[ 0 ][ 1 ] ) << ( 7 - log2CbW )(8-691) The vertical change dY of the motion vector is derived as follows: – If numCpMv equals 3, then dY is derived as follows: dY[ 0 ] = ( cpMvLX[ 2 ][ 0 ] - cpMvLX[ 0 ][ 0 ] ) << ( 7 - log2CbH )(8-692) dY[ 1 ] = ( cpMvLX[ 2 ][ 1 ] - cpMvLX[ 0 ][ 1 ] ) << ( 7 - log2CbH )(8-693) – Otherwise (numCpMv equals 2), dY is derived as follows: dY[0] = -dX[1] (8-694) dY[1] = dX[0] (8-695) <insert2> From the cpmvLX motion vector, boundary region parameters, current block parameters cbWidth, cbHeight, and local... The scaling parameters scDX and scDY are derived using partial coordinates. The dX and dY parameters are scaled proportionally to prevent the resulting MV from pointing towards the boundary. Outside the block <insertend2> dX[0] = scDX dX[0] () dX[1] = scDX dX[1] () dY[0] = scDY dY[0] () dY[1] = scDY dY[1] () The following derivation corresponds to the motion vector mvBaseScaled at the top left corner of the luminance decoding block: mvBaseScaled[ 0 ] = cpMvLX[ 0 ][ 0 ] << 7 (8-696) mvBaseScaled[ 1 ] = cpMvLX[ 0 ][ 1 ] << 7 (8-697)
[0222] In some examples, a threshold is used instead of image boundaries to crop the accessible motion vectors and / or spatial coordinates for motion compensation. Cropping motion vectors or spatial coordinates using a threshold can be performed to benefit from existing cropping procedures (e.g., the cropping procedure given in the EVC standard). For example, the cropping parameters can be calculated once per block based on parameters made into a table (where in " <highlight>"and" <highlighted>Underlined text is used between symbols to emphasize the text (e.g., " <highlight> Highlighted text <highlighted> ”)): Deviation_A[5] = { 16, 80, 224, 512, 1088}; Deviation_B[5] = { 16, 96, 240, 528, 1104}; hor_min = (center_mv_hor - Deviation_A[log2(w) - 3])<<5; ver_min = (center_mv_ver - Deviation_A[log2(h) - 3])<<5; hor_max = (center_mv_hor + Deviation_B[log2(w) - 3])<<5; ver_max = (center_mv_ver + Deviation_B[log2(h) - 3])<<5; <highlight> mvX[ 0 ] = Clip3( mv_max[ 0 ], mv_min[ 0 ], mvX[ 0 ] ) (8- 734) mvX[ 1 ] = Clip3( mv_max[ 1 ], mv_min[ 1 ], mvX[ 1 ] ) (8-735) <highlightend>
[0223] As described herein, affine sample generation can be used for video decoding (e.g., video encoding and / or decoding), including standard-based decoding (such as EVC, VVC, and / or other existing or developing decoding standards). The affine decoding pattern in video decoding allows uncorrelated motion vectors for the current block (e.g., current block 1502) being decoded using predictive processing operations. In some such systems as described above, there is no single motion vector for the entire current block (e.g., CU, CTU, PU, TU, or other blocks). Instead, some samples within a block have independent affine motion vectors. Each sample in such a block can have an independent motion vector that may point to a considerable distance around a reference image identified for that block. An affine pattern decoder operating without constraints might call or retrieve regions from large areas of the reference image, which uses significant memory resources for predictive operations in decoding (e.g., using reference image data exceeding the capacity of a DPB). In some such systems, the Enhanced Interpolation Filter (EIF) generates an independent motion vector for each sample, and retrieving the data used for such vectors individually can be bandwidth-intensive and consume significant computational resources. The retrieved reference data is stored (e.g., buffered) in memory that stores samples of the region indicated by the motion vector from the reference image. To provide acceptable performance, the reference data that can be retrieved back to memory can be limited by a cropping parameter, as illustrated in the examples described herein.
[0224] Various methods can be used to impose constraints, including restricting coordinates, restricting the motion vectors predicted by the affine pattern, modifying the affine parameters used for clipping, using a segmentation table for clipping vectors outside a defined region to restrict the clipping of horizontal and vertical motion vectors, and using other such constraints. Some examples include devices and processes that restrict the magnitude of motion vectors to be defined by specific boundaries around the central location (e.g., Figure 16 The boundary region 1660 or Figure 17 The boundary region 1760, or the boundary region 1810 surrounding the center location 1854 of Figure 18 described below, is demarcated. Some such examples can be operated by obtaining the current block, obtaining control points for affine prediction, generating synthetic motion vectors, and approximating motion vectors for samples located at the center of the block. In some such examples, the center location can be used together with the DPB to store data from a reference (e.g., data from the boundary regions 1660 or 1760 of the reference image). In some examples, a minimum-maximum (min-max) deviation allowed for the affine motion vectors used for the block can be defined, where any vectors pointing outside the restricted region for the vectors are clipped to the restricted region.
[0225] In some examples, the decoding device is configured to retrieve different reference region sizes for different block sizes using an affine decoding mode. In some examples, the size ratio is configured by the device to be computationally feasible (e.g., without performance degradation) for a particular device or system. In some examples, the decoding device is configured to retrieve a specific number of reference samples per sample using an affine decoding mode. In some of these examples described herein, a threshold for a specific block size can be indicated as part of the affine decoding mode. Other thresholds can be used in other examples. In some examples, affine motion vector clipping parameters can be derived from the center vector in the reference image and other input values as part of the affine decoding mode operation. In some of these examples that utilize center samples with motion vectors, the motion vectors point to a position in the reference image. In these examples, the position in the reference image gives the center position of the reference region. The size of the reference region is defined by a bias value, which in some examples is fixed by the center position identified by the center motion vector.
[0226] In some examples, the cropping parameters are deviations that depend on the block size. For example, deviations A and B have specified values, described above as Deviation_A[5]= { 16, 80, 224, 512, 1088}; Deviation_B[5]={ 16, 96, 240, 528, 1104}. Such values are specified based on a specific size value (such as image resolution) and will be different for images with different size values (e.g., different image resolutions).
[0227] As described above, in some examples, samples in the current block that comply with affine decoding have affine motion vectors pointing to a reference image. The motion vectors set the center position from which a referenceable region (e.g., a region such as boundary region 1660 or 1760) can be defined. As part of the affine decoding operation, affine motion vectors are defined from a standard affine motion vector generation process. As part of the affine motion model-based affine decoding operation, control point motion vectors are determined and sub-block or sample motion vectors are derived.
[0228] Figure 18A This is a diagram illustrating various aspects of clipping based on some examples of using thresholds. For example... Figure 18A As shown in the example, block 1860 is the current CU implemented using the EIF affine decoding according to the example described herein. A reference block-size region 1862, defined by a reference image pointed to by the center motion vector 1850 of sample 1852 from block 1860, defines a reference block-size region 1862 of the same size as block 1860 (CU). Boundary region 1810 and center position 1854 of reference block-size region 1862. Regions 1864, 1866, and 1868 show the permissible deviations of motion vectors 1850, 1840, and 1830 corresponding to samples 1852, 1842, and 1832. In some examples, the region sizes of regions 1864, 1866, and 1868 are defined by a deviation given by an integer number of pixels ((MV(center) – 1) / (MV(center) + 1)) (e.g., for size 8).
[0229] Using the constraints or deviations defined as described above, in some examples, the upper left position 1834 associated with sample 1832 can allow a shift associated with the motion vector, for a deviation width / 2 (w / 2) and height / 2 (h / 2), while for region 1864, the boundary is still defined by the center MV(center) - 1 / MV(center) + 1 (e.g., center motion vector 1830). When applied to all samples for the current block 1860, the deviations and boundaries shown for region 1864 and position 1834 (e.g., associated with center motion vector 1830 and sample 1832) introduce a valid boundary block 1810 for memory access of samples for the current block 1860. Figure 18A In the example, boundary block 1810 can be considered as a reference block-sized region 1862 (e.g., having the same size as the current block 1860) plus the deviation_on_mv caused by the boundary regions around the extreme positions at the edges of region 1862 (such as regions 1864 and 1868 around positions 1834 and 1844). By imposing a clipping constraint on each sample of the current block 1860 to limit the motion vector to a deviation associated with the sizes of regions 1864, 1866, and 1868 (e.g., making these regions the same size, and the deviations for all other affine motion vectors from samples of the current block 1860 would be the same size), the affine pattern described herein introduces a constraint on each individual motion vector within the current block 1860. In some examples, such constraints are applied to the motion vector even if it is within the boundary block. Such a solution effectively provides a “translation” on affine motion, which will be discussed below. Figure 18B and 18C Further description.
[0230] Figure 18B This is a diagram illustrating various aspects of clipping using thresholds based on some examples. Figure 18B An example of a current block 1860 is shown, having a specific set of affine motion vectors 1836, 1847, and 1877 for corresponding samples 1832, 1842, and 1872. As described above, each sample is associated with a region defining the maximum deviation of the affine vectors associated with that particular sample. Sample 1832 is associated with region 1864, sample 1842 with region 1868, and sample 1872 with region 1876. As described above, if an affine motion vector is outside the restricted region defined for that vector (e.g., indicated by the center vector of the sample), a clipping operation is used to adjust the affine motion vector to produce a clipped affine motion vector that does not deviate from the associated region for the sample and therefore will not deviate from boundary block 1810. Since the data for boundary block 1810 can be stored in a local buffer as described above, the decoding device can perform operations for all samples of the current block 1860 without retrieving additional reference data and without degrading device performance due to excessive memory bandwidth usage.
[0231] exist Figure 18B In the example, the affine motion vector 1836 for sample 1832 is outside region 1864 (e.g., where center position 1834 is associated with sample 1832). The affine motion vector 1836 is clipped via an affine mode operation to create a clipped affine motion vector 1838 pointing within the border of region 1864 and boundary block 1810. In contrast, the motion vector 1877 for sample 1872, associated with center position 1874 and region 1876, points to position 1875. Since position 1875, indicated by affine motion vector 1877, is within region 1876, affine motion vector 1877 is not clipped. Similarly, the motion vector 1847 for sample 1842, associated with center position 1844 and region 1868, points to position 1845, which is within both region 1868 and boundary block 1810, and therefore affine motion vector 1847 is not clipped.
[0232] As mentioned above, regarding Figure 18A The described center position 1854 for sample 1852 is used to define the center motion vector 1850. Regardless of whether the center motion vector 1850 indicates a large or small motion, the boundary regions of all other samples for this current block 1860 (including samples 1832, 1842, 1872, and others) have associated clipping regions based on a center vector that can be translated (e.g., parallel vectors of the same size but different positions or intersecting the current block). In various examples, positional changes between different sample positions will match positional changes in associated regions (e.g., the difference between the positions of samples 1872 and 1832 will be the same as the difference between positions 1874 and 1834 and the same as the difference between regions 1876 and 1864). The relationship between the restricted regions and their corresponding samples is the "translationalization" of the affine motion described above.
[0233] Figure 18C This is a diagram illustrating various aspects of clipping using thresholds based on some examples. Figure 18C It shows the relationship with Figure 18B Similar examples, but with different affine motion vectors. In Figure 18C In the example, the affine vector 1836 used for sample 1832 is... Figure 18B The motion vectors are the same, but the motion vectors 1886 and 1896 used for the corresponding samples 1872 and 1842 are different. Similar to motion vector 1836, motion vector 1896 exceeds the allowable motion vector deviation, and therefore motion vector 1896 is processed to generate a clipped motion vector 1898, which points to position 1897 within boundary block 1810 and region 1868. Figure 18C In the example, the affine motion vector 1886 is clipped even though it points within the boundary block 1810. Because the affine motion vector 1886 exceeds the permissible variation indicated by region 1876, it is processed to generate a clipped motion vector 1888 pointing to a position 1887 within the region 1876 associated with the center position 1874 of sample 1872. In the example above, even though the motion vector 1886 points to a position within the boundary block 1810, the motion vector 1886 is clipped due to the clipping parameters. Applying such clipping parameters within the boundary region 1810 simplifies the clipping operation and provides efficient utilization of system resources.
[0234] In another example, only memory regions may be defined by boundary blocks (e.g., without restrictions on motion within the boundary blocks, such as the aforementioned restrictions on motion vector 1886 associated with region 1876). In such examples, memory boundary blocks (e.g., boundary regions 1660, 1760, or 1810) may be implemented, for example, with integer precision in the final x / y coordinates. Such examples would allow unrestricted affine motion vectors within the boundary blocks. In such examples, motion vectors 1836 and 1896 would be clipped, but motion vector 1886 would not be clipped because the reference data indicated by motion vector 1886 is within boundary block 1810 and will be stored in memory and available for affine decoding without the additional restrictions associated with region 1876. In such examples, additional computational resources can be used to construct the clipping, and performance can be improved at the cost of resources used to construct more complex clipping operations, while maintaining the memory bandwidth performance of other examples (e.g., having the same boundary region 1660, 1760, or 1810 constraints on the reference data, but without separate region constraints for each motion vector within a boundary block such as regions 1864, 1876, and 1868).
[0235] Examples of deviation limits are as follows (where in " <highlight>"and" <highlighted>Underlined text is used between symbols to emphasize the text (e.g., " <highlight>Highlighted text <highlighted> ”)): Deviation_A[5] = { 16, 80, 224, 512, 1088}; Deviation_B[5] = { 16, 96, 240, 528, 1104}; MinX = center_pos_x + (center_mv_hor - Deviation_A[log2(w) - 3])<<5; MinY = center_pos_y + (center_mv_ver - Deviation_A[log2(h) - 3])<<5; MaxX = center_pos_x + (center_mv_hor + Deviation_B[log2(w) - 3])<<5; MaxY = center_pos_y + (center_mv_ver + Deviation_B[log2(h) - 3])<<5; <highlight> xInt = Clip3(MinX, MaxX - 1, xInt) (8-734) yInt = Clip3( MinY, MaxY -1, yInt ) (8-735) <highlightend>
[0236] The following uses " <insert1>"and" <insertend1>Text separated by underlined text (e.g., ") <insert1> Added text <insertend1>The document shows an example of specification text that provides examples of implementations for such solutions:
[0237] 1.1.1.4 Interpolation process for enhanced interpolation filters The input to this process is: – The position (xCb, yCb) within the entire sample cell. – Two variables, cbWidth and cbHeight, specify the width and height of the current decoded block. – Horizontal change of the motion vector dX, – The vertical change of the motion vector dY, – Motion vector mvBaseScaled – The selected reference image sample array refPicLX, – Sample bit depth – The width of the image in the sample, pic_width. – The height of the image in the sample, pic_height. The output of this process is: – The array predSamplesLX, representing the predicted sample values (cbWidth) x (cbHeight).
[0238] The variables shift1, shift2, shift3, offset1, offset2, and offset3 are derived as follows: shift0 is set to equal bitDept – 6, and offset0 is equal to 2. shift1-1 , shift1 is set to 11, and offset1 is equal to 1024. <insert> The variables hor_max, ver_max, hor_min, and ver_min are obtained by calling the procedure specified in 0. The derivation is based on the position (xCb, yCb) in the full sample cell, and two parameters specifying the width and height of the current decoded block. Variables cbWidth and cbHeight, horizontal change of motion vector dX, vertical change of motion vector dY, motion vector mvBaseScaled, the width pic_width of the sample image, and the height pic_height of the sample image are used as input. The input is hor_max, and the output is hor_min, ver_max, hor_min, and ver_min. <insert>
[0239] For x = -1.. cbWidth and y = -1.. cbHeight, the following applies: –The motion vector mvX is derived as follows: mvX[ 0 ] = ( mvBaseScaled[ 0 ] + dX[ 0 ] x + dY[0] y ) (8-728) mvX[ 1 ] = ( mvBaseScaled[ 1 ] + dX[ 1 ] x + dY[ 1 ] y ) (8-729) <insert> mvX[ 0 ] = Clip3( hor_min, hor_max, mvX[ 0 ] ) (8-730) mvX[ 1 ] = Clip3(ver_min, ver_max, mvX[ 1 ] ) (8-731) <insertend> <insert> 1.1.1.5 Derivation of clipping parameters for affine motion vectors The input to this process is: – The position (xCb, yCb) within the entire sample cell. – Two variables, cbWidth and cbHeight, specify the width and height of the current decoded block. – The horizontal change dX of the motion vector – The vertical change dY of the motion vector – Motion vector mvBaseScaled, – The width of the image in the sample, pic_width. – The height of the image in the sample, pic_height. The output of this process is: – hor_max, ver_max, hor_min, and ver_min represent the maximum and minimum permissible horizontal and vertical motion vectors, respectively. Vertical component. The center motion vector mv_center is derived as follows: mv_center [ 0 ] = ( mvBaseScaled[ 0 ] + dX[ 0 ] (cbWidth>>1) + dY [ 0 ] (cbHeight>>1) (8-743) mv_center [ 1 ] = ( mvBaseScaled[ 1 ] + dX[ 1 ] (cbWidth>>1) + dY [ 1 ] (cbHeight>>1) (8-743) Call the rounding procedure for motion vectors as specified in Clause 8.5.3.10, where mv_center, the set... The motion vector is rounded and takes the rightShift (set to 5) and leftShift (set to 0) as input. The quantity is returned as mv_center. The motion vector mv_center is clipped as follows: mv_center [ 0 ] = Clip3( -2 17 , 2 17 - 1, mv_center [ 0 ] ) (8-686) mv_center [ 1 ] = Clip3( -2 17 , 2 17 - 1, mv_center [ 1 ] ) (8-686) The variables smv_hor_min, mv_ver_min, mv_hor_max, and mv_ver_max are derived as follows: mv_hor_min = mv_center [ 0 ] – deviationA [ log2CbWidth - 3 ] (8-743) mv_ver_min = mv_center [ 1 ] – deviationA [ log2CbHeight - 3 ] (8- 743) mv_hor_max = mv_center [ 0 ] + deviationB [ log2CbWidth - 3 ] (8- 743) mv_ver_max = mv_center [ 1 ] + deviationB [ log2CbHeight – 3 ] (8- 743) For k=0..4, deviationA and deviationB are specified as: deviationA[k] = { 16, 80, 224, 512, 1088}, deviationB[k] { 16, 96, 240, 528, 1104}. The variables hor_max_pic, ver_max_pic, hor_min_pic, and ver_min_pic are derived as follows: hor_max_pic = (pic_width + 128 - xCb - cbWidth + 1)<<4 (8-743) ver_max_pic = (pic_height + 128 - yCb - cbHeight + 1)<<4 (8-743) hor_min_pic = (-128 - xCb)<<4 (8-743) ver_min_pic = (-128 - yCb)<<4 (8-743) The following derivation represents the outputs hor_max and ver_ of the maximum and minimum permissible horizontal and vertical components of the motion vector. max, hor_min and ver_min: hor_max = min(hor_max_pic, mv_hor_max)<<5 (8-743) ver_max = min(ver_max_pic, mv_ver_max)<<5 (8-743) hor_min = max(hor_min_pic, mv_hor_min)<<5 (8-743) ver_min = max(ver_min_pic, mv_ver_min)<<5 (8-743) <insert>
[0240] Figure 19 This is a flowchart illustrating an affine decoding process 1900 with clipping parameters according to the examples described herein. In some examples, process 1900 may be executed by encoding device 104 or decoding device 112. In some examples, process 1900 may be implemented as instructions in a computer-readable storage medium that, when executed by the processing circuitry of the device, cause the device to perform the operations of process 1900.
[0241] At box 1902, procedure 1900 includes the following operation: obtaining the current decoding block from video data. This operation can be part of a sequential operation processing multiple decoding blocks, where clipping parameters are determined for each block and used for each sample of the current block. When a block is decoded and the operation moves to the next block, a new set of clipping parameters can be determined for the new block and used for each sample of the new block. In some examples of procedure 1900, control data includes values from a derivation table.
[0242] At box 1904, procedure 1900 includes the following operation: determining control data for the current decoded block. In some examples, the control data may include the above-mentioned input consisting of: the position (xCb, yCb) in the full sample cell; two variables cbWidth and cbHeight specifying the width and height of the current decoded block; the horizontal variation dX of the motion vector; the vertical variation dY of the motion vector; the motion vector mvBaseScaled; the width pic_width and the height pic_height of the image in the sample. In other examples, other combinations or groupings of data may be used. In another example, the control data includes: the position having associated horizontal and associated vertical coordinates in the full sample cell; the width variable specifying the width of the current decoded block; the height variable specifying the height of the current decoded block; the horizontal variation of the motion vector; the vertical variation of the motion vector; the base scaled motion vector; the height of the image associated with the current decoded block in the sample; and the width of the image in the sample.
[0243] At box 1906, procedure 1900 includes the following operation: determining one or more affine motion vector clipping parameters based on control data. In some examples, the affine motion vector clipping parameters include: a horizontal maximum variable; a horizontal minimum variable; a vertical maximum variable; and a vertical minimum variable.
[0244] In some examples, the horizontal minimum variable is defined by the maximum value selected from the horizontal minimum image value and the horizontal minimum motion vector value. In some such examples, the horizontal minimum variable (hor_min) is defined by the maximum value selected from the horizontal minimum image value (hor_min_pic) and the horizontal minimum motion vector value (mv_hor_min): hor_min = max(hor_min_pic, mv_hor_min).
[0245] In some of these examples, the minimum horizontal image value (hor_min_pic) is determined based on the associated horizontal coordinates. In some of these examples, hor_min_pic is defined as: hor_min_pic = (-128 - xCb).
[0246] In some examples, the minimum horizontal motion vector value is determined based on an array of values of the center motion vector value, a resolution value or block size associated with the video data (e.g., current decoded block width x height), and a width variable specifying the width of the current decoded block. In some such examples, mv_hor_min is defined as: mv_hor_min = mv_center [0] – deviationA [log2CbWidth - 3]; where mv_center [0] is the center motion vector value, deviationA is an array of values of a resolution value or block size associated with the video data (e.g., current decoded block width x height), and cbWidth is a width variable specifying the width of the current decoded block.
[0247] In some examples, the center motion vector value is determined based on the base scaling motion vector, the horizontal variation of the motion vector, the width variable, and the height variable. In some such examples, the center motion vector value is defined as: mv_center [ 0 ] = ( mvBaseScaled[ 0 ] + dX[ 0 ] (cbWidth>>1) + dY[0] (cbHeight>>1) ).
[0248] In some examples, the base scaling motion vector corresponds to the top-left corner of the current decoded block and is determined based on the control point motion vector value. In some examples, mvBaseScaled corresponds to the top-left corner of the luminance decoded block and is defined as: mvBaseScaled[0]=cpMvLX[0][0]<<7; mvBaseScaled[1]=cpMvLX[0][1]<<7; where cpMvLX is the control point motion vector.
[0249] The above aspects of box 1906 primarily describe the operations used to determine the parameters associated with the horizontal minimum variable (hor_min). Each of the other combinations of horizontal, vertical, maximum, and minimum parameters used for vector clipping can have similar examples as described herein, including elements for: the horizontal maximum variable; the vertical maximum variable; and the vertical minimum variable.
[0250] In some examples, the maximum horizontal variable is defined by the minimum value selected from the maximum horizontal picture value and the maximum horizontal motion vector value. In some examples, the maximum horizontal picture value is determined based on the picture width, the associated horizontal coordinates, and the width variable. In some examples, the maximum horizontal motion vector value is determined based on the center motion vector value, an array of values based on the resolution value or block size associated with the video data (e.g., current decoded block width x height), and a width variable specifying the width of the current decoded block. In some examples, the center motion vector value is determined based on the base scaling motion vector, the horizontal variation of the motion vector, the width variable, and the height variable. In some examples, the base scaling motion vector corresponds to the corner of the current decoded block and is determined based on the control point motion vector value.
[0251] In some examples, the maximum vertical variable is defined by the minimum value selected from the maximum vertical image value and the maximum vertical motion vector value. In some examples, the maximum vertical image value is determined based on the image height, the associated vertical coordinates, and a height variable. In some examples, the maximum vertical motion vector value is determined based on an array of values for the center motion vector value, a resolution value or block size associated with the video data (e.g., current decoded block width x height), and a height variable specifying the height of the current decoded block.
[0252] In some examples, the minimum vertical variable is defined by the maximum value selected from the minimum vertical picture value and the minimum vertical motion vector value. In some examples, the minimum vertical picture value is determined based on the associated vertical coordinates. In some examples, the minimum vertical motion vector value is determined based on an array of values of the center motion vector value, a resolution value or block size (e.g., current decoded block width x height) associated with the video data, and a height variable specifying the height of the current decoded block.
[0253] As part of box 1906, additional parameter-specific derivations can be performed, including derivations that include elements for: the horizontal maximum variable; the vertical maximum variable; and the vertical minimum variable. In some examples, these variables can be determined based on the details described herein, including: hor_max = min(hor_max_pic, mv_hor_max)<<5; ver_max = min(ver_max_pic, mv_ver_max)<<5; hor_min = max(hor_min_pic, mv_hor_min)<<5; ver_min = max(ver_min_pic, mv_ver_min)<<5; mv_center [ 0 ] = ( mvBaseScaled[ 0 ] + dX[ 0 ] (cbWidth>>1) + dY[0] (cbHeight>>1) ); mv_center [ 1 ] = ( mvBaseScaled[ 1 ] + dX[ 1 ] (cbWidth>>1) + dY[1] (cbHeight>>1) ); mv_center[0] = Clip3(-2) 17 , 2 17 - 1, mv_center [ 0 ] ); mv_center [1] = Clip3(-2 17 , 2 17 - 1, mv_center [ 1 ] ); mv_hor_min = mv_center [ 0 ] – deviationA [ log2CbWidth - 3 ]; mv_ver_min = mv_center [ 1 ] – deviationA [ log2CbHeight - 3 ]; mv_hor_max = mv_center [ 0 ] + deviationB [ log2CbWidth - 3 ]; mv_ver_max = mv_center [ 1 ] + deviationB [ log2CbHeight – 3 ]; For k=0..4, deviationA and deviationB are specified as deviationA[k] = { 16,80, 224, 512, 1088} and deviationB[k]{ 16, 96, 240, 528, 1104}; hor_max_pic = (pic_width + 128 - xCb - cbWidth + 1)<<4; ver_max_pic = (pic_height + 128 - yCb - cbHeight + 1)<<4; hor_min_pic = (-128 - xCb)<<4; ver_min_pic = (-128 - yCb)<<4; And other such details described herein. In other examples, other similar procedures may be used to determine the parameters used for cropping.
[0254] At box 1908, process 1900 includes the following operation: selecting samples for the current decoding block. As mentioned above, either a selection number of samples for the current block can be used, or every sample of the current block can be used. Example affine prediction based on EVC can be implemented using different methods. One example EVC method utilizes translational motion prediction for sub-blocks. Another example of EVC affine prediction uses finer-grained y-motion prediction (e.g., by pixel). Different methods can have associated operations for selecting samples.
[0255] At box 1910, procedure 1900 includes the following operation: determining the affine motion vector of the sample for the current decoding block. In some examples, the affine motion vector for the sample for the current decoding block is determined based on the following: a first base scaled motion vector value, a first horizontal change of the motion vector value, a first vertical change of the motion vector value, a second base scaled motion vector value, a second horizontal change of the motion vector value, a second vertical change of the motion vector value, the horizontal coordinate of the sample, and the vertical coordinate of the sample. In some examples, the motion vector specified as mvX can be derived using the following equation: mvX[0] = (mvBaseScaled[0] + dX[0]) x + dY[0] y );mvX[ 1 ] = (mvBaseScaled[ 1 ] + dX[ 1 ] x + dY[ 1 ] y.
[0256] At box 1912, procedure 1900 includes the following operation: clipping an affine motion vector using one or more affine motion vector clipping parameters to generate a clipped affine motion vector. In some examples, the affine motion vector is clipped according to the following equations: mvX[0] = Clip3(hor_min, hor_max, mvX[0]); and mvX[1] = Clip3(ver_min, ver_max, mvX[1]).
[0257] In addition to the boxes mentioned above, some elements of process 1900 may also include additional operations, intermediate operations, or repetitions of operations from certain boxes. In some examples, such additional operations may include: identifying a reference picture associated with the current decoding block; and storing a portion of the reference picture defined by the affine motion vector clipping parameters. Some of these operations may take effect when this portion of the reference picture is stored in a memory buffer for use in affine motion processing operations using the current decoding block.
[0258] Similarly, some repetitive operations may include: sequentially obtaining multiple current decoded blocks from video data; determining a set of affine motion vector clipping parameters on a per-block basis for each block in the multiple current decoded blocks; and retrieving a portion of a corresponding reference image on a per-block basis using the set of affine motion vector clipping parameters for each of the multiple current decoded blocks. In any such example, the operation may also include: processing the current block using reference image data derived from a reference image indicated by clipped affine motion vectors. Such a block may be a luminance decoded block or any other such block of video data used for decoding in an affine decoding mode. Such a process 1900 may be performed by any device herein, including devices having memory and one or more processors. Such devices may include devices having a display device coupled to one or more processors and configured to display an image from the video data; and one or more wireless interfaces coupled to one or more processors, the one or more wireless interfaces including one or more baseband processors and one or more transceivers. Other such devices may include other components described herein.
[0259] In some examples, the processes described herein may be performed by a computing device or apparatus, such as encoding device 104, decoding device 112, and / or any other computing device. In some cases, the computing device or apparatus may include a processor, microprocessor, microcomputer, or other components of a device configured to perform the steps of the processes described herein. In some examples, the computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) comprising video frames. For example, the computing device may include a camera device, which may or may not include a video codec. As another example, the computing device may include a mobile device with a camera (e.g., a camera device such as a digital camera, IP camera, etc., a mobile phone or tablet device including a camera, or other types of devices with a camera). In some cases, the computing device may include a display for displaying images. In some examples, the camera or other capturing device that captures video data is separate from the computing device, in which case the computing device receives the captured video data. The computing device may also include a network interface, transceiver, and / or transmitter configured to transmit video data. The network interface, transceiver, and / or transmitter may be configured to transmit Internet Protocol (IP) based data or other network data.
[0260] The processes described herein can be implemented using hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation represents a computer-executable instruction stored on one or more computer-readable storage media, which, when executed by one or more processors, performs the described operation. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.
[0261] Furthermore, the processes described herein can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or a combination thereof. As noted above, the code can be stored, for example, on a computer-readable or machine-readable storage medium in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0262] The decoding techniques discussed herein can be implemented in an example video encoding and decoding system (e.g., System 100). In some examples, the system includes a source device that provides encoded video data to be decoded later by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source and destination devices can include any of a variety of devices, including desktop computers, laptops, tablets, set-top boxes, mobile phones (such as so-called "smartphones"), so-called "smartboards," televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source and destination devices can be equipped for wireless communication.
[0263] A destination device can receive encoded video data to be decoded via a computer-readable medium. The computer-readable medium can include any type of medium or device capable of moving encoded video data from a source device to a destination device. In one example, the computer-readable medium can include a communication medium enabling the source device to transmit encoded video data directly to the destination device in real time. The encoded video data can be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network (such as the Internet). The communication medium can include a router, a switch, a base station, or any other means that can be used to facilitate communication from the source device to the destination device.
[0264] In some examples, encoded data can be output from an output interface to a storage device. Similarly, encoded data can be accessed from a storage device via an input interface. The storage device can include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In other examples, the storage device can correspond to a file server or another intermediate storage device capable of storing encoded video generated by a source device. The destination device can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing and sending encoded video data to the destination device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device can access the encoded video data via any standard data connection, including an internet connection. The connection can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The transmission of encoded video data from a storage device can be streaming, downloading, or a combination thereof.
[0265] The technology disclosed herein is not necessarily limited to wireless applications or setups. The technology can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video (such as HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding digital video stored on data storage media, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0266] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device may include an input interface, a video decoder, and a display device. The video encoder of the source device may be configured to apply the techniques disclosed herein. In other examples, the source and destination devices may include other components or arrangements. For example, the source device may receive video data from an external video source such as an external camera. Similarly, the destination device may interface with an external display device, rather than including an integrated display device.
[0267] The example system described above is merely an example. Techniques for processing video data in parallel can be implemented by any digital video encoding and / or decoding device. While the techniques of this disclosure are typically implemented by video encoding devices, they can also be implemented by a video encoder / decoder, commonly referred to as a "CODEC". Furthermore, the techniques of this disclosure can also be implemented by a video preprocessor. The source and destination devices are merely examples of decoding devices in which the source device generates decoded video data for transmission to the destination device. In some examples, the source and destination devices can operate in a substantially symmetrical manner, such that each of these devices includes both video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0268] Video sources may include video capture devices such as cameras, video archiving units containing previously captured video, and / or video feed interfaces for receiving video from video content providers. Alternatively, a video source may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In some cases, if the video source is a camera, the source and destination devices may form a so-called camera phone or video phone. However, as mentioned above, the techniques described in this disclosure are generally applicable to video decoding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by a video encoder. The encoded video information can be output to a computer-readable medium via an output interface.
[0269] As mentioned, computer-readable media can include transient media (such as wireless broadcasting or wired network transmission), or storage media (i.e., non-transient storage media) (such as hard disks, flash drives, compressed optical discs, digital versatile optical discs, Blu-ray discs), or other computer-readable media. In some examples, a network server (not shown) can, for example, receive encoded video data from a source device via a network transmission and provide the encoded video data to a destination device. Similarly, a computing device in a media production facility (such as an optical disc stamping facility) can receive encoded video data from a source device and manufacture an optical disc containing the encoded video data. Therefore, in the various examples, computer-readable media can be understood to include one or more computer-readable media of various forms.
[0270] The input interface of the destination device receives information from a computer-readable medium. The information in the computer-readable medium may include grammatical information defined by the video encoder, which is also used by the video decoder. The grammatical information includes grammatical elements describing the characteristics of blocks and other decoding units (e.g., groups of pictures (GOPs)) and / or processing thereof. The display device displays the decoded video data to the user and may include any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device. Various embodiments of this application have been described.
[0271] exist Figure 20 and Figure 21 The specific details of the encoding device 104 and the decoding device 112 are shown in the figure. Figure 20 This is a block diagram illustrating an example encoding device 104 that can implement one or more of the techniques described in this disclosure. Encoding device 104 can, for example, generate the syntax structures described herein (e.g., syntax structures of VPS, SPS, PPS, or other syntax elements). Encoding device 104 can perform intra-frame prediction and inter-frame prediction decoding of video blocks within a video slice. As previously described, intra-frame decoding relies at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter-frame decoding relies at least in part on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames of a video sequence. Intra-frame mode (I-mode) can refer to any of several spatial-based compression modes. Inter-frame modes such as one-way prediction (P-mode) or two-way prediction (B-mode) can refer to any of several temporal-based compression modes.
[0272] Encoding device 104 includes a segmentation unit 35, a prediction processing unit 41, a filter unit 63, an image memory 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-frame prediction processing unit 46. For video block reconstruction, encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62. Filter unit 63 is intended to represent one or more loop filters, such as deblocking filters, adaptive loop filters (ALF), and sample adaptive offset (SAO) filters. Although in Figure 20 The filter unit 63 is shown as an in-loop filter, but in other configurations, it may be implemented as a post-loop filter. The post-processing device 57 may perform additional processing on the encoded video data generated by the encoding device 104. In some instances, the techniques of this disclosure may be implemented by the encoding device 104. However, in other instances, one or more of the techniques of this disclosure may be implemented by the post-processing device 57.
[0273] like Figure 20 As shown, encoding device 104 receives video data, and segmentation unit 35 segments the data into video blocks. Segmentation may also include, for example, segmentation into slices, segments, tiles, or other larger units based on the quadtree structure of LCUs and CUs, as well as video block segmentation. Encoding device 104 generally illustrates components for encoding video blocks within video slices to be encoded. Slices may be divided into multiple video blocks (and may be divided into a set of video blocks referred to as tiles). Prediction processing unit 41 may select one of several possible decoding modes for the current video block based on error results (e.g., decoding rate and distortion level, etc.), such as one of several intra-frame prediction decoding modes or one of several inter-frame prediction decoding modes. Prediction processing unit 41 may provide the obtained intra-frame decoded or inter-frame decoded blocks to summer 50 to generate residual block data, and to summer 62 to reconstruct the encoded blocks for use as reference pictures.
[0274] The intra-prediction processing unit 46 within the prediction processing unit 41 can perform intra-prediction decoding of the current video block relative to one or more adjacent blocks in the same frame or slice as the current block to be decoded, to provide spatial compression. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-prediction decoding of the current video block relative to one or more prediction blocks in one or more reference pictures, to provide temporal compression.
[0275] Motion estimation unit 42 can be configured to determine an inter-frame prediction mode for video slices based on a predetermined mode for the video sequence. The predetermined mode can designate video slices in the sequence as P-slices, B-slices, or GPB-slices. Motion estimation unit 42 and motion compensation unit 44 can be highly integrated, but are shown separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is a process of generating motion vectors that estimate the motion for video blocks. Motion vectors can, for example, indicate the displacement of a prediction unit (PU) of a video block within the current video frame or picture relative to a prediction block within a reference picture.
[0276] A predicted block is a block found to closely match the PU of the video block to be decoded in terms of pixel difference, which can be determined by the sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some examples, the encoding device 104 can compute values for pixel positions less than integers of a reference image stored in the image memory 64. For example, the encoding device 104 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, the motion estimation unit 42 can perform motion search relative to full-pixel positions and fractional pixel positions, and output a motion vector with fractional pixel precision.
[0277] The motion estimation unit 42 calculates the motion vector for the PU by comparing the position of the PU of the video block in the inter-frame decoded slice with the position of the predicted block of the reference picture. The reference picture may be selected from a first reference picture list (list 0) or a second reference picture list (list 1), where each reference picture list identifies one or more reference pictures stored in the picture memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.
[0278] Motion compensation performed by motion compensation unit 44 may involve retrieving or generating prediction blocks based on motion vectors determined by motion estimation, and performing interpolation, if possible, with subpixel precision. Upon receiving the motion vector of the PU for the current video block, motion compensation unit 44 can locate the prediction block pointed to by the motion vector in a list of reference images. Encoding device 104 forms a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being decoded. The pixel difference forms residual data for that block and may include both luminance difference components and chrominance difference components. Summer 50 represents one or more components performing the subtraction operation. Motion compensation unit 44 may also generate syntax elements associated with video blocks and video slices for use by decoding device 112 when decoding video blocks of video slices.
[0279] As described above, the intra-prediction processing unit 46 can perform intra-prediction on the current block as an alternative to the inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44. Specifically, the intra-prediction processing unit 46 can determine the intra-prediction mode to be used for encoding the current block. In some examples, the intra-prediction processing unit 46 can use various intra-prediction modes to encode the current block, for example, during a separate encoding path, and the intra-prediction processing unit 46 can select a suitable intra-prediction mode from the tested modes. For example, the intra-prediction processing unit 46 can use rate-distortion analysis to calculate rate-distortion values for various tested intra-prediction modes, and can select the intra-prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis typically determines the amount of distortion (or error) between the encoded block and the original uncoded block (which is encoded to produce the encoded block), and the bit rate (i.e., the number of bits) used to produce the encoded block. Intra-prediction processing unit 46 can calculate the ratio based on the distortion and rate for various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for that block.
[0280] In any case, after selecting an intra-prediction mode for a block, the intra-prediction processing unit 46 can provide information indicating the selected intra-prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 can encode the information indicating the selected intra-prediction mode. The encoding device 104 can include in the transmitted bitstream configuration data definitions for encoding contexts for various blocks, as well as indications of the most probable intra-prediction mode to be used for each context, an intra-prediction mode index table, and a modified intra-prediction mode index table. The bitstream configuration data may include multiple intra-prediction mode index tables and multiple modified intra-prediction mode index tables (also referred to as codeword maps).
[0281] After prediction processing unit 41 generates a prediction block for the current video block via inter-frame prediction or intra-frame prediction, encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to transform processing unit 52. Transform processing unit 52 uses a transform (such as Discrete Cosine Transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients. Transform processing unit 52 can transform the residual video data from the pixel domain to the transform domain (such as the frequency domain).
[0282] The transform processing unit 52 can send the obtained transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can perform a scan over a matrix that includes the quantized transform coefficients. Alternatively, the entropy coding unit 56 can perform the scan.
[0283] After quantization, entropy coding unit 56 performs entropy coding on the quantized transform coefficients. For example, entropy coding unit 56 can perform context-adaptive variable-length decoding (CAVLC), context-adaptive binary arithmetic decoding (CABAC), syntax-based context-adaptive binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, or another entropy coding technique. After entropy coding by entropy coding unit 56, the encoded bitstream can be sent to decoding device 112, or archived for later transmission or retrieved by decoding device 112. Entropy coding unit 56 can also perform entropy coding on motion vectors and other syntax elements for the current video slice being decoded.
[0284] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual block in the pixel domain for later use as a reference block for reference images. Motion compensation unit 44 calculates the reference block by adding the residual block to a predicted block of one of the reference images in the reference image list. Motion compensation unit 44 can also apply one or more interpolation filters to the reconstructed residual block to calculate pixel values below the integer level for motion estimation. Summer 62 adds the reconstructed residual block to the motion-compensated predicted block generated by motion compensation unit 44 to produce a reference block for storage in image memory 64. This reference block can be used by motion estimation unit 42 and motion compensation unit 44 as a reference block for inter-frame prediction of blocks in subsequent video frames or images.
[0285] Encoding device 104 can perform any of the techniques described herein. Some techniques of this disclosure have been generally described with respect to encoding device 104, but as mentioned above, some of the techniques of this disclosure can also be implemented by post-processing device 57.
[0286] Figure 20 Encoding device 104 represents an example of a video encoder configured to perform one or more of the transform decoding techniques described herein. Encoding device 104 can perform any of the techniques described herein, including those mentioned above. Figure 21 The process described.
[0287] Figure 21 This is a block diagram illustrating an example decoding device 112. Decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, a filter unit 91, and an image memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-frame prediction processing unit 84. In some examples, decoding device 112 can perform operations generally related to... Figure 20 The encoding path described by the encoding device 104 is the opposite of the decoding path.
[0288] During the decoding process, decoding device 112 receives an encoded video bitstream representing video blocks and associated syntax elements of an encoded video slice, sent by encoding device 104. In some embodiments, decoding device 112 may receive the encoded video bitstream from encoding device 104. In some embodiments, decoding device 112 may receive the encoded video bitstream from network entity 79 (such as a server, a media-aware network element (MANE), a video editor / stitcher, or other such device configured to implement one or more of the techniques described above). Network entity 79 may or may not include encoding device 104. Some of the techniques described in this disclosure may be implemented by network entity 79 before network entity 79 sends the encoded video bitstream to decoding device 112. In some video decoding systems, network entity 79 and decoding device 112 may be part of a single device, while in other instances, the functionality described with respect to network entity 79 may be performed by the same device including decoding device 112.
[0289] The entropy decoding unit 80 of the decoding device 112 performs entropy decoding on the bitstream to generate quantized coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 can receive syntax elements at the video slice level and / or video block level. The entropy decoding unit 80 can process and parse both fixed-length and variable-length syntax elements in one or more parameter sets (such as VPS, SPS, and PPS).
[0290] When a video slice is decoded into an intra-frame decoded (I) slice, the intra-frame prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video slice based on the intra-frame prediction mode notified by a signal and data from previously decoded blocks from the current frame or picture. When a video frame is decoded into an inter-frame decoded (i.e., B, P, or GPB) slice, the motion compensation unit 82 of the prediction processing unit 81 generates a prediction block for the video block of the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit 80. A prediction block can be generated based on one of the reference pictures in the reference picture list. The decoding device 112 can construct the reference frame list (list 0 and list 1) using a default construction technique based on the reference pictures stored in the picture memory 92.
[0291] Motion compensation unit 82 determines prediction information for video blocks used in the current video slice by parsing motion vectors and other syntax elements, and uses this prediction information to generate prediction blocks for the current video slice being decoded. For example, motion compensation unit 82 may use one or more syntax elements in the parameter set to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) for decoding video blocks in the video slice, the inter-frame prediction slice type (e.g., B-slice, P-slice, or GPB-slice), construction information for one or more reference picture lists for the slice, motion vectors for each inter-frame encoded video block in the slice, inter-frame prediction state for each inter-frame decoded video block in the slice, and other information for decoding video blocks in the current video slice.
[0292] The motion compensation unit 82 can also perform interpolation based on an interpolation filter. The motion compensation unit 82 can use an interpolation filter, such as that used by the encoding device 104 during the encoding of a video block, to calculate interpolated values for pixels less than an integer value in the reference block. In this case, the motion compensation unit 82 can determine the interpolation filter used by the encoding device 104 based on the received syntax elements, and can use the interpolation filter to generate a prediction block.
[0293] The inverse quantization unit 86 inverse-quantizes or dequantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include determining the degree of quantization using quantization parameters calculated by the encoding device 104 for each video block in the video slice, and similarly determining the degree of inverse quantization to be applied. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT or other suitable inverse transform), inverse integer transform, or conceptually similar inverse transform to the transform coefficients to produce residual blocks in the pixel domain.
[0294] After the motion compensation unit 82 generates a prediction block for the current video block based on motion vectors and other syntax elements, the decoding device 112 forms a decoded video block by summing the residual block from the inverse transform processing unit 88 with the corresponding prediction block generated by the motion compensation unit 82. The summer 90 represents one or more components performing this summation operation. If needed, loop filters (in or after the decoding loop) can also be used to smooth pixel transitions or otherwise improve video quality. The filter unit 91 is intended to represent one or more loop filters, such as deblocking filters, adaptive loop filters (ALF), and sample adaptive offset (SAO) filters. Although the filter unit 91 is in Figure 21 The filter unit 91 is shown as a filter in the loop, but in other configurations, it can be implemented as a post-loop filter. Decoded video blocks in a given frame or image are stored in image memory 92, which stores reference images for subsequent motion compensation. Image memory 92 also stores the decoded video for later display on a display device (such as a monitor). Figure 1 The video is displayed on the destination device 122 shown in the figure.
[0295] Figure 21 Decoding device 112 represents an example of a video decoder configured to perform one or more of the transform decoding techniques described herein. Decoding device 112 can perform any of the techniques described herein, including those mentioned above. Figure 21 The process described in 1900.
[0296] In the foregoing description, aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that the subject matter of this application is not limited thereto. Therefore, although illustrative embodiments of this application have been described in detail herein, it is to be understood that the inventive concept may be embodied and employed in other ways, and the appended claims are intended to be construed as including such variations, except those limited by the prior art. Various features and aspects of the foregoing subject matter may be used individually or collectively. Furthermore, embodiments may be used in any number of environments and applications other than those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings are to be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in a different order than that described.
[0297] Those skilled in the art will understand that, without departing from the scope of this specification, the symbols or terms less than ("<") and greater than (">") used herein can be replaced with less than or equal to ("<"), respectively. ") and greater than or equal to (" ")symbol.
[0298] When a component is described as being "configured" to perform certain operations, such configuration can be achieved, for example, by designing a circuit or other hardware to perform the operation, programming a programmable circuit (e.g., a microprocessor or other suitable circuit) to perform the operation, or any combination thereof.
[0299] The language of a claim that states "at least one" and / or "one or more" in a set, or other languages, indicates that one or more members of the set (in any combination) satisfy the claim. For example, the language of a claim stating "at least one of A and B" means A, B, or A and B. In another example, the language of a claim stating "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language of "at least one" and / or "one or more" in a set does not limit the set to items listed in the set. For example, the language of a claim stating "at least one of A and B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0300] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been generally described above in relation to their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in alternative ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.
[0301] The techniques described herein can also be implemented using electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, mobile phones with wireless communication capabilities, or integrated circuit devices with multiple uses, including applications in mobile phones with wireless communication capabilities and other devices. Any feature described as a module or component can be implemented together in an integrated logical device or implemented separately as discrete but interoperable logical devices. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code, which includes instructions for performing one or more of the methods described above when executed. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Alternatively or concurrently, the technology may be implemented at least in part by a computer-readable communication medium (such as a propagating signal or wave) that carries or transmits program code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer.
[0302] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated software or hardware modules configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).
[0303] Illustrative examples of this disclosure include:
[0304] Example 1: A method for processing video data, the method comprising: obtaining one or more blocks of video data; and determining an affine motion vector to be used to predict the blocks of video data, wherein a region of at least a reference image accessible using the affine motion vector is constrained based on constraints.
[0305] Example 2: The method described in Example 1, wherein the constraint is based on the size of the block.
[0306] Example 3, the method according to any one of Examples 1 to 2, further includes: trimming the affine motion vector according to the constraint.
[0307] Example 4, the method according to any one of Examples 1 to 2, further includes: cropping reference sample coordinates of at least one sample from the at least one reference image according to the constraints, the reference sample coordinates being determined using the affine motion vector.
[0308] Example 5, the method according to any one of Examples 1 to 4, further includes: deriving clipping parameters as a function of the size of the block.
[0309] Example 6, the method according to any one of Examples 1 to 5, further includes: calculating parameters for cutting at least one of the affine motion vector and reference sample coordinates on a per-block basis, according to one or more table-formed parameters.
[0310] Example 7: An apparatus comprising a memory configured to store video data and a processor configured to process the video data according to any one of Examples 1 to 6.
[0311] Example 8, the apparatus according to Example 7, wherein the apparatus includes a decoder.
[0312] Example 9. The apparatus according to Example 7, wherein the apparatus includes an encoder.
[0313] Example 10: The apparatus according to any one of Examples 7 to 9, wherein the apparatus is a mobile device.
[0314] Example 11, the apparatus according to any one of Examples 7 to 10, further includes a display configured to display the video data.
[0315] Example 12, the apparatus according to any one of Examples 7 to 11, further includes a camera configured to capture one or more images.
[0316] Example 13: A computer-readable medium having instructions stored thereon, which, when executed by a processor, perform the method according to any one of Examples 1 to 6.
[0317] Example 14. An apparatus for decoding video data, the apparatus comprising: a memory; and one or more processors coupled to the memory, the one or more processors being configured to: obtain a current decoding block from the video data; determine control data for the current decoding block; determine one or more affine motion vector clipping parameters based on the control data; select a sample of the current decoding block; determine an affine motion vector for the sample of the current decoding block; and clip the affine motion vector using the one or more affine motion vector clipping parameters to generate a clipped affine motion vector.
[0318] Example 15, the apparatus according to Example 14, wherein the control data includes: a position having associated horizontal coordinates and associated vertical coordinates in the full sample cell; a width variable specifying the width of the current decoded block; a height variable specifying the height of the current decoded block; a horizontal change of the motion vector; a vertical change of the motion vector; a base scaling motion vector; the height of the image in the sample associated with the current decoded block; and the width of the image in the sample.
[0319] Example 16, the apparatus according to Example 15, wherein the one or more affine motion vector clipping parameters include: a horizontal maximum variable; a horizontal minimum variable; a vertical maximum variable; and a vertical minimum variable.
[0320] Example 17, the apparatus according to Example 16, wherein the minimum horizontal variable is defined by the maximum value selected from the minimum horizontal image value and the minimum horizontal motion vector value.
[0321] Example 18, the apparatus according to Example 17, wherein the minimum horizontal image value is determined based on the associated horizontal coordinates.
[0322] Example 19, the apparatus according to Example 18, wherein the minimum horizontal motion vector value is determined based on a center motion vector value, an array of values based on resolution values or block size (e.g., current decoded block width x height) associated with the video data, and a width variable specifying the width of the current decoded block.
[0323] Example 20, the apparatus according to Example 19, wherein the center motion vector value is determined based on the base scaling motion vector, the horizontal variation of the motion vector, the width variable, and the height variable.
[0324] Example 21, the apparatus according to Example 20, wherein the base scaling motion vector corresponds to the top left corner of the current decoding block and is determined based on the control point motion vector value.
[0325] Example 22: The apparatus according to Examples 16-21 above, wherein the maximum horizontal variable is defined by the minimum value selected from the maximum horizontal image value and the maximum horizontal motion vector value.
[0326] Example 23, the apparatus according to Example 22, wherein the maximum horizontal image value is determined based on the width of the image, the associated horizontal coordinates, and the width variable.
[0327] Example 24, the apparatus according to Example 23, wherein the maximum horizontal motion vector value is determined based on a center motion vector value, an array of values based on resolution values or block size (e.g., current decoded block width x height) associated with the video data, and a width variable specifying the width of the current decoded block.
[0328] Example 25, the apparatus according to Example 24, wherein the center motion vector value is determined based on the base scaling motion vector, the horizontal variation of the motion vector, the width variable, and the height variable.
[0329] Example 26, the apparatus according to Example 25, wherein the base scaling motion vector corresponds to the angle of the current decoded block and is determined based on the control point motion vector value.
[0330] Example 27: The apparatus according to Examples 16-26 above, wherein the maximum vertical variable is defined by the minimum value selected from the maximum vertical image value and the maximum vertical motion vector value.
[0331] Example 28, the apparatus according to Example 27, wherein the maximum vertical image value is determined based on the height of the image, the associated vertical coordinates, and the height variable.
[0332] Example 29, the apparatus according to Example 28, wherein the maximum vertical motion vector value is determined based on a center motion vector value, an array of values based on resolution values or block size (e.g., current decoded block width x height) associated with the video data, and a height variable specifying the width of the current decoded block.
[0333] Example 30: The apparatus according to Examples 16-30 above, wherein the minimum vertical variable is defined by the maximum value selected from the minimum vertical image value and the minimum vertical motion vector value.
[0334] Example 31, the apparatus according to Example 30, wherein the minimum vertical image value is determined based on the associated vertical coordinates.
[0335] Example 32, the apparatus according to Example 31, wherein the minimum vertical motion vector value is determined based on a center motion vector value, an array of values based on resolution values or block size (e.g., current decoded block width x height) associated with the video data, and a height variable specifying the height of the current decoded block.
[0336] Example 33, the apparatus according to Examples 14-32, wherein the one or more processors are configured to: sequentially obtain a plurality of current decoded blocks from the video data; determine a set of affine motion vector clipping parameters on a per-block basis for each block of the plurality of current decoded blocks; and retrieve a portion of a corresponding reference image on a per-block basis using the set of affine motion vector clipping parameters for each of the plurality of current decoded blocks.
[0337] Example 34, the apparatus according to Examples 14-33, wherein the one or more processors are configured to: identify a reference image associated with the current decoding block; and store a portion of the reference image defined by the one or more affine motion vector clipping parameters.
[0338] Example 35, the apparatus according to Example 34, further includes a memory buffer coupled to the one or more processors, wherein the portion of the reference image is stored in the memory buffer for use in affine motion processing operations with the current decoding block.
[0339] Example 36, the apparatus according to Examples 14-35, wherein the one or more processors are configured to process the current decoding block using reference image data derived from the reference image indicated by the cropped affine motion vector.
[0340] Example 37. The apparatus according to Examples 14-36, wherein the affine motion vector for the sample of the current decoded block is determined based on the following: a first base scaling motion vector value, a first horizontal change of the motion vector value, a first vertical change of the motion vector value, a second base scaling motion vector value, a second horizontal change of the motion vector value, a second vertical change of the motion vector value, the horizontal coordinate of the sample, and the vertical coordinate of the sample.
[0341] Example 38, the apparatus according to Examples 14-37, wherein the control data includes values from a derivation table.
[0342] Example 39: The apparatus according to Examples 14-38, wherein the current decoding block is a luminance decoding block.
[0343] Example 40, the apparatus according to Examples 14-39, further includes: a display device coupled to the one or more processors and configured to display an image from the video data; and one or more wireless interfaces coupled to the one or more processors, the one or more wireless interfaces including one or more baseband processors and one or more transceivers.
[0344] Example 41. A method for decoding video data, the method comprising: obtaining a current decoding block from the video data; determining control data for the current decoding block; determining one or more affine motion vector clipping parameters based on the control data; selecting a sample of the current decoding block; determining an affine motion vector for the sample of the current decoding block; and clipping the affine motion vector using the one or more affine motion vector clipping parameters to generate a clipped affine motion vector.
[0345] Example 42: The method described in Example 41 according to any one of Examples 14-40.
[0346] Example 43: A non-transitory computer-readable medium including instructions that, when executed by one or more processors of a decoding device, cause the device to perform a video decoding operation on video data according to any one of Examples 14-40 above.
[0347] Example 44. An apparatus for decoding video data, the apparatus comprising: a unit for obtaining a current decoding block from the video data; a unit for determining control data for the current decoding block; a unit for determining one or more affine motion vector clipping parameters based on the control data; a unit for selecting a sample of the current decoding block; a unit for determining an affine motion vector for the sample of the current decoding block; and a unit for clipping the affine motion vector using the one or more affine motion vector clipping parameters to generate a clipped affine motion vector.
[0348] Example 45. An apparatus for decoding video data according to Example 44 of any one of Examples 14-40 above.
[0349] Example 46. A non-transitory computer-readable storage medium including instructions stored thereon, the instructions causing the one or more processors, when executed, to: obtain a current decoding block from video data; determine control data for the current decoding block; determine one or more affine motion vector clipping parameters based on the control data; select a sample of the current decoding block; determine an affine motion vector for the sample of the current decoding block; and clip the affine motion vector using the one or more affine motion vector clipping parameters to generate a clipped affine motion vector.
[0350] Example 47: A non-transitory computer-readable medium according to Example 46, including instructions for causing the one or more processors to operate according to any one of Examples 14-40 above.< / insert> < / insert> < / insertend> < / insert> < / insert> < / insert> < / insert1> < / highlightend> < / highlight> < / highlighted> < / highlight> < / highlighted> < / highlight> < / highlightend> < / highlight> < / highlighted> < / highlight> < / highlighted> < / highlight> < / insert2> < / insert1> < / insert2> < / deleteend> < / delete> < / deleteend> < / delete> < / deleteend> < / delete> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / highlightend> < / highlight> < / highlightend> < / highlight> < / highlighted> < / highlight> < / highlighted> < / highlight>
Claims
1. An apparatus for decoding video data, the apparatus comprising: Memory; as well as One or more processors coupled to the memory, the one or more processors being configured to: Obtain the current decoded block from the video data; Determine control data for the current decoding block, wherein the control data includes: The position of the current decoded block, the position having horizontal coordinate xCb and vertical coordinate yCb in the full sample cell; Specifies the width variable for the current decoded block; The height variable specifies the height of the current decoded block; Horizontal change of motion vector; The vertical change of the motion vector; Basic scaling motion vector; The height of the image in the sample; and The width of the image in the sample; Based on the control data, one or more affine motion vector clipping parameters are determined for the current decoding block, wherein the one or more affine motion vector clipping parameters include: The maximum horizontal variable is defined by the minimum value selected from the maximum horizontal image value and the maximum horizontal motion vector value, wherein the maximum horizontal image value is determined based on the width of the image, the horizontal coordinate xCb of the current decoded block, and the width variable; The minimum horizontal variable is defined by the maximum value selected from the minimum horizontal image value and the minimum horizontal motion vector value, wherein the minimum horizontal image value is determined based on the horizontal coordinate xCb of the current decoded block; The maximum vertical variable is defined by the minimum value selected from the maximum vertical image value and the maximum vertical motion vector value, wherein the maximum vertical image value is determined based on the height of the image, the vertical coordinate yCb of the current decoded block, and the height variable; and The minimum vertical variable is defined by the maximum value selected from the minimum vertical image value and the minimum vertical motion vector value, wherein the minimum vertical image value is determined based on the vertical coordinate yCb of the current decoded block; Select a sample of the current decoding block; Determine the affine motion vector for the sample used in the current decoding block; and The affine motion vectors of the sample for the current decoding block are clipped using one or more affine motion vector clipping parameters for the current decoding block to generate clipped affine motion vectors for the sample for the current decoding block.
2. The apparatus according to claim 1, wherein, The minimum horizontal motion vector value is determined based on the center motion vector value, an array of values based on the resolution values associated with the video data, and the width variable specifying the width of the current decoded block.
3. The apparatus according to claim 2, wherein, The center motion vector value is determined based on the base scaling motion vector, the horizontal variation of the motion vector, the width variable, and the height variable.
4. The apparatus according to claim 3, wherein, The base scaling motion vector corresponds to the top left corner of the current decoding block and is determined based on the control point motion vector value.
5. The apparatus according to claim 1, wherein, The maximum horizontal motion vector value is determined based on the center motion vector value, an array of values based on the resolution values associated with the video data, and the width variable specifying the width of the current decoded block.
6. The apparatus according to claim 5, wherein, The center motion vector value is determined based on the base scaling motion vector, the horizontal variation of the motion vector, the width variable, and the height variable.
7. The apparatus according to claim 6, wherein, The base scaling motion vector corresponds to the angle of the current decoding block and is determined based on the control point motion vector value.
8. The apparatus according to claim 1, wherein, The maximum vertical motion vector value is determined based on the center motion vector value, an array of values based on the block region size associated with the video data, and the height variable specifying the height of the current decoded block.
9. The apparatus according to claim 1, wherein, The minimum vertical motion vector value is determined based on the center motion vector value, the size of the data block region associated with the video data, and the height variable that specifies the height of the current decoded block.
10. The apparatus according to claim 1, wherein, The one or more processors are configured to: Multiple current decoded blocks are obtained sequentially from the video data; For each block in the plurality of current decoding blocks, a set of affine motion vector clipping parameters is determined on a per-decoding-block basis; as well as For each of the multiple current decoded blocks, the corresponding portion of the reference image is retrieved using the affine motion vector clipping parameter set on a block-by-block basis.
11. The apparatus according to claim 1, wherein, The one or more processors are configured to: Identify the reference image associated with the current decoding block; and The reference image is stored as a portion defined by the one or more affine motion vector clipping parameters.
12. The apparatus of claim 11, further comprising a memory buffer coupled to the one or more processors, wherein, The portion of the reference image is stored in the memory buffer for use in affine motion processing operations with the current decoding block.
13. The apparatus according to claim 1, wherein, The one or more processors are configured to: The current decoding block is processed using reference image data derived from the reference image indicated by the cropped affine motion vector.
14. The apparatus according to claim 1, wherein, The affine motion vector for the sample of the current decoding block is determined based on the following: a first base scaling motion vector value, a first horizontal change of the motion vector value, a first vertical change of the motion vector value, a second base scaling motion vector value, a second horizontal change of the motion vector value, a second vertical change of the motion vector value, the horizontal coordinate of the sample, and the vertical coordinate of the sample.
15. The apparatus according to claim 1, wherein, The control data includes values from the derivation table.
16. The apparatus according to claim 1, wherein, The current decoding block is a luminance decoding block.
17. The apparatus according to claim 1, further comprising: A display device coupled to the one or more processors and configured to display an image from the video data; as well as One or more wireless interfaces coupled to the one or more processors, the one or more wireless interfaces including one or more baseband processors and one or more transceivers.
18. A method for decoding video data, the method comprising: Obtain the current decoded block from the video data; Determine control data for the current decoding block, wherein the control data includes: The position of the current decoded block, the position having horizontal coordinate xCb and vertical coordinate yCb in the full sample cell; Specifies the width variable for the current decoded block; The height variable specifies the height of the current decoded block; Horizontal change of motion vector; The vertical change of the motion vector; Basic scaling motion vector; The height of the image in the sample; and The width of the image in the sample; Based on the control data, one or more affine motion vector clipping parameters are determined for the current decoding block, wherein the one or more affine motion vector clipping parameters include: The maximum horizontal variable is defined by the minimum value selected from the maximum horizontal image value and the maximum horizontal motion vector value, wherein the maximum horizontal image value is determined based on the width of the image, the horizontal coordinate xCb of the current decoded block, and the width variable; The minimum horizontal variable is defined by the maximum value selected from the minimum horizontal image value and the minimum horizontal motion vector value, wherein the minimum horizontal image value is determined based on the horizontal coordinate xCb of the current decoded block; The maximum vertical variable is defined by the minimum value selected from the maximum vertical image value and the maximum vertical motion vector value, wherein the maximum vertical image value is determined based on the height of the image, the vertical coordinate yCb of the current decoded block, and the height variable; and The minimum vertical variable is defined by the maximum value selected from the minimum vertical image value and the minimum vertical motion vector value, wherein the minimum vertical image value is determined based on the vertical coordinate yCb of the current decoded block; Select a sample of the current decoding block; Determine the affine motion vector for the sample used in the current decoding block; and The affine motion vectors of the sample for the current decoding block are clipped using one or more affine motion vector clipping parameters for the current decoding block to generate clipped affine motion vectors for the sample for the current decoding block.
19. The method according to claim 18, wherein, The minimum horizontal motion vector value is determined based on the center motion vector value, an array of values based on the block size, and the width variable specifying the width of the current decoded block.
20. The method according to claim 19, wherein, The center motion vector value is determined based on the base scaling motion vector, the horizontal variation of the motion vector, the width variable, and the height variable.
21. The method according to claim 20, wherein, The base scaling motion vector corresponds to the top left corner of the current decoding block and is determined based on the control point motion vector value.
22. The method according to claim 18, wherein, The maximum horizontal motion vector value is determined based on the center motion vector value, an array of values based on the block region size associated with the video data, and the width variable specifying the width of the current decoded block.
23. The method according to claim 22, wherein, The center motion vector value is determined based on the base scaling motion vector, the horizontal variation of the motion vector, the width variable, and the height variable.
24. The method according to claim 23, wherein, The base scaling motion vector corresponds to the angle of the current decoding block and is determined based on the control point motion vector value.
25. The method according to claim 18, wherein, The maximum vertical motion vector value is determined based on the center motion vector value, an array of values based on the block region size associated with the video data, and the height variable specifying the height of the current decoded block.
26. The method according to claim 18, wherein, The minimum vertical motion vector value is determined based on the center motion vector value, an array of values based on the block region size associated with the video data, and the height variable specifying the height of the current decoded block.
27. The method of claim 18, further comprising: Multiple current decoded blocks are obtained sequentially from the video data; For each block in the plurality of current decoding blocks, a set of affine motion vector clipping parameters is determined on a per-decoding-block basis; as well as For each of the multiple current decoded blocks, the corresponding portion of the reference image is retrieved using the affine motion vector clipping parameter set on a block-by-block basis.
28. The method of claim 18, further comprising: Identify the reference image associated with the current decoding block; as well as The reference image is stored as a portion defined by the one or more affine motion vector clipping parameters.
29. The method according to claim 28, wherein, The portion of the reference image is stored in a memory buffer for use in affine motion processing operations with the current decoding block.
30. The method of claim 18, further comprising: The current decoding block is processed using reference image data derived from the reference image indicated by the cropped affine motion vector.
31. The method according to claim 18, wherein, The affine motion vector for the sample of the current decoding block is determined based on the following: a first base scaling motion vector value, a first horizontal change of the motion vector value, a first vertical change of the motion vector value, a second base scaling motion vector value, a second horizontal change of the motion vector value, a second vertical change of the motion vector value, the horizontal coordinate of the sample, and the vertical coordinate of the sample.
32. The method according to claim 18, wherein, The control data includes values from the derivation table.
33. The method according to claim 18, wherein, The current decoding block is a luminance decoding block.
34. A non-transitory computer-readable storage medium including instructions stored thereon, the instructions causing the one or more processors to perform the following operations when executed by one or more processors: Obtain the current decoded block from the video data; Determine the control data used for the current decoding block, wherein, The control data includes: The position of the current decoded block, the position having horizontal coordinate xCb and vertical coordinate yCb in the full sample cell; Specifies the width variable for the current decoded block; The height variable specifies the height of the current decoded block; Horizontal change of motion vector; The vertical change of the motion vector; Basic scaling motion vector; The height of the image in the sample; and The width of the image in the sample; Based on the control data, one or more affine motion vector clipping parameters are determined for the current decoding block, wherein the one or more affine motion vector clipping parameters include: The maximum horizontal variable is defined by the minimum value selected from the maximum horizontal image value and the maximum horizontal motion vector value, wherein the maximum horizontal image value is determined based on the width of the image, the horizontal coordinate xCb of the current decoded block, and the width variable; The minimum horizontal variable is defined by the maximum value selected from the minimum horizontal image value and the minimum horizontal motion vector value, wherein the minimum horizontal image value is determined based on the horizontal coordinate xCb of the current decoded block; The maximum vertical variable is defined by the minimum value selected from the maximum vertical image value and the maximum vertical motion vector value, wherein the maximum vertical image value is determined based on the height of the image, the vertical coordinate yCb of the current decoded block, and the height variable; and The minimum vertical variable is defined by the maximum value selected from the minimum vertical image value and the minimum vertical motion vector value, wherein the minimum vertical image value is determined based on the vertical coordinate yCb of the current decoded block; Select a sample of the current decoding block; Determine the affine motion vector for the sample used in the current decoding block; and The affine motion vectors of the sample for the current decoding block are clipped using one or more affine motion vector clipping parameters for the current decoding block to generate clipped affine motion vectors for the sample for the current decoding block.