Adaptive Bilateral Matching for Decoder-Side Motion Vector Refinement
Patent Information
- Application Number
- JP2023575920
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-23
- Filing Date
- 2022-06-24
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2042-06-24
AI Technical Summary
Existing video coding techniques face challenges in achieving efficient compression of video data while maintaining high video quality, particularly in the context of evolving video services that demand lower bit rates without degrading image quality.
The implementation of decoder-side motion vector refinement (DMVR) using adaptive bilateral matching, which involves obtaining reference pictures, identifying motion vectors, and applying a selected search strategy to refine motion vectors, thereby improving video quality and compression efficiency.
This approach enhances the accuracy of motion vectors, leading to improved video quality and performance in video coding systems, particularly in codecs like HEVC, AVC, VVC, and future video coding standards.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] For example, aspects of the present disclosure include improving video coding techniques for decoder-side motion vector refinement (DMVR) using bilateral matching. [Background technology]
[0002] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radio telephones, so-called "smart phones," video teleconferencing devices, video streaming devices, and the like. Such devices allow video data to be processed and output for consumption. Digital video data comprises a large amount of data to satisfy the demands of consumers and video providers. For example, consumers of video data want the highest quality video with high fidelity, resolution, frame rates, and the like. As a result, the large amount of video data required to satisfy these demands places a strain on communication networks and devices that process and store the video data. Summary of the Invention [Problem to be solved by the invention]
[0003] Digital video devices may implement video coding techniques for compressing video data. Video coding is performed according to one or more video coding standards or formats. For example, video coding standards or formats include, among others, versatile video coding (VVC), high-efficiency video coding (HEVC), advanced video coding (AVC), MPEG-2 Part 2 coding (MPEG stands for moving picture experts group), as well as proprietary video codecs / formats such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media. Video coding generally employs prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit redundancy present in a video image or sequence. The goal of video coding techniques is to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation of video quality. As ever-evolving video services become available, more coding-efficient encoding techniques are needed. [Means for solving the problem]
[0004] In some examples, systems and techniques are described for decoder-side motion vector refinement (DMVR) using adaptive bilateral matching. According to at least one illustrative example, an apparatus is provided for processing video data, including at least one memory (e.g., configured to store data such as video data) and at least one processor (e.g., implemented in a circuit) coupled to the at least one memory. The at least one processor is configured to and capable of obtaining one or more reference pictures for a current picture, identifying a first motion vector and a second motion vector for a merge mode candidate, determining a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, determining one or more refined motion vectors based on at least one of the first motion vector or the second motion vector and the one or more reference pictures using the selected motion vector search strategy, and processing the merge mode candidate using the one or more refined motion vectors.
[0005] In another example, a method for processing video data is provided, the method including obtaining one or more reference pictures for a current picture, identifying a first motion vector and a second motion vector for a merge mode candidate, determining a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, determining one or more refined motion vectors based on at least one of the first motion vector or the second motion vector and the one or more reference pictures using the selected motion vector search strategy, and processing the merge mode candidate using the one or more refined motion vectors.
[0006] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to obtain one or more reference pictures for a current picture, identify a first motion vector and a second motion vector for a merge mode candidate, determine a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, determine one or more refined motion vectors based on at least one of the first motion vector or the second motion vector and the one or more reference pictures using the selected motion vector search strategy, and process the merge mode candidate using the one or more refined motion vectors.
[0007] In another example, an apparatus is provided for processing video data, the apparatus including means for obtaining one or more reference pictures for a current picture, means for identifying a first motion vector and a second motion vector for a merge mode candidate, means for determining a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, means for determining one or more refined motion vectors based on at least one of the first motion vector or the second motion vector and the one or more reference pictures using the selected motion vector search strategy, and means for processing the merge mode candidate using the one or more refined motion vectors.
[0008] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter, which subject matter should be understood by reference to the entire specification of this patent, any or all drawings, and appropriate portions of each claim.
[0009] The above, together with other features and aspects, will become more apparent with reference to the following specification, claims, and accompanying drawings.
[0010] Illustrative aspects of the present application are described in detail below with reference to the following figures: [Brief description of the drawings]
[0011] [Figure 1] 1A-1C are block diagrams illustrating examples of encoding and decoding devices in accordance with some examples of this disclosure. [Figure 2A] 1 is a conceptual diagram illustrating example spatially adjacent motion vector candidates for merge mode, in accordance with some examples of this disclosure. [Figure 2B] FIG. 2 is a conceptual diagram illustrating example spatially neighboring motion vector candidates for advanced motion vector prediction (AMVP) mode, in accordance with some examples of this disclosure. [Figure 3A] FIG. 1 is a conceptual diagram illustrating example temporal motion vector predictor (TMVP) candidates, in accordance with some examples of this disclosure. [Figure 3B] 1 is a conceptual diagram illustrating an example of motion vector scaling, according to some examples of the present disclosure. [Figure 4A] 1 is a conceptual diagram illustrating examples of neighboring samples of a current coding unit used to estimate motion compensation parameters for the current coding unit, in accordance with some examples of this disclosure. [Figure 4B] 1 is a conceptual diagram illustrating examples of neighboring samples of a reference block used to estimate motion compensation parameters for a current coding unit, in accordance with some examples of this disclosure. [Diagram 5] A diagram illustrating locations of spatial merging candidates for use in processing blocks, in accordance with some examples of this disclosure. [Figure 6] A diagram illustrating aspects of motion vector scaling for temporal merge candidates for use in processing blocks, in accordance with some examples of this disclosure. [Figure 7]11A-11C illustrate aspects of temporal merge candidates for use in processing blocks, in accordance with certain examples of this disclosure. [Figure 8] FIG. 1 illustrates aspects of bilateral matching, according to some examples of the present disclosure. [Figure 9] FIG. 2 illustrates aspects of bi-directional optical flow (BDOF), in accordance with some examples of the present disclosure. [Figure 10] A diagram showing a search area region according to some examples of the present disclosure. [Figure 11] 1 is a flowchart illustrating an example process for decoder-side motion vector refinement using adaptive bilateral matching, in accordance with some examples of this disclosure. [Figure 12] 1 is a block diagram illustrating an example video encoding device, in accordance with some examples of this disclosure. [Figure 13] 1 is a block diagram illustrating an example video decoding device, in accordance with some examples of this disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Some aspects of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects may be applied independently, and some of them may be applied in combination. In the following description, for the purpose of explanation, specific details are set forth to provide a thorough understanding of the aspects of the present application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and descriptions are not intended to be limiting.
[0013] The following description merely provides exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments provides those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.
[0014] A video coding device (e.g., an encoding device, a decoding device, or a composite encoding-decoding device) implements video compression techniques to efficiently encode and / or decode video data. Video compression techniques may include applying different prediction modes, including spatial prediction (e.g., intra-frame or intra prediction), temporal prediction (e.g., inter-frame or inter prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques to reduce or remove redundancy inherent in video sequences. A video encoder may partition each picture of an original video sequence into rectangular regions called video blocks or coding units (described in more detail below). These video blocks may be encoded using particular prediction modes.
[0015] A video block may be divided into one or more groups of smaller blocks in one or more ways. A block may include a coding tree block, a prediction block, a transform block, or other appropriate block. In general, references to a "block" may refer to such a video block (e.g., a coding tree block, a coding block, a prediction block, a transform block, or other appropriate block or sub-block as understood by one of ordinary skill in the art) unless otherwise specified. Furthermore, each of these blocks may also be referred to interchangeably herein as a "unit" (e.g., a coding tree unit (CTU), a coding unit, a prediction unit (PU), a transform unit (TU), etc.). In some cases, a unit may refer to a coding logical unit that is encoded in the bitstream, while a block may refer to a portion of a video frame buffer that is the subject of a process.
[0016] For inter-prediction modes, the video encoder may search for a block similar to the block encoded in a frame (or picture) located at another temporal location, called a reference frame or picture. The video encoder may limit this search to a certain spatial displacement from the block to be encoded. A two-dimensional (2D) motion vector, including a horizontal displacement component and a vertical displacement component, may be used to identify the best match. For intra-prediction modes, the video encoder may use spatial prediction techniques to form a predicted block based on data from previously encoded neighboring blocks in the same picture.
[0017] The video encoder may determine a prediction error. For example, the prediction may be determined as a difference between pixel values in the block being coded and pixel values in a predicted block. The prediction error may also be referred to as a residual. The video encoder may also apply a transform (e.g., a discrete cosine transform (DCT) or other appropriate transform) to the prediction error to generate transform coefficients. After the transform, the video encoder may quantize the transform coefficients. The quantized transform coefficients and motion vectors may be represented using syntax elements and, together with control information, form a coded representation of the video sequence. In some instances, the video encoder may entropy code the syntax elements, thereby further reducing the number of bits required for their representation.
[0018] A video decoder may use the syntax elements and control information discussed above to construct prediction data (e.g., a prediction block) for decoding a current frame. For example, the video decoder may add the prediction block and a compressed prediction error. The video decoder may determine the compressed prediction error by weighting the transform basis functions using the quantized coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.
[0019] Described herein are systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to herein as "systems and techniques") for increasing the accuracy of one or more motion vectors that may be used by a video coding device (e.g., a video decoder or decoding device) when performing a prediction technique (e.g., an inter-prediction mode). For example, the systems and techniques may perform bilateral matching for decoder-side motion vector refinement (DMVR). Bilateral matching is a technique that refines a pair of two initial motion vectors. Such refinement may occur with a search around the pair of initial motion vectors to derive an updated motion vector that minimizes a block matching cost. The block matching cost may be generated in various ways, including using a sum of absolute difference (SAD) criterion, a sum of absolute transformed difference (SATD) criterion, a sum of square error (SSE) criterion, or other such criteria. The aspects described herein can increase the accuracy of motion vectors for bi-predictive merge candidates, resulting in improved video quality and associated improved device performance for devices operating in accordance with the aspects described herein.
[0020] In some aspects, the systems and techniques may be used to perform adaptive bilateral matching for DMVR. For example, the systems and techniques may perform bilateral matching using different search strategies and / or search parameters for different coded blocks. As described in more detail below, adaptive bilateral matching for DMVR may be based on a selected search strategy determined or signaled for a given block. The selected search strategy may include one or more constraints for the bilateral matching search process. In some examples, the selected search strategy may additionally or alternatively include one or more constraints for the first motion vector differential and / or the second motion vector differential. In some examples, the selected search strategy may include one or more constraints between the first motion vector differential and the second motion vector differential.
[0021] In some aspects, a constraint is selected for the motion vector to be refined. The constraint may be a mirroring constraint, a zero constraint for the first vector, a zero constraint for the second vector, or other types of constraints. In some cases, the constraint is applied to merge mode coded blocks within a merge candidate that meet one or more DMVR conditions. The one or more constraints may then be used with one or more search strategies to identify candidates and select the refined motion vector.
[0022] In some aspects, different search strategies are used. The search strategies may be grouped into multiple subsets, with each subset including one or more search strategies. In some cases, the decoder may use a syntax element to determine the selected subset. For example, the encoder may include the syntax element in a bitstream. In such an example, the decoder may receive the bitstream and decode the syntax element from the bitstream. The decoder may use the syntax element to determine a selected subset and any associated constraints for a given block or blocks of video data included in the bitstream. Using the selected subset and any constraints associated with the subsets, the decoder may process the motion vectors (e.g., two motion vectors of a bi-predictive merge candidate) to identify a refined motion vector. In one illustrative aspect, an adaptive bilateral mode is provided, in which the coding device signals (e.g., together with a signaling structure as part of a new adaptive bilateral mode) a selected motion information candidate that satisfies an associated DMVR condition.
[0023] Using the search strategies and associated constraints described above can provide improvements to decoder-side motion vector refinement, such as by providing adaptive bilateral motion vector refinement using selectable search algorithms and associated constraints. Such improvements to decoder-side motion vector refinement can be used with various video codecs, such as implementing enhanced compression models (ECMs). Examples described herein include implementations applied to multi-pass DMVR to improve ECM systems operating according to one or more video coding standards. The techniques described herein can be implemented using one or more coding devices, including one or more encoding devices, decoding devices, or composite encoding-decoding devices. The coding devices can be implemented by one or more of the player devices, such as a mobile device, an augmented reality (XR) device, a vehicle or a vehicle's computing system, a server device or system (e.g., a distributed server system including multiple servers, a single server device or system, etc.), or other devices or systems.
[0024] The systems and techniques described herein may be applied to any existing, developing, and / or future video coding standards, including, but not limited to, High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Versatile Video Coding (VVC), VP9, AOMedia Video1 (AV1) formats / codecs, and / or other existing, developing, or future developed video coding standards. The systems and techniques described herein may improve the operation of communication systems and devices in the system by improving the performance of video data transfer by the devices along with improved compression and associated improved video quality based on the selection of improved motion vectors from adaptive bilateral matching as described herein.
[0025] FIG. 1 is a block diagram illustrating an example of a system 100 including an encoding device 104 and a decoding device 112. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device. The source device and / or the receiving device may include electronic devices, such as a mobile or fixed telephone handset (e.g., a smartphone, a mobile phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video gaming console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the source device and the receiving device may include one or more wireless transceivers for wireless communication. The coding techniques described herein are applicable to video coding in various multimedia applications, including streaming video transmission (e.g., over the Internet), television broadcast or transmission, encoding digital video for storage on a data storage medium, decoding digital video stored on a data storage medium, or other applications. The term coding as used herein may refer to encoding and / or decoding. In some examples, the system 100 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.
[0026] Encoding device 104 (or encoder) may be used to encode the video data using a video coding standard, format, codec, or protocol to generate an encoded video bitstream. Examples of video coding standards and formats / codecs include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, including its Scalable Video Coding (SVC) and Multiview Video Coding (MVC) extensions, ITU-T H.264 (also called ISO / IEC MPEG-4 AVC), High Efficiency Video Coding (HEVC) or ITU-T H.265, and Versatile Video Coding (VVC) or ITU-T H.266. There are various extensions of HEVC that deal with multi-layer video coding, including range and screen content coding extensions, 3D video coding (3D-HEVC), and multiview extensions (MV-HEVC) and scalable extensions (SHVC). HEVC and its extensions are developed by the Joint Collaboration Team on Video Coding (JCT-VC), as well as the Joint Collaboration Team on 3D Video Coding Extension Development (JCT-3V) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). VP9, AOMedia Video1 (AV1) developed by the Alliance for Open Media Alliance of Open Media (AOMedia), and Essential Video Coding (EVC) are other video coding standards to which the techniques described herein may be applied.
[0027] The techniques described herein may be applied to existing video codecs (e.g., High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), or other suitable existing video codecs) and / or may be efficient coding tools for any video coding standard, including video coding standards being developed and / or developing and / or future video coding standards, such as, for example, VVC, and / or other video coding standards being developed or to be developed in the future. For example, examples described herein may be performed using video codecs such as VVC, HEVC, AVC, and / or extensions thereof. However, the techniques and systems described herein may also be applicable to other coding standards, codecs, or formats, such as MPEG, JPEG (or other coding standards for still images), VP9, AV1, extensions thereof, or other suitable coding standards that are already available or not yet available or developed. For example, in some examples, the encoding device 104 and / or the decoding device 112 may operate according to a proprietary video codec / format, such as AV1, an extension of AVI, and / or a successor version of AV1 (e.g., AV2), or other proprietary formats or industry standards. Thus, those skilled in the art will understand that although the techniques and systems described herein may be described with respect to a particular video coding standard, the description should not be construed as applying only to that particular standard.
[0028] 1, video source 102 may provide video data to encoding device 104. Video source 102 may be part of a source device or part of a device other than the source device. Video source 102 may include a video capture device (e.g., a video camera, a camera phone, a video phone, etc.), a video archive containing stored video, a video server or content provider providing video data, a video feed interface receiving video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.
[0029] The video data from the video source 102 may include one or more input pictures or frames. A picture or frame is a still image that is possibly part of a video. In some examples, the data from the video source 102 may be a still image that is not part of a video. In HEVC, VVC, and other video coding specifications, a video sequence may include a series of pictures. A picture may include three sample arrays, denoted SL, SCb, and SCr. SL is a two-dimensional array of luma samples, SCb is a two-dimensional array of Cb chrominance samples, and SCr is a two-dimensional array of Cr chrominance samples. The chrominance samples are sometimes referred to herein as "chroma" samples. A pixel may refer to all three components (luma samples and chroma samples) for a given location in the array of a picture. In other cases, a picture may be monochrome and include only an array of luma samples, in which case the terms pixel and sample may be used interchangeably. For example techniques described herein that refer to individual samples for illustration purposes, the same techniques may be applied to pixels (e.g., all three sample components for a given location in an array of pictures). For example techniques described herein that refer to pixels (e.g., all three sample components for a given location in an array of pictures) for illustration purposes, the same techniques may be applied to individual samples.
[0030] The encoder engine 106 (or encoder) of the encoding device 104 encodes video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a sequence of one or more coded video sequences. A coded video sequence (CVS) includes a sequence of access units (AUs) starting from an AU having a random access point picture with some characteristics in a base layer to just before a next AU having a random access point picture with some characteristics in a base layer. For example, some characteristics of a random access point picture that starts a CVS may include a RASL flag (e.g., NoRaslOutputFlag) equal to 1. Otherwise, a random access point picture (having a RASL flag equal to 0) does not start a CVS. An access unit (AU) includes one or more coded pictures and control information corresponding to coded pictures that share the same output time. At the bitstream level, coded slices of a picture are encapsulated into data units called network abstraction layer (NAL) units. For example, an HEVC video bitstream may contain one or more CVSs that contain NAL units. Each of the NAL units has a NAL unit header. In one example, the header is 1 byte for H.264 / AVC (excluding multi-layer extensions) and 2 bytes for HEVC. Syntax elements in the NAL unit header take designated bits and are therefore recognizable to all kinds of systems and transport layers, such as Transport Stream, Real-time Transport (RTP) protocol, file formats, among others.
[0031] Two classes of NAL units exist in the HEVC standard, including video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain coded picture data that form a coded video bitstream. For example, a sequence of bits that form a coded video bitstream exists in a VCL NAL unit. A VCL NAL unit may contain one slice or slice segment (described below) of coded picture data, and a non-VCL NAL unit contains control information about one or more coded pictures. In some cases, a NAL unit may be referred to as a packet. A HEVC AU includes VCL NAL units that contain coded picture data and non-VCL NAL units that correspond to the coded picture data (if any). A non-VCL NAL unit may include, in addition to other information, a parameter set that has high-level information about the coded video bitstream. For example, the parameter sets may include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). In some cases, each slice or other portion of the bitstream may reference a single active PPS, SPS, and / or VPS to allow the decoding device 112 to access information that can be used to decode the slice or other portion of the bitstream.
[0032] A NAL unit may include a sequence of bits that form a coded representation of video data (e.g., an encoded video bitstream, a CVS of a bitstream, etc.), such as a coded representation of a picture in a video. The encoder engine 106 generates the coded representation of a picture by partitioning each picture into a number of slices. A slice is independent of other slices such that information in a slice is coded without depending on data from other slices in the same picture. A slice includes one or more slice segments, including an independent slice segment and, if present, one or more dependent slice segments that depend on a previous slice segment.
[0033] In HEVC, a slice is then partitioned into coding tree blocks (CTBs) of luma samples and chroma samples. A CTB of luma samples and one or more CTBs of chroma samples, together with syntax for the samples, are called a coding tree unit (CTU). A CTU may also be called a "tree block" or a "largest coding unit" (LCU). A CTU is the basic processing unit for HEVC encoding. A CTU may be divided into multiple coding units (CUs) of various sizes. A CU contains a luma sample array and a chroma sample array, called a coding block (CB).
[0034] The luma CB and the chroma CB may be further divided into prediction blocks (PBs). A PB is a block of luma or chroma component samples that uses the same motion parameters for inter prediction or intra block copy (IBC) prediction (when available or enabled for use). The luma PB and one or more chroma PBs, together with associated syntax, form a prediction unit (PU). For inter prediction, a set of motion parameters (e.g., one or more motion vectors, reference indexes, etc.) is signaled in the bitstream for each PU and is used for inter prediction of the luma PB and one or more chroma PBs. The motion parameters may also be referred to as motion information. The CB may also be partitioned into one or more transform blocks (TBs). A TB represents a rectangular block of samples of a color component to which a residual transform (e.g., possibly the same two-dimensional transform) is applied to code the prediction residual signal. A transform unit (TU) represents a TB of luma and chroma samples, as well as corresponding syntax elements. Transform coding is described in more detail below.
[0035] The size of a CU corresponds to the size of a coding mode and may be square in shape. For example, the size of a CU may be 8×8 samples, 16×16 samples, 32×32 samples, 64×64 samples, or any other suitable size up to the size of a corresponding CTU. The phrase “N×N” is used herein to refer to pixel dimensions of a video block in terms of vertical and horizontal dimensions (e.g., 8 pixels×8 pixels). The pixels in a block may be arranged in rows and columns. In some implementations, a block may not have the same number of pixels in the horizontal direction as in the vertical direction. Syntax data associated with a CU may, for example, describe the partitioning of the CU into one or more PUs. The partitioning mode may differ whether the CU is intra-prediction mode coded or inter-prediction mode coded. The PUs may be partitioned to be non-square in shape. Syntax data associated with a CU may also, for example, describe the partitioning of a CU into one or more TUs according to a CTU. The TUs may be square or non-square in shape.
[0036] According to the HEVC standard, the transform may be performed using transform units (TUs). The TUs may be different for different CUs. The TUs may be sized based on the size of the PUs in a given CU. The TUs may be the same size or may be smaller than the PUs. In some examples, the residual samples corresponding to a CU may be subdivided into smaller units using a quadtree structure known as a residual quadtree (RQT). Leaf nodes of the RQT may correspond to TUs. Pixel difference values associated with the TUs may be transformed to generate transform coefficients. The transform coefficients may then be quantized by the encoder engine 106.
[0037] Once a picture of video data is partitioned into CUs, the encoder engine 106 predicts each PU using a prediction mode. The prediction unit or prediction block is then subtracted from the original video data to obtain a residual (described below). For each CU, a prediction mode may be signaled inside the bitstream using syntax data. The prediction mode may include intra prediction (or intra-picture prediction) or inter prediction (or inter-picture prediction). Intra prediction exploits the correlation between spatially adjacent samples in a picture. For example, with intra prediction, each PU is predicted from adjacent image data in the same picture, for example, using DC prediction to find the average value for the PU, planar prediction to fit a flat surface to the PU, directional prediction to extrapolate from adjacent data, or any other suitable type of prediction. Inter prediction uses temporal correlation between pictures to derive a motion-compensated prediction for a block of image samples. For example, with inter prediction, each PU is predicted using motion-compensated prediction from image data in one or more reference pictures (before or after the current picture in output order). The decision whether to code a picture area using inter-picture prediction or intra-picture prediction may be made, for example, at the CU level.
[0038] The encoder engine 106 and the decoder engine 116 (described in more detail below) may be configured to operate according to VVC. According to VVC, a video coder (such as the encoder engine 106 and / or the decoder engine 116) partitions a picture into multiple coding tree units (CTUs) (a CTB for luma samples and one or more CTBs for chroma samples, together with a syntax for the samples, are referred to as a CTU). The video coder may partition the CTUs according to a tree structure, such as a quad-tree binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple partition types, such as the distinction between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels, including a first level partitioned according to a quad-tree partition and a second level partitioned according to a binary tree partition. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to coding units (CUs).
[0039] In the MTT partition structure, blocks may be partitioned using quadtree partitions, binary tree partitions, and one or more types of ternary tree partitions. Ternary tree partitions are partitions in which a block is divided into three subblocks. In some examples, ternary tree partitions divide a block into three subblocks without splitting the original block through the center. Partition types in MTT (e.g., quadtree, binary tree, and ternary tree) may be symmetric or asymmetric.
[0040] When operating according to the AV1 codec, encoding device 104 and decoding device 112 may be configured to code video data in blocks. In AV1, the largest coding block that may be processed is called a superblock. In AV1, a superblock may be either 128×128 luma samples or 64×64 luma samples. However, in successor video coding formats (e.g., AV2), a superblock may be defined by a different (e.g., larger) luma sample size. In some examples, a superblock is the top of a block quadtree. Encoding device 104 may further partition the superblock into smaller coding blocks. Encoding device 104 may partition the superblock and other coding blocks into smaller blocks using rectangular or non-rectangular partitions. Non-rectangular blocks may include N / 2×N, N×N / 2, N / 4×N, and N×N / 4 blocks. Encoding device 104 and decoding device 112 may perform separate prediction and transformation processes for each of the coding blocks.
[0041] AV1 also defines tiles of video data. A tile is a rectangular array of superblocks that may be coded independently of other tiles. That is, encoding device 104 and decoding device 112 may encode and decode coding blocks in a tile, respectively, without using video data from other tiles. However, encoding device 104 and decoding device 112 may perform filtering across tile boundaries. Tiles may be uniform or non-uniform in size. Tile-based coding may enable parallel processing and / or multithreading for encoder and decoder implementations.
[0042] In some examples, the video coder may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video coder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luma component and another QTBT or MTT structure for both chroma components (or two QTBT and / or MTT structures for each chroma component).
[0043] The video coder may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning, or other partitioning structures.
[0044] In some examples, one or more slices of a picture are assigned a slice type. The slice types include intra-coded slices (I slices), inter-coded P slices, and inter-coded B slices. An I slice (independently decodable intra-coded frame) is a slice of a picture that is coded only by intra prediction and is therefore independently decodable because an I slice only needs data in a frame to predict any prediction unit or prediction block of the slice. A P slice (unidirectionally predicted frame) is a slice of a picture that can be coded using intra prediction and using unidirectional inter prediction. Each prediction unit or prediction block in a P slice is coded using either intra prediction or inter prediction. When inter prediction is applied, a prediction unit or prediction block is predicted by only one reference picture, and therefore the reference samples are from only one reference area of a frame. A B slice (bidirectionally predicted frame) is a slice of a picture that can be coded using intra prediction and using inter prediction (e.g., either bidirectional or unidirectional). A prediction unit or prediction block of a B slice may be bidirectionally predicted from two reference pictures, where each picture contributes to one reference region, and the sample sets of the two reference regions are weighted (for example, with equal weights or with different weights) to generate a prediction signal of the bidirectionally predicted block. As described above, the slices of a picture are coded independently. In some cases, a picture may be coded as just one slice.
[0045] As mentioned above, intra-picture prediction of a picture exploits correlation between spatially adjacent samples in a picture. There are multiple intra-prediction modes (also called "intra modes"). In some examples, intra-prediction of a luma block includes 35 modes, including a planar mode, a DC mode, and 33 angle modes (e.g., a diagonal intra-prediction mode and an angle mode adjacent to the diagonal intra-prediction mode). The 35 modes of intra-prediction are indexed as shown in Table 1 below. In other examples, more intra-modes may be defined, including prediction angles that may not yet be represented by the 33 angle modes. In other examples, the prediction angles associated with the angle modes may differ from those used in HEVC.
[0046] [Table 1]
[0047] Inter-picture prediction uses temporal correlation between pictures to derive motion-compensated predictions for blocks of image samples. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is indicated by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement and Δy specifies the vertical displacement of the reference block relative to the position of the current block. In some cases, the motion vector (Δx, Δy) may be integer sample precision (also called integer precision), in which case the motion vector points to an integer pel grid (or integer pixel sampling grid) of a reference frame. In some cases, the motion vector (Δx, Δy) may be fractional sample precision (also called fractional pel precision or non-integer precision) to more accurately capture the motion of the underlying object without being restricted to the integer pel grid of the reference frame. The precision of the motion vector may be represented by the quantization level of the motion vector. For example, the quantization level may be integer precision (e.g., 1 pixel) or fractional pel precision (e.g., ¼ pixel, ½ pixel, or other sub-pixel value). When the corresponding motion vector has fractional sample precision, interpolation is applied to the reference picture to derive a prediction signal. For example, samples available at integer positions may be filtered (e.g., using one or more interpolation filters) to estimate values at fractional positions. A previously decoded reference picture is indicated by a reference index (refIdx) to a reference picture list. The motion vector and the reference index may be referred to as motion parameters. Two types of inter-picture prediction may be performed, including uni-prediction and bi-prediction.
[0048] In inter prediction using bi-prediction (also called bidirectional inter prediction), two sets of motion parameters (Δx 0 , y 0 , refIdx 0 and Δx 1 , y 1 , refIdx1 ) is used. For example, in bi-prediction, each prediction block uses two motion compensated prediction signals to generate a B prediction unit. The two motion compensated predictions are then combined to obtain a final motion compensated prediction. For example, the two motion compensated predictions may be combined by averaging. In another example, weighted prediction may be used, in which case a different weight may be applied to each motion compensated prediction. Reference pictures that may be used in bi-prediction are stored in two separate lists, denoted as list 0 and list 1. Using a motion estimation process, motion parameters may be derived in the encoder.
[0049] In inter prediction using uni-prediction (also called unidirectional inter prediction), one set of motion parameters (Δx 0 , y 0 , refIdx 0 For example, in uni-prediction, each prediction block uses at most one motion compensated prediction signal to generate P prediction units.
[0050] A PU may include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when a PU is encoded using intra prediction, the PU may include data describing the intra prediction mode of the PU. As another example, when a PU is encoded using inter prediction, the PU may include data defining a motion vector of the PU. The data defining the motion vector of the PU may describe, for example, the horizontal component (Δx) of the motion vector, the vertical component (Δy) of the motion vector, the resolution of the motion vector (e.g., integer precision, ¼ pixel precision, or ⅛ pixel precision), the reference picture to which the motion vector points, the reference index, the reference picture list (e.g., list 0, list 1, or list C) of the motion vector, or any combination thereof.
[0051] AV1 includes two common techniques for encoding and decoding coding blocks of video data. The two common techniques are intra prediction (e.g., intraframe prediction or spatial prediction) and inter prediction (e.g., interframe prediction or temporal prediction). In the context of AV1, when predicting a block of a current frame of video data using an intra prediction mode, the encoding device 104 and the decoding device 112 do not use video data from other frames of the video data. In most intra prediction modes, the video encoding device 104 encodes a block of a current frame based on a difference between a sample value in the current block and a predicted value generated from a reference sample in the same frame. The video encoding device 104 determines the predicted value generated from the reference sample based on the intra prediction mode.
[0052] After performing prediction using intra prediction and / or inter prediction, the encoding device 104 may perform transformation and quantization. For example, following prediction, the encoder engine 106 may calculate a residual value corresponding to the PU. The residual value may comprise pixel difference values between a current block of pixels being coded (PU) and a predictive block (e.g., a predicted version of the current block) used to predict the current block. For example, after generating a predictive block (e.g., issuing an inter prediction or an intra prediction), the encoder engine 106 may generate a residual block by subtracting the predictive block generated by the prediction unit from the current block. The residual block includes a set of pixel difference values that quantify differences between pixel values of the current block and pixel values of the predictive block. In some examples, the residual block may be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of pixel values.
[0053] Any residual data that may remain after prediction is performed is transformed using a block transform, which may be based on a discrete cosine transform, a discrete sine transform, an integer transform, a wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., size 32×32, 16×16, 8×8, 4×4, or other suitable size) may be applied to the residual data in each CU. In some aspects, TUs may be used for the transform and quantization processes implemented by the encoder engine 106. A given CU having one or more PUs may also include one or more TUs. As described in more detail below, the residual values may be transformed into transform coefficients using a block transform, and then quantized and scanned using the TUs to generate serialized transform coefficients for entropy coding.
[0054] In some aspects, following intra-predictive coding or inter-predictive coding using a PU of a CU, the encoder engine 106 may calculate residual data for the TU of the CU. The PU may comprise pixel data in the spatial domain (or pixel domain). The TU may comprise coefficients in the transform domain after applying a block transform. As previously mentioned, the residual data may correspond to pixel difference values between pixels of an uncoded picture and predicted values corresponding to the PU. The encoder engine 106 may form a TU including the residual data for the CU and then transform the TU to generate transform coefficients for the CU.
[0055] The encoder engine 106 may perform quantization of the transform coefficients. Quantization achieves further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization may reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient having an n-bit value may be truncated to an m-bit value during quantization, where n is greater than m.
[0056] Once quantization is performed, the coded video bitstream includes the quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), partition information, and any other suitable data, such as other syntax data. Various elements of the coded video bitstream may then be entropy coded by the encoder engine 106. In some examples, the encoder engine 106 may scan the quantized transform coefficients utilizing a predefined scan order to generate serialized vectors that may be entropy coded. In some examples, the encoder engine 106 may perform an adaptive scan. After scanning the quantized transform coefficients to form vectors (e.g., one-dimensional vectors), the encoder engine 106 may entropy code the vectors. For example, the encoder engine 106 may use context-adaptive variable length coding, context-adaptive binary arithmetic coding, syntax-based context-adaptive binary arithmetic coding, probability interval partition entropy coding, or another suitable entropy coding technique.
[0057] An output 110 of the encoding device 104 may transmit the NAL units constituting the encoded video bitstream data to a decoding device 112 of a receiving device via a communication link 120. An input 114 of the decoding device 112 may receive the NAL units. The communication link 120 may include channels provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network may include any wireless interface or combination of wireless interfaces, and may include any suitable wireless network (e.g., the Internet or other wide area network, a packet-based network, WiFi™, radio frequency (RF), ultra-wideband (UWB), WiFi-Direct, cellular, Long-Term Evolution (LTE), WiMax™, etc.). The wired network may include any wired interface (e.g., fiber, Ethernet, powerline Ethernet, Ethernet over coaxial cable, Digital Signal Line (DSL), etc.). Wired and / or wireless networks may be implemented using a variety of equipment, such as base stations, routers, access points, bridges, gateways, switches, etc. The encoded video bitstream data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to a receiving device.
[0058] In some examples, the encoding device 104 may store the encoded video bitstream data in the storage 108. The output 110 may retrieve the encoded video bitstream data from the encoder engine 106 or from the storage 108. The storage 108 may include any of a variety of distributed or locally accessed data storage media. For example, the storage 108 may include a hard drive, a storage disk, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. The storage 108 may also include a decoded picture buffer (DPB) for storing reference pictures for use in inter prediction. In a further example, the storage 108 may correspond to a file server or another intermediate storage device that may store the encoded video generated by the source device. In such a case, a receiving device including the decoding device 112 may access the stored video data from the storage device via streaming or download. The file server may be any type of server capable of storing the encoded video data and transmitting the encoded video data to the receiving device. Exemplary file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device may access the encoded video data through any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored in the file server. The transmission of the encoded video data from storage 108 may be a streaming transmission, a download transmission, or a combination thereof.
[0059] An input 114 of the decoding device 112 may receive the encoded video bitstream data and provide the video bitstream data to the decoder engine 116 or to the storage 118 for later use by the decoder engine 116. For example, the storage 118 may include a decoded picture buffer (DPB) for storing reference pictures for use in inter prediction. A receiving device including the decoding device 112 may receive the encoded video data to be decoded via the storage 108. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the receiving device. A communication medium for transmitting the encoded video data may comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from a source device to a receiving device.
[0060] The decoder engine 116 may decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more coded video sequences that make up the encoded video data. The decoder engine 116 may then rescale the encoded video bitstream data and perform an inverse transform on the encoded video bitstream data. The residual data is then passed to a prediction stage of the decoder engine 116. The decoder engine 116 then predicts a block of pixels (e.g., a PU). In some examples, the prediction is added to the output of the inverse transform (the residual data).
[0061] Video decoding device 112 may output the decoded video to video destination device 122, which may include a display or other output device for displaying the decoded video data to a content consumer. In some aspects, video destination device 122 may be part of a receiving device that includes decoding device 112. In some aspects, video destination device 122 may be part of a separate device other than the receiving device.
[0062] In some aspects, the video encoding device 104 and / or the video decoding device 112 may be integrated with an audio encoding device and an audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 may also include other hardware or software necessary to implement the coding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. The video encoding device 104 and the video decoding device 112 may be integrated in the respective devices as part of a combined encoder / decoder (codec).
[0063] The exemplary system shown in FIG. 1 is one illustrative example that may be used herein. Techniques for processing video data using techniques described herein may be performed by any digital video encoding and / or decoding device. Generally, the techniques of this disclosure are performed by a video encoding device or a video decoding device, but these techniques may also be performed by a composite video encoder-decoder, usually referred to as a "codec." In addition, the techniques of this disclosure may also be performed by a video preprocessor. The source device and the receiving device are merely examples of coding devices, such that the source device generates coded video data for transmission to the receiving device. In some examples, the source device and the receiving device may operate substantially symmetrically, such that each of the devices includes video encoding and decoding components. Thus, the exemplary system may support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0064] This disclosure may generally refer to "signaling" some information, such as syntax elements. The term "signaling" may generally refer to communication of values for syntax elements and / or other data used to decode encoded video data. For example, video encoding device 104 may signal values for syntax elements in a bitstream. In general, signaling refers to generating values in a bitstream. As noted above, video source 102 may transport the bitstream to video destination device 122 substantially in real time or not in real time, such as may occur when storing syntax elements in storage 108 for later retrieval by video destination device 122.
[0065] A video bitstream may also include Supplemental Enhancement Information (SEI) messages. For example, an SEI NAL unit may be part of the video bitstream. In some cases, the SEI message may include information that is not required by the decoding process. For example, the information in the SEI message may not be essential for a decoder to decode a video picture of the bitstream, but the decoder may use the information to improve the display or processing of the picture (e.g., the decoded output). The information in the SEI message may be embedded metadata. In one illustrative example, the information in the SEI message may be used by a decoder-side entity to improve the viewability of the content. In some instances, some application standards may mandate the presence of such SEI messages in the bitstream so that quality improvements can be provided to all devices that comply with the application standard (e.g., the carrying of a frame-packing SEI message for a frame compatible plano-stereoscopic 3DTV video format in which an SEI message is carried for every frame of the video, the processing of a recovery point SEI message, the use of the pan-scan scan rectangle SEI message in DVB, in addition to many other examples).
[0066] As described above, for each block, a set of motion information (also referred to herein as motion parameters) may be available. The set of motion information includes motion information for a forward prediction direction and a backward prediction direction. The forward prediction direction and the backward prediction direction are two prediction directions of a bidirectional prediction mode, in which case the terms "forward" and "backward" do not necessarily have a geometric meaning. Instead, "forward" and "backward" correspond to reference picture list 0 (RefPicList0 or L0) and reference picture list 1 (RefPicList1 or L1) of the current picture. In some examples, when only one reference picture list is available for a picture or slice, only RefPicList0 is available, and the motion information of each block of the slice is always forward.
[0067] In some cases, a motion vector, together with its reference index, is used in the coding process (e.g., motion compensation). Such motion vectors with associated reference indexes are denoted as a uni-predictive set of motion information. For each prediction direction, the motion information may include a reference index and a motion vector. In some cases, the motion vector itself may be referenced in such a way that, for simplicity, it is assumed that the motion vector has an associated reference index. The reference index is used to identify a reference picture in the current reference picture list (RefPicList0 or RefPicList1). The motion vector has horizontal and vertical components that provide an offset from a coordinate location of the current picture to a coordinate in the reference picture identified by the reference index. For example, the reference index may indicate a particular reference picture to be used for a block in the current picture, and the motion vector may indicate where in the reference picture is the best matching block (the block that best matches the current block).
[0068] In H.264 / AVC, each inter macroblock (MB) may be partitioned in four different ways, including one 16×16 MB partition, two 16×8 MB partitions, two 8×16 MB partitions, and four 8×8 MB partitions. Different MB partitions within one MB may have different reference index values (RefPicList0 or RefPicList1) for each direction. In some cases, when an MB is not partitioned into four 8×8 MB partitions, there may be only one motion vector per MB partition in each direction. In some cases, when an MB is partitioned into four 8×8 MB partitions, each 8×8 MB partition may be further partitioned into sub-blocks, in which case each sub-block may have a different motion vector in each direction. In some examples, there are four different ways to obtain sub-blocks from an 8×8 MB partition, including one 8×8 sub-block, two 8×4 sub-blocks, two 4×8 sub-blocks, and four 4×4 sub-blocks. Each sub-block may have a different motion vector in each direction. Thus, motion vectors exist at the sub-block level and above.
[0069] In AVC, for skip and / or direct modes in B slices, temporal direct mode can be enabled either at the MB level or at the MB partition level. For each MB partition, the motion vectors of the blocks co-located with the current MB partition in the current block's RefPicList1[0] are used to derive a motion vector. Each motion vector in the co-located block is scaled based on the POC distance.
[0070] A spatial direct mode may also be implemented in AVC, for example, in which the direct mode may also predict motion information from spatial neighbors.
[0071] As mentioned above, in HEVC, the largest coding unit in a slice is called a coding tree block (CTB). The CTB contains a quadtree, and the nodes of the quadtree are coding units. The size of the CTB can range from 16x16 to 64x64 in the HEVC Main Profile. In some cases, a CTB size of 8x8 can be supported. The coding units (CUs) may be the same size of the CTB or as small as 8x8. In some cases, each coding unit is coded using one mode. When a CU is inter-coded, it may be further partitioned into two or four prediction units (PUs), or there may be only one PU when no further partitioning is applied. When there are two PUs in one CU, they may be a rectangle with half the size, or two rectangles with 1 / 4 or 3 / 4 the size of the CU. When a CU is inter-coded, one set of motion information exists for each PU. In addition, each PU is coded with a unique inter prediction mode to derive a set of motion information.
[0072] For example, for motion prediction in HEVC, there are two inter-prediction modes, including merge mode and advanced motion vector prediction (AMVP) mode for a prediction unit (PU). Skip is considered as a special case of merge. In either AMVP mode or merge mode, a motion vector (MV) candidate list is maintained for multiple motion vector predictors. The motion vector of the current PU, as well as the reference index in merge mode, are generated by taking one candidate from the MV candidate list. In some examples, one or more scaling window offsets may be included in the MV candidate list along with the stored motion vector.
[0073] In an example where an MV candidate list is used for motion prediction of a block, the MV candidate list may be constructed separately by the encoding device and the decoding device. For example, the MV candidate list may be generated by the encoding device when encoding the block and by the decoding device when decoding the block. Information about the motion information candidates in the MV candidate list (e.g., information about one or more motion vectors, information about one or more LIC flags that may possibly be stored in the MV candidate list, and / or other information) may be signaled between the encoding device and the decoding device. For example, in a merge mode, index values for the stored motion information candidates may be signaled from the encoding device to the decoding device (e.g., in a syntax structure such as a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), a slice header, a supplemental enhancement information (SEI) message in the video bitstream, or transmitted separately from the video bitstream, and / or in other signaling). The decoding device may construct an MV candidate list and obtain one or more motion information candidates from the constructed MV candidate list to use for motion compensated prediction using the signaled reference or index. For example, the decoding device 112 may build an MV candidate list and use the motion vector from the indexed location (and possibly the LIC flag) for motion prediction of the block. In the case of the AMVP mode, in addition to the reference or index, a difference or residual value may be signaled as a delta. For example, in the case of the AMVP mode, the decoding device may build one or more MV candidate lists and apply the delta value to one or more motion information candidates obtained using the signaled index value in performing motion compensated prediction of the block.
[0074] In some examples, the MV candidate list includes up to five candidates for merge mode and two candidates for AMVP mode. In other examples, a different number of candidates for merge mode and / or AMVP mode may be included in the MV candidate list. A merge candidate may include a set of motion information. For example, the set of motion information may include a motion vector corresponding to both the reference picture list (list 0 and list 1) and a reference index. If a merge candidate is identified by a merge index, the reference picture is used for prediction of the current block, and the associated motion vector is determined. However, in AMVP mode, since the AMVP candidate includes only a motion vector for each possible prediction direction from either list 0 or list 1, a reference index needs to be explicitly signaled along with the MVP index for the MV candidate list. In AMVP mode, the predicted motion vector may be further refined.
[0075] Thus, a merge candidate corresponds to a complete set of motion information, while an AMVP candidate contains only one motion vector for a particular prediction direction and reference index. The candidates for both modes are derived in the same way from the same spatial and temporal neighboring blocks.
[0076] In some examples, the merge mode allows an inter-predicted PU to inherit one or more of the same motion vectors, prediction direction, and one or more reference picture indexes from an inter-predicted PU that includes motion data positions selected from one group of spatially adjacent motion data positions and two temporally co-located motion data positions. For the AMVP mode, one or more motion vectors of the PU may be predictively coded with respect to one or more motion vector predictors (MVPs) from an AMVP candidate list constructed by the encoder and / or decoder. In some cases, in unidirectional inter prediction of the PU, the encoder and / or decoder may generate a single AMVP candidate list. In some cases, in bidirectional prediction of the PU, the encoder and / or decoder may generate two AMVP candidates, one using the motion data of the spatial and temporal neighboring PUs from the forward prediction direction and one using the motion data of the spatial and temporal neighboring PUs from the backward prediction direction.
[0077] Candidates for both modes are derived from spatial and / or temporal neighboring blocks. For example, Figures 2A and 2B include conceptual diagrams illustrating spatial neighboring candidates. Figure 2A illustrates spatial neighboring motion vector (MV) candidates for merge mode. Figure 2B illustrates spatial neighboring motion vector (MV) candidates for AMVP mode. Although spatial MV candidates are derived from neighboring blocks for a particular PU (PU0), the method of generating candidates from blocks differs between merge mode and AMVP mode.
[0078] In the merge mode, the encoder and / or decoder may form a merge candidate list by considering merge candidates from various motion data positions. For example, as shown in FIG. 2A, up to four spatial MV candidates may be derived for spatially adjacent motion data positions indicated with numbers 0-4 in FIG. 2A. The MV candidates may be ordered in the merge candidate list in the order indicated by the numbers 0-4. For example, the positions and orders may include a left position (0), an upper position (1), a top right position (2), a bottom left position (3), and a top left position (4).
[0079] In the AVMP mode shown in Fig. 2B, the neighboring blocks are divided into two groups: a left group including blocks 0 and 1, and an upper group including blocks 2, 3 and 4. For each group, possible candidates among the neighboring blocks that refer to the same reference picture as indicated by the signaled reference index have the highest priority in being chosen to form the final candidate of the group. It may be that all the neighboring blocks do not contain motion vectors pointing to the same reference picture. Therefore, if such a candidate cannot be found, the first available candidate is scaled to form the final candidate, so that the difference in temporal distance can be compensated.
[0080] 3A and 3B include conceptual diagrams illustrating temporal motion vector prediction. If a temporal motion vector predictor (TMVP) candidate is enabled and available, it is added to the MV candidate list after the spatial motion vector candidate. The process of motion vector derivation for a TMVP candidate is the same for both merge mode and AMVP mode. However, in some cases, the target reference index for a TMVP candidate in merge mode may be set to 0 or may be derived from the target reference index of a neighboring block.
[0081] The location of the primary block for TMVP candidate derivation is the bottom right block outside the co-located PU, as shown in FIG. 3A as block "T", to compensate for the bias towards the upper and left blocks used to generate spatial neighboring candidates. However, if the block is currently located outside the CTB (or LCU) row or motion information is not available, the block is replaced with the center block of the PU. The motion vectors for TMVP candidates are derived from the co-located PU of the co-located picture, shown at the slice level. Similar to the temporal direct mode in AVC, the motion vectors of TMVP candidates may undergo motion vector scaling, which is performed to compensate for distance differences.
[0082] Other aspects of motion prediction are covered in the HEVC standard and / or other standards, formats, or codecs. For example, some other aspects of merge mode and AMVP mode are covered. Other aspects include motion vector scaling. For motion vector scaling, it can be assumed that the value of the motion vector is proportional to the distance of the pictures in presentation time. A motion vector relates two pictures, namely, a reference picture and a picture that contains the motion vector (i.e., a stored picture). When a motion vector is utilized to predict another motion vector, the distance between the stored picture and the reference picture is calculated based on a picture order count (POC) value.
[0083] For a motion vector to be predicted, both its associated stored picture and reference picture may be different. Therefore, a new distance (based on POC) is calculated. Then, the motion vector is scaled based on these two POC distances. For spatially adjacent candidates, the stored pictures for the two motion vectors are the same, but the reference pictures are different. In HEVC, the motion vector scaling is applied to both TMVP and AMVP for spatially and temporally adjacent candidates.
[0084] Another aspect of motion prediction includes the generation of artificial motion vector candidates. For example, if the motion vector candidate list is not complete, an artificial motion vector candidate is generated and inserted at the end of the list until all candidates are obtained. In merge mode, there are two types of artificial MV candidates: composite candidates that are derived only for B slices, and zero candidates that are used only for AMVP when the first type does not provide enough artificial candidates. For each pair of candidates that are already in the candidate list and have the necessary motion information, a bidirectional composite motion vector candidate is derived by combining the motion vector of the first candidate that refers to a picture in list 0 and the motion vector of the second candidate that refers to a picture in list 1.
[0085] In some implementations, a pruning process may be performed when adding or inserting a new candidate into the MV candidate list. For example, in some cases, MV candidates from different blocks may contain the same information. In such cases, storing duplicated motion information of multiple MV candidates in the MV candidate list may lead to redundancy and reduced efficiency of the MV candidate list. In some examples, the pruning process may eliminate or minimize redundancy in the MV candidate list. For example, the pruning process may include comparing a potential MV candidate to be added to the MV candidate list with MV candidates already stored in the MV candidate list. In one illustrative example, the horizontal displacement (Δx) and vertical displacement (Δy) of a stored motion vector (indicating the position of a reference block relative to the position of a current block) may be compared with the horizontal displacement (Δx) and vertical displacement (Δy) of a motion vector of a potential candidate. If the comparison reveals that the motion vector of the potential candidate does not match any of the one or more stored motion vectors, the potential candidate may not be considered as a candidate to be pruned and may be added to the MV candidate list. If a match is found based on this comparison, the potential MV candidate is not added to the MV candidate list to avoid inserting the same candidate. In some cases, to reduce complexity, instead of comparing each potential MV candidate with all existing candidates, only a limited number of comparisons are performed during the pruning process.
[0086] Some coding schemes such as HEVC support weighted prediction (WP), where a scaling factor (denoted by a), a shift number (denoted by s), and an offset (denoted by b) are used in motion compensation. If p(x,y) is a pixel value at position (x,y) of a reference picture, p'(x,y) = ((a*p(x,y) + (1 << (s-1))) >> s) + b is used as a predicted value in motion compensation instead of p(x,y).
[0087] When WP is enabled, for each reference picture of the current slice, a flag is signaled to indicate whether WP is applied to the reference picture. If WP is applied to a reference picture, a set of WP parameters (i.e., a, s, and b) is sent to the decoder and used for motion compensation from the reference picture. In some examples, to flexibly turn on / off WP for luma and chroma components, the WP flag and WP parameters are signaled separately for luma and chroma components. In WP, one and the same set of WP parameters is used for all pixels in one reference picture.
[0088] FIG. 4A is a diagram illustrating an example of reconstructed neighboring samples of a current block 402 and neighboring samples of a reference block 404 used for unidirectional inter prediction. A motion vector MV may be coded for the current block 402, and the MV may include a reference index to a reference picture list and / or other motion information for identifying the reference block 404. For example, the MV may include a horizontal component and a vertical component that provide an offset from a coordinate location in the current picture to a coordinate in the reference picture identified by the reference index. FIG. 4B is a diagram illustrating an example of reconstructed neighboring samples of a current block 422 and neighboring samples of a first reference block 424 and a second reference block 426 used for bidirectional inter prediction. In this case, two motion vectors MV0 and MV1 may be coded for the current block 422 to identify the first reference block 424 and the second reference block 426, respectively.
[0089] Bilateral matching (BM) is a technique that may be used to refine a pair of two initial motion vectors (e.g., a first motion vector MV0 and a second motion vector MV1). For example, BM may be performed by searching around a pair of initial motion vectors MV0 and MV1 to derive a refined motion vector (e.g., refined motion vectors MV0' and MV1'). The refined motion vectors MV0' and MV1' may subsequently be used to replace the first motion vector MV0 and the second motion vector MV1, respectively. The refined motion vector may be selected in the search as the motion vector identified in the search that minimizes the block matching cost.
[0090] In some examples, a block matching cost may be generated based on the similarity of two motion compensated predictors generated for two MVs. Exemplary criteria for the block matching cost include, but are not limited to, sum of absolute differences (SAD), sum of absolute transformed differences (SATD), sum of squared error (SSE), etc. The block matching cost criterion may also include a regularization term derived based on the MV difference between a current MV pair (e.g., the MV pair being considered for selection as refined motion vectors MV0′ and MV1′) and an initial MV pair (e.g., MV0 and MV1).
[0091] In some examples, one or more constraints may be applied to the MV difference terms MVD0 and MVD1 (e.g., MVD0=MV0'-MV0 and MVD1=MV1'-MV1). For example, in some cases, the constraints may be applied based on the assumption that MVD0 and MVD1 are proportional to the temporal distance (TD) between the current picture (e.g., current block) and the reference picture (e.g., reference block) pointed to by the two MVs. In some examples, the constraints may be applied based on the assumption that MVD0=-MVD1 (e.g., MVD0 and MVD1 are equal in magnitude but opposite in sign).
[0092] In some examples, an inter-predicted CU may be associated with one or more motion parameters. For example, in a versatile video coding standard (VVC), each inter-predicted CU may be associated with one or more motion parameters, which may include, but are not limited to, a motion vector, a reference picture index, and a reference picture list usage index. The motion parameters may further include additional information related to the coding features of VVC to be used for inter-predicted sample generation. The motion parameters may be signaled explicitly or implicitly. For example, when a CU is coded in skip mode, the CU is associated with one PU, has no significant residual coefficients, and has no coded motion vector delta or reference picture index.
[0093] In some aspects, a merge mode may be defined such that motion parameters for a current CU are obtained from neighboring CUs including spatial and temporal candidates. Additionally or alternatively, one or more merge modes may be defined based on additional schedules provided in the VVC standard. In some examples, the merge mode may be applied to any inter-predicted CU (e.g., the merge mode may be applied to more than just skip mode). In some examples, an alternative to the merge mode may include explicit transmission of one or more motion parameters. For example, a motion vector, a corresponding reference picture index for each reference picture list, a reference picture list usage flag, and other related information may be explicitly signaled for each CU.
[0094] In addition to the inter-coding features in HEVC, VVC includes several new refined inter-prediction coding techniques, including enhanced merge prediction, merge mode with motion vector difference (MMVD), Symmetric MVD (SMVD) signaling, affine motion compensated prediction, sub-block based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage: 1 / 16 luma sample MV storage and 8x8 motion field compression, bi-prediction with CU-level weight (BCW), bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), geometric partitioning mode (GPM), and combined inter and intra prediction (CIIP).
[0095] For enhanced merge prediction in VVC merge mode (e.g., referred to as normal or default merge mode), a merge candidate list may be constructed by including five types of candidates in order: spatial motion vector prediction (MVP) from spatially neighboring CUs, temporal MVP from co-located CUs, history-based MVP from a first-in-first-out (FIFO) table, pairwise average MVP, and zero MV. The size of the merge candidate list may be signaled in the sequence parameter set header. The maximum allowed size of the merge candidate list may be six (e.g., six entries or six candidates). For each CU coded in the merge mode, the index of the best merge candidate is coded using truncated unary (TU). In some examples, VVC may also support parallel derivation of merge candidate lists for all CUs in a certain size or area (e.g., as done in HEVC). The five aforementioned types of merge candidates, as well as the associated example derivation process for each category of merge candidates, are then described below.
[0096] FIG. 5 illustrates locations or positions of spatial merge candidates for use in processing a block according to some examples of the present disclosure. For example, FIG. 5 illustrates example locations of spatial merge candidates (also referred to as "spatial neighborhoods") A0, A1, B0, B1, and B2 for use in processing block 500 according to some examples of the present disclosure. The spatial neighborhoods A0, A1, B0, B1, and B2 are illustrated in FIG. 5 based on their relationship to block 500. The derivation of spatial merge candidates in VVC may be the same as in HEVC, with the positions of the first two merge candidates swapped. In some examples, the largest of the four merge candidates may be selected from the five spatial merge candidates (e.g., A0, A1, B0, B1, and B2) located at the positions illustrated in FIG. 5.
[0097] The order of derivation may be B0, A0, B1, A1, and B2. For example, merge candidate position B2 may be considered only when one or more CUs associated with positions B0, A0, B1, A1 are unavailable or are intra-coded. CUs associated with positions B0, A0, B1, or A1 may be unavailable because the CUs belong to different slices or tiles. In some aspects, after the merge candidate at position A1 is added, the addition of the remaining merge candidates may be subject to a redundancy check. The redundancy check may be performed such that merge candidates with the same motion information are excluded from the merge candidate list (e.g., to improve coding efficiency).
[0098] FIG. 6 illustrates an aspect of motion vector scaling 600 for a temporal merge candidate for use in processing a block, according to some examples of this disclosure. In some examples, a temporal merge candidate derivation may be performed, and one merge candidate (e.g., a temporal merge candidate) is added to a merge candidate list. The temporal merge candidate derivation may be performed based on a scaled motion vector. The scaled motion vector may be derived based on a co-located CU included in a co-located reference picture. The reference picture list to be used for derivation of the co-located CU may be explicitly signaled in the slice header.
[0099] For example, FIG. 6 illustrates a current picture 610 and a co-located picture 630, which may be associated with a current reference picture 615 and a co-located reference picture 635, respectively. FIG. 6 also illustrates a current CU 612 (e.g., associated with the current picture 610) and a co-located CU (e.g., associated with the co-located picture 630). In some examples, scaled motion vectors for derivation of temporal merge candidates may be derived or obtained as shown in FIG. 6. For example, FIG. 6 illustrates a dotted line 611, which is scaled from the motion vector of the co-located CU 632 using Picture Order Count (POC) distances tb and td. In some examples, tb is the POC difference between the current reference picture 615 and the current picture 610, and td is the POC difference between the co-located reference picture 635 and the co-located picture 630. The reference picture index of the temporal merge candidate may be set equal to 0.
[0100] 7 illustrates aspects of a temporal merge candidate 700 for use in processing a block, according to some examples of this disclosure. In some examples, after a single temporal merge candidate is selected as discussed above in connection with FIG. 6, the candidate position C 0 and C. 1 A location of a temporal merge candidate may be selected from among the locations C 0 If the CU in is not available, is intra-coded, or is outside the current row of the CTU, then the candidate position C 1 Otherwise, position C 0 is used.
[0101] In some aspects, a history-based motion vector prediction (HMVP) merge candidate may be added to the merge candidate list after the spatial MVP merge candidate (e.g., as described above with respect to FIG. 5) and the TMVP merge candidate (e.g., as described above with respect to FIG. 6 and FIG. 7). The HMVP merge candidate may be derived based on the motion information of a previously coded block. For example, the motion information of a previously coded block may be stored (e.g., in a table) and used as a motion vector prediction (MVP) for the current CU. In some examples, a table with multiple HMVP candidates may be maintained during the encoding and / or decoding process. When a new CTU row is encountered, the table is reset (e.g., emptied). Whenever there is a non-subblock inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0102] In some examples, the HMVP table size S may be set to a value of 6 (e.g., up to 6 history-based MVP (HMVP) candidates may be added to the HMVP table). When inserting a new HMVP candidate into the HMVP table, a constrained first-in-first-out (FIFO) rule may be utilized. The constrained FIFO rule may include a redundancy check that is applied to determine whether an identical HMVP is in the table (e.g., to determine that the newly inserted HMVP candidate is the same as an existing HMVP candidate in the table). If the redundancy check for the newly inserted HMVP candidate finds that the identical HMVP is already included in the table, the identical HMVP may be removed from the table and all subsequent HMVP candidates are moved forward.
[0103] In some aspects, the HMVP candidates (e.g., included in an HMVP list or HMVP table) may then be used to build or otherwise generate a merge candidate list. For example, to generate the merge candidate list, the most recent several HMVP candidates in the table may be reviewed in order and inserted into the merge candidate list after the TMVP candidates. A redundancy check may be applied to the HMVP candidates added to the merge candidate list, which is used to determine whether the HMVP candidate is the same or identical to any of the spatial or temporal merge candidates previously added to or already included in the merge candidate list.
[0104] In some examples, the number of redundancy check operations performed in connection with generating the merge candidate list and / or the HMVP table may be reduced. For example, the number of HMVP candidates used for merge list generation may be set as (N <= 4) ? M: (8 - N), where N indicates the number of existing candidates in the merge candidate list and M indicates the number of available HMVP candidates in the HMVP table. In other words, if the condition N <= 4 evaluates to true (e.g., the merge candidate list contains four or fewer candidates), all M HMVP candidates in the HMVP table are used for merge list generation. If the condition N <= 4 evaluates to false (e.g., the merge candidate list contains more than four candidates), 8-N HMVP candidates in the HMVP table are used for merge list generation. In some examples, the merge candidate list construction process from HMVP may be terminated when the total number of available merge candidates reaches the maximum allowed number of merge candidates minus one.
[0105] Derivation of a pairwise average merge candidate may be performed based on a predefined pair of merge candidates in an existing merge candidate list. For example, a pairwise average merge candidate may be generated by averaging a predefined pair of merge candidates in an existing merge candidate list, where the predefined pairs are given as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where these numbers represent merge indexes for the merge candidate list. In some cases, the averaged motion vector may be calculated separately for each reference list. If both motion vectors (e.g., of a predefined pair) are included in the same list, these two motion vectors may be averaged even when the two motion vectors point to different reference pictures. In some cases, if only one motion vector (e.g., of a predefined pair) is available, the one available motion vector may be used as the averaged motion vector. If a motion vector (e.g., of a predefined pair) is not available, the list may be identified as invalid. In some examples, as described above, if the merge candidate list is not full after the pairwise average merge candidates are added, one or more zero MVPs may be inserted at the end of the merge candidate list until the maximum number of merge candidates is reached.
[0106] In some aspects, bi-prediction with CU-level weight (BCW) may be utilized. For example, a bi-predictive signal may be generated by averaging two prediction signals obtained from two different reference pictures. Additionally or alternatively, a bi-predictive signal may be generated using two different motion vectors. In some examples, a bi-predictive signal may be generated by averaging two prediction signals obtained from two different reference signals and / or by using two different motion vectors using the HEVC standard.
[0107] In other aspects, the bi-prediction mode may be extended from the simple average to include a weighted average of two prediction signals. For example, the bi-prediction mode may include a weighted average of two prediction signals using the VVC standard. In some examples, the bi-prediction mode may include a weighted average of two prediction signals as follows: P bi-pred = ((8-w)*P 0 +w*P 1 +4)≫3 Formula (1)
[0108] The weighted average bi-prediction given in Equation (1) may include five weights w ∈ {-2, 3, 4, 5, 10}. For each bi-predicted CU, the weight w may be determined according to one or more of the following: In one example, for non-merged CUs, the weight index may be signaled after the motion vector differential (MVD). In another example, for merged CUs, the weight index may be inferred from the neighboring blocks based on the merge candidate index. In some cases, the BCW may be applied only to CUs with 256 or more luma samples (e.g., CUs with a product of the CU width and the CU height of 256 or more). In some cases, all five of the weights w may be used for low latency pictures. For non-low latency pictures, only three of the five weights w may be used (e.g., three weights w ∈ {3, 4, 5}).
[0109] At the encoder side, a fast search algorithm may be applied to find the weight index without significantly increasing the encoder complexity, as described in more depth below. In some examples, a fast search algorithm may be applied to find the weight index, as described at least in part in JVET-L0646. For example, when combined with adaptive motion vector resolution (AMVR), unequal weights are only conditionally checked for 1-pel and 4-pel motion vector precision if the current picture is a low-latency picture. When combined with affine mode, affine motion estimation (ME) may be performed for unequal weights only if the affine mode is selected as the current best mode. When two reference pictures in bi-predictive mode are the same, unequal weights may only be conditionally checked. In some cases, unequal weights are not searched for when certain conditions are met (e.g., depending on the picture order count (POC) distance between the current picture and its reference picture, the coding quantization parameter (QP), and / or the temporal level, etc.).
[0110] In some examples, the BCW weight index may be coded using one context-coded bin followed by one or more bypass-coded bins. For example, the first context-coded bin may be used to indicate whether equal weights are used. If unequal weights are used, additional bins may be signaled using bypass coding to indicate which unequal weights are used.
[0111] Weighted prediction (WP) is a coding tool supported by H.264 / AVC and HEVC standards for efficiently coding video content with fading. Support for WP has also been added to the VVC standard. WP may be used to allow one or more weighting parameters (e.g., weights and offsets) to be signaled for each reference picture in each of the reference picture lists L0 and L1. Subsequently, the weights and offsets of the corresponding reference pictures are applied during motion compensation.
[0112] The WP and BCW may be utilized with different types of video content. In some examples, when a CU utilizes a WP, the BCW weight index may not be signaled and w may be inferred to have a value of 4 (e.g., equal weights are applied). For example, to avoid interactions between the WP and BCW (e.g., which may complicate the design of a VVC decoder), the BCW weight index may not be signaled when a CU utilizes a WP.
[0113] For a merge CU, a weight index may be inferred from neighboring blocks in both normal merge mode and inherited affine merge mode based on the merge candidate index. In constructed affine merge mode, affine motion information may be constructed based on the motion information of up to three blocks. For example, the BCW index for a CU using constructed affine merge mode may be set equal to the BCW index of the first control point MV. In some examples, using the VVC standard, synthetic inter and intra prediction (CIIP) and bi-prediction with CU level weight (BCW) cannot be applied together for a CU. When a CU is coded using a CIIP mode, the BCW index of the current CU may be set to a value of 2 (e.g., set to equal weight).
[0114] FIG. 8 is a diagram 800 illustrating an aspect of bilateral matching according to some examples of the present disclosure. As previously mentioned, bilateral matching (BM) may be used to refine a pair of two initial motion vectors MV0 and MV1. For example, BM may be performed by searching around MV0 and MV1 to derive refined motion vectors MV0′ and MV1′, respectively, that minimize the block matching cost. The block matching cost may be calculated based on the similarity of two motion compensated predictors generated using a pair of initial motion vectors (e.g., MV0 and MV1). For example, the block matching cost may be based on a sum of absolute differences (SAD). Additionally or alternatively, the block matching cost may be based on or include a regularization term based on the motion vector difference (MVD) between the current MV pair (e.g., MV0′ and MV1′ currently being tested) and the initial MV pair (e.g., MV0 and MV1). As described in more depth below, one or more constraints may be applied based on MVD0 (eg, MVD0=MV0'-MV0) and MVD1 (eg, MVD1=MV1'-MV1).
[0115] As previously mentioned, in the versatile video coding standard (VVC), bilateral matching-based decoder-side motion vector refinement (DMVR) may be applied to increase (e.g., refine) the accuracy of the MV of a bi-predictive merge candidate. For example, as shown in the example of FIG. 8, bilateral matching-based DMVR may be applied to increase or otherwise refine the accuracy of the MV of a bi-predictive merge candidate 812. The bi-predictive merge candidate 812 may be included in a current picture 810 and may be associated with an initial pair of motion vectors MV0 and MV1. Prior to performing the bilateral matching-based DMVR, initial motion vectors MV0 and MV1 may be obtained, identified, or otherwise determined for the bi-predictive merge candidate 812. Subsequently, as described in more depth below, the initial motion vectors MV0 and MV1 may be used to identify refined motion vectors MV0′ and MV1′ for the bi-predictive merge candidate 812.
[0116] As shown in the example of Figure 8, a first initial motion vector MV0 may point to a first reference picture 830. The first reference picture 830 may be associated with a backward direction (e.g., with respect to the current picture 810) and / or may be included in a reference picture list L0. A second initial motion vector MV1 may point to a second reference picture 820. The second reference picture 820 may be associated with a forward direction (e.g., with respect to the current picture 810) and / or may be included in a reference picture list L1.
[0117] The first initial motion vector MV0 may be used to determine or generate a first predictor 832, which may be a block included in the first reference picture 830. The first predictor 832 may also be referred to as a first candidate block (e.g., the first predictor 832 is a candidate block in the first reference picture 830 and / or the reference picture list L0). The second initial motion vector MV1 may be used to determine or generate a second predictor 822, which may be a block included in the second reference picture 820. The second predictor 822 may also be referred to as a second candidate block (e.g., the second predictor 822 is a candidate block in the second reference picture 820 and / or the reference picture list L1).
[0118] Subsequently, a search may be performed around the first predictor 832 and the second predictor 822 to identify or determine a first refined motion vector MV0′ and a second refined motion vector MV1′, respectively. As shown, a surrounding area associated with each of the first predictor 832 (e.g., an area in the first reference picture 830 and / or reference picture list L0) and a surrounding area associated with the second predictor 822 (e.g., an area in the reference picture 820 and / or reference picture list L1) may be searched. For example, a surrounding area associated with the first predictor 832 may be searched to identify or probe one or more refined candidate blocks 834, and a surrounding area associated with the second predictor 822 may be searched to identify or probe one or more refined candidate blocks 824. The search may be performed based on one or more of a distortion (e.g., SAD, SATD, SSE, etc.) and / or regularization term calculated between one of the initial predictors 832 or 822 and the corresponding refined candidate block 834 or 824. In some examples, the distortion and / or regularization may be calculated based on a distance traveled from an initial point (e.g., the distance between an initial point associated with the initial predictor 832 or 822 and a searched point associated with the refined candidate block 834 or 824, respectively).
[0119] As the search moves around the initial point associated with each of the initial predictors 832 and 822, new refined candidate blocks (e.g., refined candidate blocks 834 and 824) are obtained. Each new refined candidate block may be associated with a new cost (e.g., the calculated SAD value, one of the motion vector difference MVD0 or MVD1, etc.). The search may be associated with a search range and / or search interval. After searching each candidate block included in the search range and / or search window for the initial predictors 832 and 822 (e.g., and determining a corresponding cost for each searched candidate block), the candidate block with the lowest determined cost may be identified and used to generate refined motion vectors MV0′ and MV1′.
[0120] In some examples, bilateral matching (BM) based DMVR may be performed by calculating the SAD between two candidate blocks in the reference picture list L0 and list L1. As shown in FIG. 8, the SAD between blocks based on each MV' candidate (e.g., blocks 834 and 824) around the initial MV (e.g., around predictors 832 and 822, respectively) may be calculated. The MV' candidate with the smallest SAD may be selected as the refined MV and used to generate the bi-predicted signal. In some examples, the SAD of the initial MV is subtracted by 1 / 4 of the SAD value to become a regularization term. In some cases, the temporal distances (e.g., picture order count (POC) difference) from the two reference pictures to the current picture may be the same, and MVD0 and MVD1 may be of the same magnitude but opposite sign (e.g., MVD0=-MVD1).
[0121] In some cases, the bilateral matching-based DMVR may be performed using a refinement search range of two integer luma samples from the initial MV. For example, in the context of FIG. 8, the bilateral matching-based DMVR may be performed using a refinement search range of two integer luma samples from the initial motion vectors MV0 and MV1 (e.g., from the initial predictors 832 and 822, respectively). This search may include an integer sample offset search stage and a fractional sample refinement stage.
[0122] In some cases, a 25-point full search may be applied for the integer sample offset search. The 25-point full refinement search may be performed by first calculating the SAD of an initial MV pair (e.g., an initial MV pair of MV0 and MV1 and / or the initial predictors 832 and 822). If the SAD of the initial MV pair is less than a threshold, the integer sample stage of the DMVR may be terminated. If the SAD of the initial MV pair is not less than a threshold, the SAD of the remaining 24 points may be calculated and checked in raster scan order. The point with the smallest SAD may then be selected as the output of the integer sample offset search stage.
[0123] As mentioned above, the integer sample offset search may be followed by fractional sample refinement. In some cases, computational complexity may be reduced by deriving the fractional sample refinement using one or more parametric error surface equations (e.g., rather than performing an additional search with a SAD comparison). The fractional sample refinement may be conditionally invoked based on the output of the integer sample offset search stage. For example, when the integer sample offset search stage ends with a center having the smallest SAD in either the first or second iteration search, non-integer sample refinement may be further applied in response thereto.
[0124] As mentioned above, a parametric error surface may be used to derive the fractional sample refinement. For example, in a parametric error surface based sub-pixel offset estimation, the center location cost and the costs at the four neighboring locations (e.g., relative to the center) may be used to fit a 2D parabolic error surface equation of the form: E(x,y) = A(x - x min ) 2 + B(y - y min ) 2 + C formula (2)
[0125] Here, (x min ,y min ) corresponds to the lowest cost fractional position, and C corresponds to the minimum cost value. By solving equation (2) using the cost values of five search points (e.g., the center position and four adjacent positions), (x min , y min ) can be calculated as follows: x min = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0))) Equation (3) y min = (E(0,-1) - E(0,1)) / (2((E(0,-1) + E(0,1) - 2E(0,0))) Equation (4)
[0126] x min and y min The values of may be automatically constrained to be between -8 and 8 (e.g., all cost values are positive and the minimum is E(0,0)), which may correspond to half-pel offsets in 1 / 16-pel MV precision in VVC. min ,y min ) may be added to the integer distance refinement MV (eg, from the integer sample offset search described above) to obtain or determine a sub-pixel precision refinement delta MV.
[0127] In VVC, the motion vector resolution may be 1 / 16 luma samples. In some examples, samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, refinement search points may surround the initial fractional pel MV with integer sample offsets, and the DMVR search process may include interpolating samples at fractional positions. In some examples, a bilinear interpolation filter may be used to generate fractional samples for the DMVR search process. The use of a bilinear interpolation filter to generate fractional samples for the DMVR search process can reduce computational complexity. In some cases, when a bilinear interpolation filter is utilized with a search range of two samples, the DMVR search process may not need to access additional reference samples (e.g., compared to existing motion compensation processes).
[0128] After the refined MV is determined using the DMVR search process, an 8-tap interpolation filter may be applied to generate a final prediction. In some examples, additional reference samples may not be required to perform the interpolation process based on the original MV (e.g., as described above). In some examples, additional samples may be utilized to perform the interpolation based on the refined MV. Rather than accessing additional reference samples (e.g., accessing additional reference samples compared to an existing motion compensation process), existing or available reference samples may be padded to generate additional reference samples for performing the interpolation based on the refined MV. In some cases, when the width and / or height of the CU is greater than 16 luma samples, as described in more depth below, the DMVR process may further include dividing the CU into sub-blocks, each having a width and / or height equal to 16 luma samples.
[0129] In VVC, DMVR may be applied to CUs coded with one or more of the following modes or features. In one illustrative example, the modes or features associated with CUs to which DMVR may be applied may also be referred to as "DMVR conditions." In some examples, the DMVR conditions may include, but are not limited to, CUs coded with one or more of the following modes and / or features: CU level merge mode with bi-predictive MV, one reference picture in the past (e.g., relative to the current picture) and another reference picture in the future (e.g., relative to the current picture), the distance (e.g., POC distance) from the two reference pictures to the current picture is the same, both reference pictures are short-term reference pictures, CUs that include more than 64 luma samples, both the CU height and the CU width are 8 luma samples or more, the BCW weight index indicates equal weights, weighted prediction (WP) is not enabled for the current block, a combined inter- and intra-prediction (CIIP) mode is not utilized for the current block, etc.
[0130] FIG. 9 illustrates an example extended CU region 900 that may be utilized to perform bidirectional optical flow (BDOF) according to some examples of the present disclosure. For example, BDOF may be used to refine the bi-predictive signal of luma samples in a CU at a 4×4 sub-block level (e.g., 4×4 sub-block 910). In some cases, BDOF may be used for sub-blocks of other sizes (e.g., 8×8 sub-blocks, 4×8 sub-blocks, 8×4 sub-blocks, 16×16 sub-blocks, and / or sub-blocks of other sizes). The BDOF mode may be based on optical flow that assumes that object motion is smooth. As shown in FIG. 9, the BDOF mode may utilize one extended row and column (e.g., extended row 970 and extended column 980) around the boundary of the extended CU region 900.
[0131] For each 4×4 sub-block (e.g., 4×4 sub-block 910), motion refinement (v x ,v y ) can be calculated. x ,v y ) may then be used to adjust the bi-predicted sample values within a 4×4 sub-block (e.g., 4×4 sub-block 910). In one illustrative example, BDOF may be performed as described below.
[0132] First, the horizontal and vertical gradients of the two prediction signals
[0133]
number
[0134] and
[0135]
number
[0136] (k=0,1) can each be calculated by directly calculating the difference between two adjacent samples.
[0137]
number
[0138] Here I (k) (i,j) represents the sample value at coordinate (i,j) of the prediction signal in list k, where k = 0, 1. Since shift1 may be set equal to 6, shift1 may be calculated based on the luma bit depth (e.g., bitDepth).
[0139] Then, the autocorrelation and cross-correlation of the gradients S 1 , S 2 , S 3 , S 5, and S 6 can be calculated as follows: S 1 =Σ (i,j)∈Ω │ψ x (i,j)│ Equation (7) S 2 =Σ (i,j)∈Ω ψ x (i,j)·sign(ψ y (i,j)) Equation (8) S 3 =Σ (i,j)∈Ω θ(i,j)·(-sign(ψ x (i,j))) Equation (9) S 5 =Σ (i,j)∈Ω │ψ y (i,j)│ Equation (10) S 6 =Σ (i,j)∈Ω θ(i,j)·(-sign(ψ y (i,j))) Equation (11)
[0140] Where:
[0141]
number
[0142] θ(i,j)=(I (0) (i,j)≫shift2)-(I (1) (i,j)≫shift2) Equation (14).
[0143] Here, Ω may be a 6×6 window (or other size window) around the 4×4 sub-block 910 (or other size sub-block). The value of shift2 may be set equal to 4 (or other appropriate value) and the value of shift3 may be set equal to 1 (or other appropriate value).
[0144] Movement elaboration (v x ,v y ) can then be derived using cross-correlation and autocorrelation terms using:
[0145]
number
[0146] Here, th' BIO = 1 << 4,
[0147]
number
[0148] is the floor function.
[0149] Based on the motion refinement and gradients, the following adjustments may be calculated for each sample in the 4×4 sub-block 910 (or other sized sub-block):
[0150]
number
[0151] Finally, the BDOF samples for the extended CU region 900 shown in FIG. 9 may be calculated by adjusting the bi-predictive samples as follows: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset )≫shift5 Formula (18)
[0152] Here, shift5 may be set equal to Max(3,15 - BitDepth), and the variable ο offset may be set equal to (1 << (shift5 - 1)).
[0153] In some cases, the above values may be selected such that the multipliers in the BDOF process (e.g., described above) do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process (e.g., described above) is kept within 32 bits.
[0154] In some examples, the gradient value is calculated for one or more predicted samples I in a list k, where k=0,1. (k) (i,j), where one or more prediction samples are currently outside the boundary of the CU. As shown in FIG. 9 and described above, the BDOF can utilize one extended row and one extended column (e.g., extended row 970 and extended column 980) around the boundary of the extended CU region 900. In some cases, the computational complexity of generating out-of-bounds prediction samples can be controlled at least in part based on generating prediction samples within the extended area (e.g., non-shaded blocks included in extended row 970 and extended column 980 along the perimeter of the extended CU region 900) by directly using reference samples at nearby integer positions without interpolation. For example, a floor() operation can be used on the coordinates of the reference samples in the nearby reference samples. An 8-tap motion compensated interpolation filter may then be used to generate prediction samples within the shaded box of the extended CU region 900 (e.g., the box inside the unshaded perimeter of the extended area prediction samples of the extended row 970 and extended column 980). In some examples, the extended sample values may be used only in the gradient calculation. The remaining steps of the BDOF process may be performed, as necessary, by padding (e.g., repeating) any samples and gradient values outside the boundary of the extended CU region 900 based on their nearest neighbors.
[0155] In some examples, BDOF may be used to refine the bi-predictive signal of a CU at a 4×4 sub-block level (or other sub-block level), as mentioned above. In one illustrative example, BDOF may be applied to a CU if the CU satisfies some or all of the following conditions: the CU is coded using a "true" bi-predictive mode (e.g., one of two reference pictures is located before the current picture in display order and the other reference picture is located after the current picture in display order), the CU is not coded using an affine mode or an SbTMVP merge mode, the CU has more than 64 luma samples, both the CU height and the CU width are 8 luma samples or more, the BCW weight index indicates equal weights, the WP is not currently valid for the CU, and / or the CIIP mode is not currently used for the CU, etc.
[0156] In some aspects, multi-pass decoder-side motion refinement may be used. For example, at the JVET-V meeting, the Enhanced Compression Model (ECM) was established to study compression techniques beyond VVC (https: / / vcgit.hhi.fraunhofer.de / ecm / VVCSoftware_VTM / - / tree / ECM). In the ECM, "Multi-pass Decoder-side Motion Refinement" [JVET U0100] is adopted to replace DMVR in VVC. Multi-pass decoder-side motion refinement may include multiple passes. In some examples, multi-pass decoder-side motion refinement may be performed by applying bilateral matching (BM) multiple times. For example, in each pass (e.g., of multi-pass decoder-side motion refinement), BM may be applied to different block sizes.
[0157] In a first pass, bilateral matching (BM) may be applied to entire coding blocks or coding units (e.g., similar to those of DMVR in VVC and / or as described above). For example, first pass BM may be applied to coding blocks or CUs with sizes of 64×64, 64×32, 32×32, 32×16, 16×16, etc., and combinations thereof.
[0158] In the second pass, the BM may be applied again, this time to each 16×16 sub-block contained within (or generated from) the entire coding block on which the BM was performed in the first pass. For example, if the BM of the first pass was applied to a 64×64 block, the 64×64 block may be divided into a 4×4 grid of sub-blocks, each of which is 16×16. In the second pass, in the above example, the BM may be applied to each of the 16 sub-blocks. In another example, if the BM of the first pass was applied to a 32×32 block, the 32×32 block may be divided into a 2×2 grid of sub-blocks, each of which is 16×16 in size. In this example, the second pass may be performed by applying the BM to each of the four 16×16 sub-blocks of the 32×32 block. The refined MV generated or obtained from the BM of the first pass may be utilized as the initial MV for each 16×16 sub-block on which the BM of the second pass is performed.
[0159] In the third pass, multiple 8×8 sub-blocks may be obtained or generated either based on the original block (e.g., from the first BM pass) and / or based on the 16×16 sub-blocks (e.g., from the second BM pass). In some examples, each 16×16 sub-block from the second BM pass may be divided into a 2×2 grid of sub-blocks of size 8×8. In the third BM pass, one or more MVs associated with each 8×8 sub-block may be further refined by applying bidirectional optical flow (BDOF).
[0160] In one illustrative example, multi-pass decoder-side motion refinement may be performed for a 64×64 block or CU. In a first pass, a BM may be applied for the 64×64 block to generate or obtain a pair of refined MVs. In a second pass, the 64×64 block may be divided into 16 sub-blocks, each of which is 16×16 in size. In the second pass, a BM may be applied again to each of the 16 sub-blocks, this time using the refined MV from the first pass as the initial MV. In a third pass, each 16×16 sub-block may be divided into four sub-blocks of size 8×8 (e.g., a total of 16*4=64 8×8 sub-blocks for the original 64×64 block). In the third pass, one or more MVs associated with each 8×8 sub-block may be refined by applying BDOF, as previously described.
[0161] Examples of first, second, and third passes that may be included or utilized to perform multi-pass decoder-side motion refinement are described below.
[0162] In some examples, the first pass may include performing block-based bilateral matching (BM) motion vector (MV) refinement. For example, in the first pass, a refined MV is derived by applying a BM to a coding block. Similar to decoder-side motion vector refinement (DMVR), in bi-predictive operations involving BM, the refined MV is searched around two initial MVs (e.g., MV0 and MV1) in reference picture lists L0 and L1. The refined MV (e.g., MV0) is derived by applying a BM to a coding block. pass1 and MV1 pass1 ) is derived around a pair of initial MVs (eg, MV0 and MV1, respectively) based on the lowest bilateral matching cost between two reference blocks in L0 and L1.
[0163] The first pass BM may include performing a local search to derive the integer sample precision intDeltaMV. The local search may be performed by repeatedly applying a 3×3 rectangular search pattern (or other search pattern) to the search range [−sHor, sHor] horizontally and to the search range [−sVer, sVer] vertically. The values of sHor and sVer may be determined by the dimensions of the block. In some cases, the maximum value of sHor and / or sVer may be 8 (or other suitable value).
[0164] The bilateral matching cost can be calculated as follows: bilCost = mvDistanceCost + sadCost Formula (19)
[0165] When the block size cbW*cbH is larger than 64 (or other block size threshold), a mean-removed sum of absolute difference (MRSAD) cost function may be applied to remove the DC contribution of distortion between reference blocks. When bilCost at the center point of the 3×3 search pattern (or other search pattern) has the lowest cost, the intDeltaMV local search ends. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern (or other search pattern) and the local search continues, searching for the minimum cost until the end of the search range (e.g., [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction) is reached.
[0166] In some cases, the existing fractional sample refinement may be further applied to derive the final deltaMV. The refined MV after the first pass is then MV0 pass1 = MV0 + deltaMV Equation (20) MV1 pass1 = MV1 - deltaMV Equation (21) can be derived as follows.
[0167] In some examples, the second pass may include performing sub-block-based bilateral matching (BM) motion vector (MV) refinement. For example, in the second pass, refined MVs may be derived by applying BM to 16×16 (or other size) grid sub-blocks. For each sub-block, the refined MV is derived by applying BM to two MVs (e.g., MV0, MV1, MV2, MV3, MV4, MV5, MV6, MV7, MV8, MV9, MV10, MV11, MV12, MV13, MV14, MV15, MV16, MV17, MV18, MV19, MV20, MV21, MV22, MV23, MV24, MV25, MV26, MV27, MV28, MV29, MV30, MV31, MV32, MV33, MV34, MV35, MV36, MV37, MV38, MV39, MV40, MV41, MV42, MV43, MV44, MV45, MV46, MV47, MV48, MV49, MV50, MV51, MV52, MV53, MV54, MV55, MV56, MV57, MV58, MV59, MV60, MV61, MV62, MV63, MV64, MV65, MV66, MV67, MV68, MV69, MV70, MV71, MV72, MV73, MV74, MV75, MV76, MV77, MV78, MV79, MV80, MV81, MV82, MV83, MV84, pass1 and MV1 pass1 ) is explored around.
[0168] Based on the exploration, we have two refined MVs: MV0 pass2 (sbIdx2) and MV1 pass2 (sbIdx2) may be derived based on the lowest bilateral matching cost between two reference sub-blocks in L0 and L1, where sbIdx2=0,...,N-1 is an index of the sub-block (e.g., since the second pass BM may be applied to each sub-block generated from the original block used in the first pass BM). For example, as previously described, a total of 16 sub-blocks each having a dimension of 16x16 may be generated or obtained for an input block of size 64x64. In this example, sbIdx2 may be an index to a respective one of the 16 sub-blocks.
[0169] For each subblock, the second pass BM may include performing a full search to derive the integer sample precision intDeltaMV. The full search may have a search range in the horizontal direction [-sHor, sHor] and a search range in the vertical direction [-sVer, sVer]. The values of sHor and sVer may be determined by the dimensions of the block, and the maximum value of sHor and sVer may be 8 (or other suitable value).
[0170] 10 is a diagram illustrating example search area regions in a coding unit (CU) 1000 according to some examples of this disclosure. For example, FIG. 10 illustrates four different search regions (e.g., a first search region 1020, a second search region 1030, a third search region 1040, a fourth search region 1050, etc.) in the coding unit 1000. In some cases, the bilateral matching cost may be calculated by applying a cost factor to the SATD cost (or other cost function) between two reference blocks, as follows: bilCost = satdCost * costFactor Formula (22)
[0171] As shown in FIG. 10, the search area (2*sHor+1)*(2*sVer+1) may be divided into five diamond-shaped search regions. In other embodiments, other search regions may be used. Each search region is assigned a costFactor value, which is determined by the distance (intDeltaMV) between each search point and the starting MV. Each diamond-shaped search region (e.g., search regions 1020, 1030, 1040, and 1050) may be processed in order starting from the center of the search area. Within each search region, the search points are processed in raster scan order starting from the top left corner of the region toward the bottom right corner.
[0172] In some examples, the first search area 1020 is searched first, followed by the second search area 1030, followed by the third search area 1040, followed by the fourth search area 1050. The full search of integer pels ends when the best bilCost in the current search area is less than a threshold equal to sbW*sbH. Otherwise, the full search of integer pels continues with the next search area until all search points have been examined.
[0173] In some cases, VVC DMVR fractional sample refinement may be further applied to derive the final deltaMV(sbIdx2). The refined MV in the second pass may then be derived as follows: MV0 pass2 (sbIdx2) = MV0 pass1+ deltaMV(sbIdx2) Equation (23) MV1 pass2 (sbIdx2) = MV1 pass1 - deltaMV(sbIdx2) Equation (24)
[0174] In some examples, the third pass may include performing sub-block-based bidirectional optical flow (BDOF) motion vector (MV) refinement. For example, in the third pass, refined MVs may be derived by applying BDOF to 8×8 (or other size) grid sub-blocks. For each 8×8 sub-block, BDOF refinement may be applied to derive scaled Vx and Vy without clipping, starting from the refined MV of the parent sub-block of the second pass. For example, each parent sub-block of the second pass may be 16×16 in size (e.g., each parent sub-block of the second pass may be associated with four 8×8 sub-blocks used in the third pass). The third pass BDOF refinement may be applied to derive a scaled Vx and Vy without clipping, starting from the refined MV of the parent sub-block of the second pass. For example, each parent sub-block of the second pass may be 16×16 in size (e.g., each parent sub-block of the second pass may be associated with four 8×8 sub-blocks used in the third pass). pass2 (sbIdx2) and MV1 pass2 (sbIdx2), where sbIdx2 is the index of one of the parent sub-blocks of the second pass.
[0175] The derived bioMv(Vx, Vy) may then be rounded to 1 / 16 sample precision and clipped between -32 and 32 (or other sample precision and / or other clipping value or range).
[0176] The refined MV in the third pass (MV0 pass3 (sbIdx3) and MV1 pass3 (sbIdx3) can be derived as follows: MV0 pass3 (sbIdx3) = MV0 pass2 (sbIdx2) + bioMv formula (25) MV1 pass3 (sbIdx3) = MV0 pass2(sbIdx2) - bioMv formula (26)
[0177] Here, sbIdx3 is an index to a particular one of the 8×8 sub-blocks used in the third pass, and sbIdx2 is an index to a particular one of the 16×16 sub-blocks used in the second pass. In some examples, sbIdx3 may have a range such that each 8×8 sub-block is uniquely identifiable by its sbIdx3 index. For example, a 64×64 block may be divided into 64 8×8 sub-blocks, and sbIdx3 may include 64 unique index values for different 8×8 sub-blocks. In some examples, sbIdx3 may have a range such that each 8×8 sub-block is uniquely identifiable by its sbIdx3 in combination with the sbIdx2 value of the corresponding 16×16 parent block. For example, a 64×64 block may be divided into a total of 16 sub-blocks, each having a size of 16×16, and each 16×16 sub-block may be further divided into four sub-blocks, each having a size of 8×8. In this example, sbIdx2 can take on one of 16 unique values and sbIdx3 can take on one of four unique values, so that each of the 64 8x8 sub-blocks is identifiable based on the corresponding sbIdx2 and sbIdx3 indices.
[0178] Aspects for improving the search strategy mentioned above are described herein. The aspects described herein may be applied to one or more coding (e.g., encoding, decoding, or composite encoding-decoding) techniques, such as one or more coding techniques in which a block of a current picture is predicted from one or more reference pictures using a motion vector (e.g., two reference blocks from two respective reference pictures), and the motion vector is refined by a refinement technique. These improvements may be applied to any suitable video coding standard or format (e.g., HEVC, VVC, AV1) as described above, other existing standards or formats that apply coding for a block based on a reference block from two respective reference pictures, and any future standard that uses such techniques. In general, when two reference blocks from two reference pictures are used, such techniques are generally referred to as bi-predicted merge mode and bilateral matching techniques.
[0179] Rather than strictly following the search strategies described above, various aspects described herein may enable different coded blocks to have different search strategies (or methods) for bilateral matching. The selected search strategy for a block may be signaled in one or more syntax elements coded in the bitstream. The search strategy includes constraints / relationships between MVD0 and MVD1 imposed during the bilateral matching search process, and may also be associated with some combination of search patterns, search ranges or maximum search times, cost criteria, etc. In some cases, a constraint may also be referred to as a constraint.
[0180] The systems and techniques described herein may be utilized to perform adaptive bilateral matching (BM) for decoder-side motion vector refinement (DMVR). For example, the systems and techniques may perform adaptive bilateral matching for DMVR by applying different search strategies and / or search methods for different coded blocks. In some aspects, a selected search strategy for a block may be signaled using one or more syntax elements coded in the bitstream. In some examples, a selected search strategy for a block may be signaled explicitly, implicitly, or using a combination thereof. As described in more depth below, a selected search strategy (e.g., to be used to perform adaptive BM for DMVR for a given block) may include a constraint or relationship between MVD0 and MVD1, and the search strategy constraint is utilized or applied during the bilateral matching search process. In some examples, additionally or alternatively, a search strategy may include one or more combinations of a search pattern, a search range, a maximum number of searches, a cost criterion, and the like.
[0181] In one illustrative example, the systems and techniques described herein may perform adaptive bilateral matching for DMVR using one or more motion vector differential (MVD) constraints. As previously described, a motion vector differential may be used to represent the difference between an initial motion vector and a refined motion vector (e.g., MVD0=MV0′-MV0 and MVD1=MV1′-MV1).
[0182] In some examples, an MVD constraint may be selected (e.g., included in a selected search strategy) for a given bilateral matching block. For example, the MVD constraint may be a mirroring constraint, where MVD0 and MVD1 are of the same magnitude but opposite sign (e.g., MVD0=-MVD1). The MVD mirroring constraint may also be referred to herein as a "first constraint."
[0183] In another example, the MVD constraint can be set to MVD0=0 (e.g., both x and y components of MVD0 are 0). For example, the MVD0=0 constraint may be utilized by keeping MV0 fixed while searching around MV1 to derive a refined MV1', where MV0'=MV0. The MVD0=0 constraint may also be referred to herein as the "second constraint."
[0184] In another example, the MVD constraint can be set to MVD1=0 (e.g., both x and y components of MVD1 are 0). For example, the MVD1=0 constraint may be utilized by keeping MV1 fixed while searching around MV0 to derive a refined MV0', where MV1'=MV1. The MVD1=0 constraint may also be referred to herein as the "third constraint."
[0185] In another example, the MVD constraint may be utilized to independently search around MV0 to derive MV0' and to independently search around MV1 to derive MV1'. The MVD independent search constraint may also be referred to herein as the "fourth constraint."
[0186] As described in more detail below, in some cases, only the first constraint and the second constraint may be provided as options for each bilateral matching block (e.g., included in the search strategy selected or signaled). In some cases, only the first constraint and the third constraint are provided as options for each bilateral matching block (e.g., included in the search strategy selected or signaled). In some cases, only the first constraint and the fourth constraint are provided as options for each bilateral matching block (e.g., included in the search strategy selected or signaled). In some cases, only the second constraint and the third constraint are provided as options for each bilateral matching block (e.g., included in the search strategy selected or signaled). In some cases, the first constraint, the second constraint, and the third constraint are provided as options for each bilateral matching block (e.g., included in the search strategy selected or signaled). Any other combination of constraints may be provided as options for each bilateral matching block (e.g., included in the search strategy selected or signaled).
[0187] In some aspects, one or more syntax elements may be signaled (e.g., in a bitstream) that include a value to indicate whether one or more of the constraints apply. In some examples, the syntax element may be used to indicate or determine a particular one of the above MVD constraints to apply to a given bilateral matching block.
[0188] In one illustrative example, the first syntax element may be used to indicate whether a first constraint applies. For example, the first syntax element may be used to indicate whether a mirroring constraint (e.g., MVD0=-MVD1) should be applied to a given bilateral matching block for which the first syntax element is signaled. In some cases, the first syntax element may have a first value when a mirroring constraint should be applied and a second value when a mirroring constraint should not be applied. In some cases, the presence of the first syntax element may be used to infer (e.g., implicitly signal) that a mirroring constraint should be applied to a given or current bilateral matching block, while the absence of the first syntax element may be used to infer (e.g., implicitly signal) that a mirroring constraint should not be applied.
[0189] Continuing with the above example, if the first syntax element indicates that the first constraint (e.g., MVD mirroring constraint, MVD0=-MVD1) does not apply, the second syntax element may be used to indicate which of the remaining constraints should be applied. For example, the second syntax element may be used to indicate whether the second constraint (e.g., MVD0=0) or the third constraint (e.g., MVD1=0) should be applied to the given or current bilateral matching block for which the syntax element is signaled. In some examples, the second syntax element may have a first value when the second constraint should be applied and a second value when the third constraint should be applied. In some cases, the presence of the second syntax element may be used to infer (e.g., used to implicitly signal) that a given one of the second constraint or the third constraint should be applied, while the absence of the second syntax element may be used to infer (e.g., used to implicitly signal) that a remaining one of the second constraint or the third constraint should be applied instead.
[0190] In some examples, one or more syntax elements (e.g., the first syntax element and / or the second syntax element) may include mode information, and a selected constraint from the aforementioned MVD constraints may be determined based on a mode (e.g., a merge mode) indicated by the mode information. For example, the one or more syntax elements may include mode information indicating a different merge mode for the current bilateral matching block. In one illustrative example, a first constraint (e.g., a mirroring constraint, MVD0=-MVD1) may be applied to a coded block in a normal (e.g., standard or default, such as enhanced merge prediction in VVC merge mode) merge mode when a normal merge candidate satisfies the DMVR condition set forth above.
[0191] As previously described, the DMVR condition may indicate a CU coded with one or more of the following modes or features. In some examples, the DMVR condition may include, but is not limited to, a CU coded with one or more of the following modes and / or features: CU level merge mode with bi-predictive MV, one reference picture in the past (e.g., relative to the current picture) and another reference picture in the future (e.g., relative to the current picture), the distance (e.g., POC distance) from the two reference pictures to the current picture is the same, both reference pictures are short-term reference pictures, a CU that includes more than 64 luma samples, both the CU height and the CU width are equal to or greater than 8 luma samples, the BCW weight index indicates equal weights, weighted prediction (WP) is not enabled for the current block, a combined inter- and intra-prediction (CIIP) mode is not utilized for the current block, etc.
[0192] In another illustrative example, when a coded block uses a specified new merge mode (e.g., adaptive bilateral matching mode as described herein), one of the second constraint (e.g., MVD0=0) or the third constraint (e.g., MVD1=0) may be applied, where all merge candidates satisfy the DMVR condition. In some cases, the second constraint and / or the third constraint may additionally or alternatively be indicated by a mode flag or a merge index. For example, an indication of a selection between the second constraint and the third constraint may be determined based on the mode flag or the merge index.
[0193] In some examples, one or more syntax elements described herein may include an index of a merge candidate list. The selected constraint may then be determined by the index indicating a selected merge candidate from the merge candidate list (e.g., the constraint depends on the selected merge candidate). In yet another example, the syntax element may include a combination of a mode flag and an index.
[0194] In some aspects, the systems and techniques described herein can signal (e.g., explicitly and / or implicitly) a selected search strategy that can be utilized to perform the multi-level (e.g., multi-path) DMVR described above. In some examples, the selected search strategy can be applied in one path of the multi-path DMVR. In another example, the selected search strategy can be applied in multiple levels or paths of the multi-path DMVR (e.g., but possibly not to all levels or paths of the process).
[0195] In the three-pass DMVR example described above, in one illustrative example, the selected strategy may be applied only in the first pass (e.g., PU-level bilateral matching). The second and third passes (e.g., performing bilateral matching for a first subblock size and performing BDOF for a second subblock size smaller than the first subblock size, respectively) may be performed utilizing a default strategy, e.g., the one described above with respect to FIG. 10 and a standardized three-pass structure. In one illustrative example, a multi-pass DMVR (e.g., the three-pass DMVR described above) may utilize a second search strategy (e.g., second constraint MVD0=0) and / or a third search strategy (e.g., third constraint MVD1=0) in the first pass to perform PU-level bilateral matching. Subsequent passes (e.g., the second and third passes) may use the default search strategy including the first constraint (e.g., mirroring constraint MVD0=-MVD1).
[0196] In some aspects, the search strategies may be grouped into a plurality of subsets. In some examples, one or more syntax elements may be used to determine the selected subset. In some cases, the selected strategy in a given subset may be implicitly determined. In one illustrative example, a first constraint (e.g., a mirroring constraint MVD0=-MVD1) may be included in the first subset, and both a second constraint (e.g., MVD0=0) and a third constraint (e.g., MVD1=0) may be included in the second subset. A syntax element may be used to indicate whether the second subset is used. If the second subset is used (e.g., based on the corresponding syntax element), the selection or decision between applying the second constraint and applying the third constraint may be implicitly determined. For example, an implicit decision between using the second constraint for bilateral matching and using the third constraint for bilateral matching may be made based on a minimum matching cost. If it is determined that bilateral matching using the second constraint results in a smaller matching cost than bilateral matching using the third constraint, the second constraint may be selected; otherwise, the third constraint is selected.
[0197] In some examples, one or more aspects of the systems and techniques described herein may be utilized with or applied based on an enhanced compression model (ECM). For example, in an ECM, multi-pass DMVR may be applied to a regular (e.g., standard or default, such as enhanced merge prediction in a VVC merge mode) merge mode candidate with one or more features the same or similar to those described above. For example, the one or more features may be the same or similar to some (or all) of the DMVR conditions described above. As mentioned above, a merge candidate refers to a candidate block from which information (e.g., one or more motion vectors, a prediction mode, etc.) is inherited for use in coding (e.g., encoding and / or decoding) a current block, and the candidate block may be a block adjacent to the current block. For example, a merge candidate may be an inter-predicted PU that includes a motion data location selected from one of a group of spatially adjacent motion data locations and two temporally co-located motion data locations.
[0198] Multi-path DMVR in ECM may be performed based on applying the first constraint (e.g., mirroring constraint MVD0=-MVD1) by default. In one illustrative example, the systems and techniques described herein may utilize an adaptive bilateral matching mode as a new mode for multi-path DMVR. In some examples, the adaptive bilateral matching mode may also be referred to as "adaptive_bm_mode". In adaptive_bm_mode, a merge index may be signaled to indicate the selected motion information candidate. However, all candidates in the candidate list may satisfy the DMVR condition. In some cases, a flag (e.g., bm_merge_flag) may be used to signal or indicate the use of the adaptive bilateral matching mode. For example, if the flag is true (e.g., bm_merge_flag is equal to 1), adaptive_bm_mode may be used or applied (e.g., as described below). In some aspects, when a flag is true (eg, bm_merge_flag is equal to 1), an additional flag (eg, bm_dir_flag) may be used to signal or indicate the bm_dir value to be used in adaptive_bm_mode.
[0199] In some aspects, when adaptive_bm_mode=1, adaptive bilateral matching may be performed, and when adaptive_bm_mode=0, adaptive bilateral matching is not performed. In one illustrative example, the adaptive bilateral matching process (e.g., associated with adaptive_bm_mode=1) may be performed based at least in part on applying either the second constraint or the third constraint to the selected candidate (e.g., fixing either MVD0 or MVD1 as equal to 0, respectively). In some cases, the variable bm_dir may be used to indicate which constraint is applied and / or to indicate the selected constraint. For example, when adaptive bilateral matching is performed (e.g., signaled or determined based on adaptive_bm_mode), bm_dir=1 may be used to indicate or signal that adaptive bilateral matching should be performed by fixing MVD1 to 0 (e.g., the third constraint should be applied). In some examples, bm_dir=2 may be used to indicate or signal that adaptive bilateral matching should be performed by fixing MVD0 to 0 (e.g., the second constraint should be applied).
[0200] In some cases, when adaptive bilateral matching is not performed, normal merge mode may be utilized. As previously mentioned, normal merge mode may be determined or signaled based on adaptive_bm_mode=0, which indicates that adaptive bilateral matching is not performed. In some examples, when normal merge mode is utilized (e.g., because adaptive bilateral matching is not performed), systems and techniques may utilize an inferred value of bm_dir=3, which indicates or signals that MVD0 and MVD1 are not fixed and MVD0=-MVD1 (e.g., a first mirroring constraint should be applied). In some examples, bm_dir=3 may be explicitly signaled or used to indicate normal merge mode.
[0201] In some examples, the systems and techniques may perform bilateral matching with one or more modified bilateral matching operations. In some aspects, when bm_dir=3, the bilateral matching process may be the same as the three-pass bilateral matching process previously described above. For example, given an initial pair of MVs, the first predictor is generated by the first MV referencing the first reference picture, and the second predictor is generated by the second MV referencing the second reference picture. A refined pair of MVs is then derived by minimizing the BM cost between the refined first predictor (e.g., generated using the first refined MV) and the refined second predictor (e.g., generated using the second refined MV), and the motion vector difference between the refined MV and the initial MV is MVD0 and MVD1, and MVD0=-MVD1.
[0202] In some aspects, when bm_dir=1, only the first MV is refined and the second MV remains fixed. For example, the refined first MV may be derived by minimizing the BM cost between the second predictor generated by the second MV and the refined first predictor generated by the refined first MV. For example, when bm_dir=1, MV0 may be refined but MV1 remains fixed. The refined motion vector MV0' may be derived by minimizing the BM cost between the predictor generated based on MV1 and the refined predictor generated based on MV0'.
[0203] In some examples, when bm_dir=2, only the second MV is refined and the first MV remains fixed. For example, the refined second MV may be derived by minimizing the BM cost between the first predictor generated by the first MV and the refined second predictor generated by the refined second MV. For example, when bm_dir=2, MV1 may be refined but MV0 remains fixed. The refined motion vector MV1' may be derived by minimizing the BM cost between the predictor generated based on MV1 and the refined predictor generated based on MV1'.
[0204] In some aspects, the BM cost (e.g., as described above) may additionally or alternatively include a regularization term based on or derived from a motion vector differential (MVD). In one illustrative example, a multi-pass DMVR search process may include a regularization term determined based on a MV cost that depends on a refined MV position.
[0205] In some aspects, one or more of the bilateral matching modifications described above may be applied only in the first pass of a multi-pass bilateral matching process (e.g., PU-level DMVR). In other aspects, one or more of the bilateral matching modifications described above may be applied in both the first and second passes (e.g., in the PU-level DMVR pass and the sub-PU-level DMVR pass).
[0206] In some examples, the systems and techniques described herein can perform multi-path bilateral matching DMVR using more or fewer paths than those described in the above examples. For example, fewer than three paths may be utilized and / or more than three paths may be utilized. In some examples, any number of paths may be used, and the paths may be constructed in any manner. In some aspects, in adaptive_bm_mode, a multi-path design may be applied as well. In some aspects, the second path may be skipped in adaptive_bm_mode. In some aspects, both the second and third paths may be skipped in adaptive_bm_mode. In other aspects, any combination of paths (e.g., path operations from the three-path system described above) may be combined with repeated paths or other path types based on certain adaptive bilateral matching criteria.
[0207] In some aspects of multi-pass DMVR, different search patterns may be used for different search levels and / or search precision. For example, a rectangular search may be used for both integer and half-pel searches that may be applied to perform a PU level DMVR (e.g., a first pass). In some examples, a full search may be used for integer searches and a rectangular search may be used for half-pel searches in a sub-PU level DMVR (e.g., a second pass and / or a third pass). In some examples, one or more (or all) of the search patterns described above may be utilized based on a determination that bm_dir is equal to a first value (a value of 1), a second value (a value of 2), or a third value (a value of 3). In one example, one or more (or all) of the search patterns described above may be utilized based on a determination that bm_dir=3.
[0208] The following aspects describe example search patterns and / or search processes when bm_dir is equal to 1 or 2. In one aspect, the same search pattern when bm_dir=3 may be used when bm_dir=1 and / or when bm_dir=2. In another aspect, bm_dir=1 and bm_dir=2 may use a different search pattern than when bm_dir=3. For example, in one example, a full search may be used for integer searches in the PU level DMVR and a rectangular search may be used for half-pel searches in the PU level DMVR.
[0209] In some aspects, the same search range and / or maximum search number may be used for different bm_dir values. In other aspects, different search ranges and / or different maximum search numbers may be used when bm_dir=1 or 2. For example, in the case of a full search where different cost factors are assigned to different MVDs, one or more MVD regions may be skipped. For example, as shown in FIG. 10, the search area for CU 1000 is divided into multiple search areas (e.g., first search area 1020, second search area 1030, third search area 1040, fourth search area 1050). In some cases, regions far from the center of the search area (e.g., far from first search area 1020) may be skipped.
[0210] In some examples, in the normal merge mode, all search areas may be searched. For example, with respect to FIG. 10, in the normal merge mode, the four search areas 1020, 1030, 1040, and 1050 may be searched, along with a fifth search area that includes the remaining blocks of the CU 1000 that are not yet included in one of the four search areas 1020-1050. As mentioned above, in some cases, search areas far from the center of the search area may be skipped. For example, in the adaptive bilateral matching mode (e.g., when bm_dir=1 or 2), only the first three search areas associated with the CU 1000 in FIG. 10 may be searched (e.g., in the adaptive bilateral matching mode, the first search area 1020, the second search area 1030, and the third search area 1040 may be searched).
[0211] In some aspects of multipath DMVR, SAD or average removed SAD (e.g., depending on PU size) may be used for integer search and half-pel search associated with PU level DMVR path. In some cases, SATD may be used for sub-PU level DMVR. In some aspects, the same cost criterion may be used for different values of bm_dir. For example, the current cost criterion selection in ECM may be used for all values of bm_dir. In other aspects, the cost criterion selection may be different for different values of bm_dir. For example, when bm_dir=3, the cost criterion selection in the current ECM may be applied. SAD or average removed SAD may be used for integer search or half-pel search in PU level DMVR depending on PU size. When bm_dir=1 or 2, the PU level DMVR process may use SATD for integer search and SAD for half-pel search.
[0212] As previously mentioned, the candidates in the candidate list for some aspects satisfy the DMVR condition. In one additional aspect, bm_dir may be set equal to 1 or 2 as indicated by an additional bm_dir_flag included in the adaptive_bm_mode mode. In some cases, the candidate list for adaptive_bm_mode is generated in addition to the regular merge candidate list. For example, one or more candidates in the regular merge candidate list that are determined to satisfy the DMVR condition may be inserted into the candidate list for adaptive_bm_mode.
[0213] In another additional aspect, whether to set bm_dir equal to 1 or 2 may be indicated by a merge index in adaptive_bm_mode mode. A candidate list for adaptive_bm_mode may be generated in addition to a normal merge candidate list. For each candidate in the normal merge candidate list that satisfies the DMVR condition, a pair of two candidates may be inserted into the adaptive_bm_mode candidate list, where one candidate has bm_dir=1 and the other candidate has bm_dir=2, and the two candidates in the pair have identical motion information. In some examples, bm_dir may be determined by determining whether the merge index is even or odd.
[0214] In one illustrative example, the candidate list for adaptive_bm_mode may be generated independently of the regular merge candidate list. In some cases, the generation of the candidate list for adaptive_bm_mode may follow the same or similar process as generating the candidate list for the regular merge mode (e.g., checking the same spatial, temporal neighboring locations, history-based candidates, pairwise candidates, etc.). In some cases, pruning may be applied during the list building process.
[0215] In yet another aspect, one or more candidates associated with bi-predictive (BCW) weight indices with CU-level weights indicating unequal weights may also be added to the candidate list (e.g., compared to some systems with DMVR conditions, where BCW weight indices may indicate equal weights).
[0216] In some examples, if the number of candidates in the candidate list for adaptive_bm_mode and / or in the candidate list for normal merge mode is less than a predetermined maximum, a padding process may be applied. For example, when a candidate list for adaptive_bm_mode is generated, there may be fewer candidates in the list than a predetermined maximum number of candidates. In such examples, a padding process may be applied to generate a number of padded candidates for the candidate list such that the padded candidate list includes a predetermined number of candidates. In one illustrative example, when adaptive bilateral matching is enabled (e.g., adaptive_bm_mode=1), one or more default candidates for padding in merge list construction may be used. The default candidates may be derived to satisfy the DMVR condition.
[0217] In some examples, the MV may be set to 0 for the default candidate. For example, a zero MV candidate may be added during the padding process. In some examples, a reference picture may be selected according to a DMVR condition. In some cases, the reference index may be iterated through all possible values until the number of candidates in the candidate list reaches a maximum number of candidates (e.g., a predefined maximum). In another aspect, a BCW weight index may be used to indicate equal weights for the normal candidates, and one or more unequal weight BCW candidates may be added thereafter and before adding the zero candidate.
[0218] In some aspects, a reference picture assigned to a default candidate (e.g., a reference picture assigned to a padded zero MV candidate described above) may be selected to satisfy one or more conditions associated with adaptive_bm_mode. Illustrative examples of such conditions may include one or more of the following conditions: at least one pair of reference pictures is selected that includes one reference picture in the past and one reference picture in the future relative to the current picture, the respective distances from both reference pictures to the current picture are equal, both of the reference pictures are not long-term reference pictures, both of the reference pictures have the same resolution as the current picture, weighted prediction (WP) is not applied to any of the reference pictures, any combination of these, and / or other conditions not enumerated herein.
[0219] In some aspects, one or more reference pictures may be assigned to a default candidate (e.g., to the padded zero MV candidate described above) based on selecting a reference picture that satisfies one or more conditions associated with or based on a specified or selected constraint (e.g., the second constraint MVD0=0, or the third constraint MVD1=0). In some examples, a reference picture may be selected to satisfy one or more conditions associated with a given constraint, the given constraint being determined based on bm_dir (e.g., as previously described). For example, one or more conditions based on bm_dir may apply only to the MV for which refinement is performed (e.g., when bm_dir=1). Illustrative examples of such conditions may include one or more of the following conditions: a reference picture in a given list X is not a long-term reference picture; a reference picture in list X has the same resolution as the current picture; weighted prediction (WP) is not applied to the reference pictures in list X; a respective distance from a first reference picture in list X to the current picture is not smaller than a respective distance from another reference picture (e.g., a second reference picture in list X) to the current picture; any combination of these; and / or other conditions; and / or other conditions not enumerated herein.
[0220] In some cases, if bm_dir indicates that the MVs in list 0 are refined, then list X may be the same as list 0 (e.g., list L0). In some cases, if bm_dir indicates that the MVs in list 1 are refined, then list X may be the same as list 1 (e.g., list L1). In some aspects, all potential zero MV candidates may be discovered by iterating over all possible combinations of reference pictures and identifying reference pictures that satisfy a predetermined condition in an order. In one illustrative example, a first loop may be performed for list 0 and a second loop may be performed for list 1. In another example, a first loop may be performed for list 1 and a second loop may be performed for list 0. Other orders are possible and should be considered within the scope of this disclosure. The process of determining possible zero MV candidates (e.g., by iterating over combinations of reference pictures and identifying reference pictures that satisfy a predetermined condition in an order) may be performed at the slice level, picture level, or other levels. The list of identified default MV candidates (e.g., zero MV candidates) may be stored as default candidates. In some cases, when determining potential zero MV candidates at the block level, when the number of candidates is less than a predetermined maximum number of candidates, the systems and techniques described herein may iterate through the default candidates and add one or more of the default candidates to the candidate list until the number of candidates reaches the predetermined maximum.
[0221] In some aspects, one or more size constraints may be included in and / or utilized by adaptive_bm_mode described herein. In one aspect, the same size constraints as in normal DMVR may be applied in adaptive_bm_mode. In another aspect, if neither the width nor the height of the current block is larger than the minimum block size for DMVR, adaptive_bm_mode is not applied.
[0222] In some aspects, adaptive_bm_mode may be signaled as an additional merge mode to the normal merge mode. In some examples, various signaling methods may be applied or utilized to signal adaptive_bm_mode as an additional merge mode. For example, adaptive_bm_mode may be considered or signaled as a variant of the normal merge mode. In one illustrative example, one or more syntax elements may be signaled initially to indicate the normal merge mode, and one or more additional flags and / or syntax elements may be signaled to indicate adaptive_bm_mode and / or to indicate a particular one of the constraints that may be applied in conjunction with the use of adaptive_bm_mode (e.g., the second constraint MDV0=0 or the third constraint MDV1=0).
[0223] In another aspect, adaptive_bm_mode may be indicated by one or more flags before the indication of normal merge mode. For example, if the syntax (e.g., one or more syntax elements) indicates that the current block does not use adaptive_bm_mode, one or more additional syntax elements may be signaled to indicate whether the current block uses normal merge mode or other merge modes. For example, if the current block does not use adaptive_bm_mode or normal merge mode, one or more additional syntax elements may be signaled to indicate that the current block uses other merge modes such as synthetic inter- and intra-prediction (CIIP), geometric partition mode (GPM), etc.
[0224] In yet another aspect, adaptive_bm_mode may be signaled in other merge mode branches. For example, adaptive_bm_mode may be signaled in a template matching merge mode branch. In some cases, one or more syntax elements may be signaled first to indicate whether one of adaptive_bm_mode and template matching merge mode is used. If one or more syntax elements indicate that adaptive_bm_mode or template matching merge mode is used, one or more additional flags or syntax elements may be signaled to indicate whether template matching merge mode or adaptive_bm_mode is used.
[0225] In some aspects, the merge index in adaptive_bm_mode can use the same signaling method as the normal merge mode. In one aspect, adaptive_bm_mode can use the same (or similar) context model as the normal merge mode. In another aspect, a separate context model may be used for adaptive_bm_mode. In some examples, for adaptive_bm_mode, the maximum number of merge candidates may be different from the maximum number of merge candidates for the normal merge mode.
[0226] In one illustrative example, one or more high-level syntax elements may be used to indicate whether adaptive_bm_mode may or will be applied. In one aspect, the same high-level syntax used to indicate whether normal DMVR is to be applied may also be used to indicate whether adaptive_bm_mode is to be applied. In another aspect, one or more separate (e.g., additional) high-level syntax elements may be used to indicate whether adaptive_bm_mode should be applied. In yet another aspect, a separate high-level syntax element may be used to indicate whether adaptive_bm_mode is utilized, with a separate high-level syntax element for adaptive_bm_mode only being present if normal DMVR is enabled. For example, if a separate high-level syntax element related to normal DMVR is determined to indicate that normal DMVR is not enabled or utilized, then a separate or additional high-level syntax element related to adaptive_bm_mode is not signaled and adaptive_bm_mode is inferred to be off (e.g., not enabled or utilized).
[0227] In some aspects, in addition to one or more high-level syntax elements described above, adaptive_bm_mode may be disabled for a coded picture or slice according to available reference pictures. In some cases, if it is determined that no combination of reference pictures satisfies or cannot satisfy the reference picture condition, adaptive_bm_mode may be disabled and the corresponding syntax element (e.g., at the block level) is not signaled. In some cases, there must be at least one pair of reference pictures that satisfy the reference picture condition to utilize adaptive_bm_mode. Illustrative examples of such conditions may include one or more of the following conditions: one reference picture in the past and one reference picture in the future relative to the current picture, the respective distances from both reference pictures to the current picture are equal, both reference pictures are not long-term reference pictures, both reference pictures have the same resolution as the current picture, weighted prediction (WP) is not applied to either reference picture, any combination of these, and / or other conditions, and / or other conditions not enumerated herein.
[0228] The conditions listed above may be used separately or in combination.
[0229] In some aspects, only a subset of the adaptive bilateral matching modes described herein (e.g., a first adaptive bilateral matching mode associated with the second condition MDV0=0 and a second adaptive bilateral matching mode associated with the third condition MDV1=0) may be enabled depending on the reference picture. Illustrative examples of such conditions for allowing only bm_dir=1 (e.g., associated with the second condition MDV0=0) or bm_dir=2 (e.g., associated with the third condition MDV1=0) may include one or more of the following conditions: one of the reference pictures is a long-term reference picture and the other reference picture is not a long-term reference picture; one of the reference pictures has the same resolution as the current picture but the other reference picture has a different resolution than the current picture; weighted prediction (WP) is applied to one of the reference pictures; the respective distances from one of the reference pictures to the current picture are not shorter than the respective distances from the other reference pictures to the current picture; any combination of these; and / or other conditions; and / or may include other conditions not enumerated here.
[0230] In some cases, a syntax element (e.g., at the block level) specifying the bilateral matching mode may not be signaled and may be inferred accordingly. In some examples, the syntax element may be implicitly signaled (e.g., inferred) rather than explicitly signaled (e.g., in the bitstream as part of a particular syntax table). In some cases, the syntax element may not be explicitly signaled or implicitly signaled and may be inferred. For example, if a first value bm_dir is enabled (e.g., bm_dir=1 is enabled) but a second value bm_dir is disabled (e.g., bm_dir=2 is disabled), a syntax element is used to indicate that bm_dir at the block level may not be signaled. In the absence of this syntax element, bm_dir may be inferred to be the first value (e.g., bm_dir is inferred to be 1). In another example, a syntax element is used to indicate that if a second value of bm_dir is enabled (e.g., bm_dir=2) but a first value of bm_dir is disabled (e.g., bm_dir=1 is disabled), then bm_dir at the block level may not be signaled. In the absence of this syntax element, bm_dir may be inferred to be the second value (e.g., bm_dir is inferred to be 2).
[0231] In another example, a slice-level flag and / or a picture-level flag may be utilized for adaptive_bm_mode, e.g., as a bitstream conformance constraint where if at least one of the above conditions is not met, the flag is set to 0 (e.g., DMVR mode is disabled).
[0232] In yet another example, a bitstream adaptation constraint may be introduced into the existing signaling, where if at least one of the above conditions is not met, the bitstream adaptation constraint indicates that adaptive_bm_mode should not be applied and the corresponding overhead is set to 0 (e.g., indicating that adaptive_bm_mode is not used).
[0233] In some aspects of ECM, multiple hypothesis prediction (MHP) may be utilized. In MHP, an inter-prediction technique may be used to obtain or determine a weighted overlap of more than two motion-compensated prediction signals. Based on performing a weighted overlap on a sample-by-sample basis, a resulting overall prediction signal may be obtained. For example, a uni / bi-prediction signal p uni / bi and the first additional inter-prediction signal / hypothesis h 3 Based on this, the resulting predicted signal p 3 can be obtained as follows: p 3 =(1-α)p uni / bi +αh 3 Formula (27)
[0234] Here, the weighting factor α may be specified by the syntax element add_hyp_weight_idx according to the following mapping:
[0235] [Table 2]
[0236] In some examples, more than one additional prediction signal may be used. In some cases, more than one additional prediction signal may be utilized in the same or similar manner as above. For example, when utilizing multiple additional prediction signals, the resulting overall prediction signal may be accumulated iteratively for each additional prediction signal as follows: p n+1 =(1-α n+1 ) pn +α n+1 h n+1 Formula (28)
[0237] Now the resulting whole prediction signal is the final p n (For example, p with the largest index n n )
[0238] In some aspects, MHP may not be applied (e.g., disabled) for any adaptive_bm_mode. In some aspects, MHP may be applied in addition to adaptive_bm_mode in the same or similar manner as MHP is applied for standardized (e.g., normal) merge mode.
[0239] FIG. 11 is a flowchart illustrating an example of a process 1100 for processing video data. In some examples, the process 100 may be used to perform decoder-side motion vector refinement (DMVR) using adaptive bilateral matching according to some examples of the present disclosure. In some aspects, the process 1100 may be implemented in an apparatus for processing video data comprising a memory and one or more processors coupled to the memory configured to perform operations of the process 1100. In other aspects, the process 1100 may be implemented in a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of the device, cause the device to perform operations of the process 1100.
[0240] At block 1102, process 1100 includes obtaining one or more reference pictures for a current picture (e.g., a current block of the current picture). For example, the one or more reference pictures may be obtained based on one or more of the inputs 114 to the decoding device 112 shown in FIG. 1. In some examples, the one or more reference pictures and the current picture may be obtained from video data obtained by or provided to the decoding device 112 shown in FIG. 1.
[0241] At block 1104, the process 1100 includes identifying a first motion vector and a second motion vector for the merge mode candidate. For example, the first motion vector and / or the second motion vector may be identified by the decoding device 112 shown in FIG. 1. In some examples, the first motion vector and / or the second motion vector may be identified using the decoder engine 116 of the decoding device 112 shown in FIG. 1. In some cases, one or more (or both) of the first motion vector and the second motion vector may be identified using signaled information. For example, the encoding device 104 shown in FIG. 1 may include signaling information that may be used by the decoding device 112 and / or the decoding engine 116 to identify one or more (or both) of the first motion vector and the second motion vector. In some cases, the process 1100 may include determining a merge mode candidate for the current picture. As described herein, the merge mode candidate may include a neighboring block of a block from which predictive data may be inherited for a block of the current picture. For example, the merge mode candidate may be determined by the decoding device 112 shown in FIG. 1. In some examples, the merge mode candidate may be determined using the decoder engine 116 of the decoding device 112 shown in FIG. 1. In some examples, signaling information may be used by the decoding device 112 and / or the decoding engine 116 to determine the merge mode candidate for the current picture. In some cases, the merge mode candidate may be determined using the same signaling information used to identify one or more (or both) of the first and second motion vectors associated with the merge mode candidate. In some cases, the merge mode candidate and the first and second motion vectors may be determined using separate signaling information.
[0242] At block 1106, the process 1100 includes determining a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies. In some aspects, the selected motion vector search strategy is associated with one or more constraints based on or corresponding to the first motion vector and / or the second motion vector. In one illustrative example, the selected motion vector search strategy may be a bilateral matching (BM) motion vector search strategy. In some cases, the motion vector search strategy for the merge mode candidate may be selected from a plurality of motion vector search strategies including at least two of a multi-pass decoder-side motion vector refinement strategy, a fractional sample refinement strategy, a bidirectional optical flow strategy, or a sub-block-based bilateral matching motion vector refinement strategy. In some examples, the selected motion vector search strategy may be determined before the merge mode candidate is determined. For example, the selected motion vector search strategy may be determined (e.g., as described above) and used to generate a merge candidate list. The merge mode candidate may be determined based on a selection from the generated merge candidate list. In some examples, the selected merge mode candidate may be determined before the selected motion vector search strategy is determined. For example, in some cases, the merge candidate list may be generated without using the selected search strategy (e.g., the generated merge candidate list may be the same for each respective search strategy of the multiple search strategies), and the selected merge candidate may be determined before the selected search strategy.
[0243] In some examples, the selected motion vector search strategy may be a multi-pass decoder-side motion vector refinement (DMVR) search strategy. For example, the multi-pass DMVR search strategy may include one or more block-based bilateral matching motion vector refinement passes and may also include one or more sub-block-based motion vector refinement passes. In some examples, the one or more block-based bilateral matching motion vector refinement passes may be performed using a first constraint associated with the first motion vector differential and / or the second motion vector differential. The first motion vector differential may be a difference determined between the first motion vector and the refined first motion vector. The second motion vector differential may be a difference determined between the second motion vector and the refined second motion vector. In some examples, the one or more sub-block-based motion vector refinement passes may be performed using a second constraint different from the first constraint. As described above, the second constraint may be associated with at least one of the first motion vector differential and / or the second motion vector differential. In some cases, the one or more sub-block-based motion vector refinement passes may include at least one of a sub-block-based bilateral matching (BM) motion vector refinement pass and / or a sub-block-based bidirectional optical flow (BDOF) motion vector refinement pass.
[0244] In some examples, the selected motion vector search strategy is associated with one or more constraints corresponding to at least one of the first motion vector or the second motion vector (e.g., as described above). The one or more constraints may be determined based on one or more signaled syntax elements. For example, the one or more constraints may be determined for a block of the current picture based on syntax elements signaled for the block. In some aspects, the one or more constraints are associated with at least one of a first motion vector differential associated with the first motion vector (e.g., the difference between the first motion vector and the refined first motion vector) and a second motion vector differential associated with the second motion vector (e.g., the difference between the second motion vector and the refined second motion vector). In some examples, the one or more constraints may include a mirroring constraint for the first motion vector differential and the second motion vector differential. The mirroring constraint may set the first motion vector differential and the second motion vector differential to be equal in magnitude (e.g., absolute value) but opposite in sign. In some cases, the one or more constraints may include a zero value constraint for the first motion vector differential (e.g., setting the first motion vector differential equal to 0). In some examples, the one or more constraints may include a zero value constraint for the second motion vector differential (e.g., setting the second motion vector differential equal to 0). In some aspects, the zero value constraint may be indicative of keeping the motion vector differential constant. For example, based on the zero value constraint, the process 1100 may include determining one or more refined motion vectors using a selected motion vector search strategy by keeping a first one of the first motion vector differential or the second motion vector differential at a constant value and searching against a second one of the first motion vector differential or the second motion vector differential. For example, the first motion vector differential may be fixed and a search may be performed around the second motion vector differential to derive the refined motion vector.
[0245] At block 1108, process 1100 includes determining one or more refined motion vectors based on the first motion vector, the second motion vector, and / or one or more reference pictures (e.g., based on the first motion vector and one or more reference pictures, based on the second motion vector and one or more reference pictures, or based on the first motion vector, the second motion vector, and one or more reference pictures) using the selected motion vector search strategy. In some cases, determining the one or more refined motion vectors may include determining one or more refined motion vectors for the block of video data. In some examples, the one or more refined motion vectors may include a first refined motion vector and a second refined motion vector determined for the first motion vector and the second motion vector, respectively. In some examples, a first motion vector differential is determined as a difference between the first refined motion vector and the first motion vector, and a second motion vector differential is determined as a difference between the second refined motion vector and the second motion vector.
[0246] In some examples, the selected motion vector search strategy is a bilateral matching (BM) motion vector search strategy, as previously mentioned. When the selected motion vector search strategy is a BM motion vector search strategy, determining one or more refined motion vectors may include determining a first refined motion vector by searching a first reference picture around the first motion vector. The first reference picture may be searched around the first motion vector based on the selected motion vector search strategy. The second refined motion vector may be determined by searching a second reference picture around the second motion vector based on the selected motion vector search strategy. The selected motion vector search strategy may include a motion vector difference constraint (e.g., a mirroring constraint that the first motion vector difference and the second motion vector difference are equal in magnitude but opposite in sign, a constraint setting the first motion vector difference equal to 0, a constraint setting the second motion vector difference equal to 0, etc.). In some examples, the first refined motion vector and the second refined motion vector may be determined by minimizing a difference between a first reference block associated with the first refined motion vector and a second reference block associated with the second refined motion vector.
[0247] At block 1110, the process 1100 includes processing the merge mode candidate using the one or more refined motion vectors. For example, the decoding device 112 shown in FIG. 1 may process the merge mode candidate using the one or more refined motion vectors. In some examples, the decoder engine 116 of the decoding device 112 shown in FIG. 1 may process the merge mode candidate using the one or more refined motion vectors.
[0248] In some implementations, the processes (or methods) described herein may be performed by a computing device or apparatus, such as the system 100 shown in Figure 1. For example, the processes may be performed by the encoding device 104 shown in Figures 1 and 12, by another video source side device or video transmission device, by the decoding device 112 shown in Figures 1 and 13, and / or by another client side device, such as a player device, a display, or any other client side device. In some cases, the computing device or apparatus may include one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components configured to perform the steps of the process 1100.
[0249] In some examples, the computing device may include a wireless communication device, such as a mobile device, a tablet computer, an extended reality (XR) device (e.g., a virtual reality (VR) device such as a head-mounted display (HMD), an AR device such as an HMD or augmented reality (AR) glasses, an MR device such as an HMD or mixed reality (MR) glasses, etc.), a desktop computer, a server computer and / or server system, a computing system of a vehicle or a component of a vehicle, or other type of computing device. Components of a computing device (e.g., one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) may be implemented with circuits. For example, a component may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein. In some examples, a computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) including video frames. In some examples, the camera or other capture device that captures the video data is separate from the computing device, in which case the computing device receives or obtains the captured video data.The computing device may include a network interface configured to communicate video data. The network interface may be configured to communicate Internet Protocol (IP) based data or other types of data. In some examples, the computing device or apparatus may include a display for displaying output video content, such as samples of pictures of a video bitstream.
[0250] The processes are described with respect to logic flow diagrams, whose operations represent sequences of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement a process.
[0251] Additionally, the process may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors, by hardware, or a combination thereof. As mentioned above, the code may be stored in a computer-readable or machine-readable storage medium, for example in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0252] The coding techniques discussed herein may be implemented in an exemplary video encoding and decoding system (e.g., system 100). In some examples, the system includes a source device that provides encoded video data to be subsequently decoded by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source device and destination device may comprise any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, and the like. In some cases, the source device and destination device may be capable of wireless communication.
[0253] The destination device may receive the encoded video data to be decoded via a computer-readable medium. The computer-readable medium may comprise any type of medium or device capable of moving the encoded video data from the source device to the destination device. In one example, the computer-readable medium may comprise a communication medium for enabling the source device to transmit the encoded video data directly to the destination device in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device. The communication medium may comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from the source device to the destination device.
[0254] In some examples, the encoded data may be output from the output interface to a storage device. Similarly, the encoded data may be accessed from the storage device by the input interface. The storage device may include any of a variety of distributed or locally accessed data storage media, such as hard drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In further examples, the storage device may correspond to a file server or another intermediate storage device that may store the encoded video generated by the source device. The destination device may access the stored video data from the storage device via streaming or download. The file server may be any type of server capable of storing the encoded video data and transmitting the encoded video data to the destination device. Exemplary file servers include a web server (e.g., for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. The destination device may access the encoded video data through any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.
[0255] The techniques of this disclosure are not necessarily limited to wireless applications or settings. The techniques may be applied to video coding to support any of a variety of multimedia applications, such as over-the-air television broadcast, cable television transmission, satellite television transmission, Internet streaming video transmission such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system may be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0256] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device may include an input interface, a video decoder, and a display device. The video encoder of the source device may be configured to apply the techniques disclosed herein. In other examples, the source device and the destination device may include other components or configurations. For example, the source device may receive video data from an external video source, such as an external camera. Similarly, the destination device may interface with an external display device rather than including an integrated display device.
[0257] The above exemplary system is only an example. The techniques for processing video data in parallel may be performed by any digital video encoding and / or decoding device. In general, the techniques of this disclosure are performed by a video encoding device, but the techniques may also be performed by a video encoder / decoder, usually referred to as a "codec." Moreover, the techniques of this disclosure may also be performed by a video processor. The source device and the destination device are only examples of coding devices, such that the source device generates coded video data for transmission to the destination device. In some examples, the source device and the destination device may operate substantially symmetrically, such that each of the devices includes video encoding and decoding components. Thus, the exemplary system may support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0258] The video source may include a video capture device such as a video camera, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source may generate computer graphics-based data as the source video, or a combination of live, archived, and computer-generated video. In some cases, when the video source is a video camera, the source device and destination device may form a so-called camera phone or video phone. However, as noted above, the techniques described in this disclosure may be applicable to video coding in general and may be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by a video encoder. The encoded video information may then be output onto a computer-readable medium by an output interface.
[0259] As stated, computer-readable media may include transitory media, such as wireless broadcast or wired network transmission, or storage media (i.e., non-transitory storage media), such as hard disks, flash drives, compact discs, digital video discs, Blu-ray discs, or other computer-readable media. In some examples, a network server (not shown) may receive the encoded video data from a source device, e.g., via a network transmission, and provide the encoded video data to a destination device. Similarly, a computing device at a media production facility, such as a disc stamping facility, may receive the encoded video data from a source device and produce discs including the encoded video data. Thus, computer-readable media may be understood to include one or more computer-readable media of various forms in various examples.
[0260] The input interface of the destination device receives information from the computer-readable medium. The information of the computer-readable medium may include syntax information defined by a video encoder, including syntax elements that describe characteristics and / or processing of blocks and other coded units, e.g., groups of pictures (GOPs), which are also used by a video decoder. The display device displays the decoded video data to a user and may comprise any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device. Various aspects of the present application have been described.
[0261] Specific details of the encoding device 104 and the decoding device 112 are illustrated in Figures 12 and 13, respectively. Figure 12 is a block diagram illustrating an example encoding device 104 that may implement one or more of the techniques described in this disclosure. The encoding device 104 may generate, for example, a syntax structure described herein (e.g., a syntax structure for a VPS, SPS, PPS, or other syntax elements). The encoding device 104 may perform intra-prediction and inter-prediction coding of video blocks within a video slice. As previously described, intra-coding relies at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter-coding relies at least in part on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames of a video sequence. An intra-mode (I-mode) may refer to any of several spatial-based compression modes. An inter-mode, such as unidirectional prediction (P-mode) or bidirectional prediction (B-mode), may refer to any of several temporal-based compression modes.
[0262] Encoding device 104 includes partition unit 35, prediction processing unit 41, filter unit 63, picture memory 64, adder 50, transform processing unit 52, quantization unit 54, and entropy coding unit 56. Prediction processing unit 41 includes motion estimation unit 42, motion compensation unit 44, and intra-prediction processing unit 46. For video block reconstruction, encoding device 104 also includes inverse quantization unit 58, inverse transform processing unit 60, and adder 62. Filter unit 63 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although filter unit 63 is illustrated in FIG. 12 as being an in-loop filter, in other configurations filter unit 63 may be implemented as a post-loop filter. Post-processing device 57 may perform additional processing on the encoded video data generated by encoding device 104. Techniques of this disclosure may be implemented by encoding device 104 in some instances. However, in other cases, one or more of the techniques of this disclosure may be performed by post-processing device 57.
[0263] As shown in FIG. 12, encoding device 104 receives video data, and partitioning unit 35 partitions the data into video blocks. Partitioning may also include partitioning into slices, slice segments, tiles, or other larger units, as well as video block partitioning, for example, according to a quadtree structure of LCUs and CUs. Encoding device 104 generally refers to a component that encodes video blocks within a video slice to be encoded. A slice may be divided into multiple video blocks (and possibly into a set of video blocks called tiles). Prediction processing unit 41 may select one of multiple possible coding modes, such as one of multiple intra-predictive coding modes or one of multiple inter-predictive coding modes, for a current video block based on error results (e.g., coding rate, level of distortion, etc.). Prediction processing unit 41 may provide the resulting intra- or inter-coded block to summer 50 to generate residual block data and to summer 62 to reconstruct the encoded block for use as a reference picture.
[0264] Intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-predictive coding of the current video block relative to one or more neighboring blocks in the same frame or slice as the current block to be coded to provide spatial compression. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 perform inter-predictive coding of the current video block relative to one or more predictive blocks in one or more reference pictures to provide temporal compression.
[0265] Motion estimation unit 42 may be configured to determine an inter-prediction mode for a video slice according to a predetermined pattern for the video sequence. The predetermined pattern may designate a video slice in the sequence as a P slice, a B slice, or a GPB slice. Motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of a video block. A motion vector may indicate, for example, the displacement of a prediction unit (PU) of a video block in a current video frame or picture relative to a predictive block in a reference picture.
[0266] The predictive block is a block found to closely match the PU of the video block to be coded in terms of pixel differences, which may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference measures. In some examples, encoding device 104 may calculate values for sub-integer pixel locations of reference pictures stored in picture memory 64. For example, encoding device 104 may interpolate values for quarter-pixel locations, eighth-pixel locations, or other fractional pixel locations of the reference pictures. Thus, motion estimation unit 42 may perform motion searches for full-pixel locations and fractional pixel locations and output motion vectors with fractional pixel locations.
[0267] Motion estimation unit 42 calculates a motion vector for a PU of a video block in an inter-coded slice by comparing the position of the PU to the position of a predictive block of a reference picture. The reference picture may be selected from a first reference picture list (List 0) or a second reference picture list (List 1), each of which identifies one or more reference pictures stored in picture memory 64. Motion estimation unit 42 sends the calculated motion vector to entropy encoding unit 56 and motion compensation unit 44.
[0268] Motion compensation performed by motion compensation unit 44 may involve fetching or generating a predictive block based on a motion vector determined by motion estimation, possibly performing interpolation to sub-pixel precision. Upon receiving the motion vector of the PU of the current video block, motion compensation unit 44 may locate the predictive block to which the motion vector points in the reference picture list. Encoding device 104 forms a residual video block by subtracting pixel values of the predictive block from pixel values of the current video block being coded to form pixel difference values. The pixel difference values form residual data for the block and may include both luma and chroma difference components. Adder 50 represents one or more components that perform this subtraction operation. Motion compensation unit 44 may also generate syntax elements related to the video block and video slice for use by decoding device 112 in decoding the video block of the video slice.
[0269] Intra-prediction processing unit 46 may intra-predict the current block as an alternative to inter-prediction performed by motion estimation unit 42 and motion compensation unit 44, as described above. Specifically, intra-prediction processing unit 46 may determine an intra-prediction mode to use to encode the current block. In some examples, intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, e.g., during separate encoding passes, and intra-prediction processing unit 46 may select an appropriate intra-prediction mode to use from the tested modes. For example, intra-prediction processing unit 46 may calculate rate-distortion values using a rate-distortion analysis for the various tested intra-prediction modes and may select the intra-prediction mode with the best rate-distortion characteristics from among the tested modes. The rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original uncoded block that was coded to generate the encoded block, as well as the bitrate (i.e., number of bits) used to generate the encoded block. Intra-prediction processing unit 46 may calculate ratios from the distortions and rates for the various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.
[0270] In either case, after selecting an intra-prediction mode for a block, intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to entropy encoding unit 56. Entropy encoding unit 56 may encode the information indicating the selected intra-prediction mode. Encoding device 104 may include in the transmitted bitstream configuration data definitions of the encoding contexts for the various blocks, as well as an indication of the most probable intra-prediction mode, intra-prediction mode index table, and modified intra-prediction mode index table to use for each of the contexts. The bitstream configuration data may include multiple intra-prediction mode index tables and multiple modified intra-prediction mode index tables (also referred to as codeword mapping tables).
[0271] After prediction processing unit 41 generates a predictive block for a current video block via either inter-prediction or intra-prediction, encoding device 104 forms a residual video block by subtracting the predictive block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to transform processing unit 52. Transform processing unit 52 converts the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform. Transform processing unit 52 may convert the residual video data from the pixel domain to a transform domain, such as the frequency domain.
[0272] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan of a matrix including the quantized transform coefficients. Alternatively, entropy encoding unit 56 may perform the scan.
[0273] Following quantization, entropy coding unit 56 entropy codes the quantized transform coefficients. For example, entropy coding unit 56 may perform context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding, or another entropy coding technique. Following entropy coding by entropy coding unit 56, the coded bitstream may be transmitted to decoding device 112 or archived for later transmission or retrieval by decoding device 112. Entropy coding unit 56 may also entropy code motion vectors and other syntax elements of the current video slice being coded.
[0274] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual block in the pixel domain for later use as a reference block of a reference picture. Motion compensation unit 44 may calculate a reference block by adding the residual block to a predictive block of one of the reference pictures in the reference picture list. Motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for use in motion estimation. Adder 62 adds the reconstructed residual block to a motion compensated predictive block generated by motion compensation unit 44 to generate a reference block for storage in picture memory 64. The reference block may be used by motion estimation unit 42 and motion compensation unit 44 as a reference block for inter predicting blocks in subsequent video frames or pictures.
[0275] In this manner, encoding device 104 of Figure 12 represents an example of a video encoder configured to perform any of the techniques described herein, including the process described above with respect to Figure 11. In some cases, some of the techniques of this disclosure may also be performed by post-processing device 57.
[0276] 13 is a block diagram illustrating an example decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, a filter unit 91, and a picture memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-prediction processing unit 84. The decoding device 112 may, in some examples, perform a decoding path that is generally inverse to the encoding path described with respect to the encoding device 104 from FIG.
[0277] During the decoding process, the decode device 112 receives an encoded video bitstream representing video blocks and associated syntax elements of an encoded video slice transmitted by the encoding device 104. In some aspects, the decode device 112 may receive the encoded video bitstream from the encoding device 104. In some aspects, the decode device 112 may receive the encoded video bitstream from a network entity 79, such as a server, a media aware network element (MANE), a video editor / splicer, or other such device configured to implement one or more of the techniques described above. The network entity 79 may or may not include the encoding device 104. Some of the techniques described in this disclosure may be performed by the network entity 79 before the network entity 79 transmits the encoded video bitstream to the decode device 112. In some video decoding systems, the network entity 79 and the decode device 112 may be part of separate devices, while in other cases, the functions described with respect to the network entity 79 may be performed by the same device that comprises the decode device 112.
[0278] Entropy decoding unit 80 of decoding device 112 entropy decodes the bitstream to generate quantized coefficients, motion vectors, and other syntax elements. Entropy decoding unit 80 forwards the motion vectors and other syntax elements to prediction processing unit 81. Decoding device 112 may receive syntax elements at the video slice level and / or the video block level. Entropy decoding unit 80 may process and parse both fixed-length and variable-length syntax elements in one or more parameter sets, such as VPS, SPS, and PPS.
[0279] When a video slice is coded as an intra-coded (I) slice, intra-prediction processing unit 84 of prediction processing unit 81 may generate predictive data for video blocks of the current video slice based on the signaled intra-prediction mode and data from previously decoded blocks of the current frame or picture. When a video frame is coded as an inter-coded (i.e., B, P, or GPB) slice, motion compensation unit 82 of prediction processing unit 81 generates predictive blocks for video blocks of the current video slice based on the motion vectors and other syntax elements received from entropy decoding unit 80. The predictive blocks may be generated from one of the reference pictures in the reference picture list. Decoding device 112 may construct the reference frame lists, i.e., List 0 and List 1, using a default construction technique based on the reference pictures stored in picture memory 92.
[0280] Motion compensation unit 82 determines prediction information for video blocks of the current video slice by parsing the motion vectors and other syntax elements, and uses the prediction information to generate predictive blocks for the current video block being decoded. For example, motion compensation unit 82 may use one or more syntax elements in a parameter set to determine a prediction mode (e.g., intra prediction or inter prediction) used to code the video blocks of the video slice, an inter-prediction slice type (e.g., a B slice, a P slice, or a GPB slice), construction information about one or more reference picture lists for the slice, motion vectors for each inter-coded video block of the slice, inter-prediction status for each inter-coded video block of the slice, and other information for decoding video blocks in the current video slice.
[0281] Motion compensation unit 82 may perform the interpolation based on an interpolation filter. Motion compensation unit 82 may calculate interpolated values for sub-integer pixels of the reference block using an interpolation filter as used by encoding device 104 during encoding of the video block. In this case, motion compensation unit 82 may determine the interpolation filter used by encoding device 104 from a received syntax element and may use that interpolation filter to generate the predictive block.
[0282] Inverse quantization unit 86 inverse quantizes, or dequantizes, the quantized transform coefficients provided in the bitstream and decoded by entropy decoding unit 80. The inverse quantization process may include determining the degree of quantization using a quantization parameter calculated by encoding device 104 for each video block in a video slice, as well as determining the degree of inverse quantization to be applied. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT or other suitable inverse transform), an inverse integer transform, or a conceptually similar inverse transform process to the transform coefficients to generate residual blocks in the pixel domain.
[0283] After motion compensation unit 82 generates a predictive block for a current video block based on the motion vector and other syntax elements, decoding device 112 forms a decoded video block by adding a residual block from inverse transform processing unit 88 with a corresponding predictive block generated by motion compensation unit 82. Adder 90 represents one or more components that perform this addition operation. If desired, a loop filter (either in the coding loop or after the coding loop) may also be used to smooth pixel transitions or otherwise improve video quality. Filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although filter unit 91 is illustrated in FIG. 8 as being an in-loop filter, in other configurations filter unit 91 may be implemented as a post-loop filter. The decoded video blocks in a given frame or picture are then stored in picture memory 92, which stores reference pictures used for subsequent motion compensation. Picture memory 92 also stores decoded video for later presentation on a display device, such as video destination device 122 shown in FIG.
[0284] In this manner, the decoding device 112 of FIG. 13 represents an example of a video decoder configured to perform any of the techniques described herein, including the process described above with respect to FIG. 11.
[0285] The term "computer-readable medium" as used herein includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, storing, or transporting instructions and / or data. Computer-readable media may include non-transitory media on which data may be stored and does not include carrier waves and / or transitory electronic signals propagating wirelessly or via wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact disks (CDs) or digital versatile disks (DVDs), flash memory, memory, or memory devices. A computer-readable medium may store code and / or machine-executable instructions, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0286] In some aspects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when stated, non-transitory computer-readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.
[0287] Specific details are provided in the above description to provide a thorough understanding of the aspects and examples provided herein. However, it will be understood by those skilled in the art that the aspects may be practiced without these specific details. For clarity of explanation, in some instances, the technology may be presented as including individual functional blocks, including functional blocks comprising devices, device components, and steps or routines in a method embodied in software, or a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, or other components may be shown as components in block diagram form so as not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as to avoid obscuring the aspects.
[0288] Individual aspects may be described above as a process or method that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although the flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or to the main function.
[0289] The processes and methods according to the examples described above may be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a processing device to perform a particular function or group of functions, or that otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a particular function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, network-attached storage devices, and the like.
[0290] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, and the like. The functions described herein may also be embodied in peripheral devices or add-in cards. Such functions may also be implemented on a circuit board among different chips, or on different processes executing within a single device, as further examples.
[0291] The instructions, media for carrying such instructions, computing resources for executing such instructions, and other structures for supporting such computing resources are exemplary means for providing the functionality described in this disclosure.
[0292] In the above description, the aspects of the present application are described with reference to specific aspects thereof, but those skilled in the art will recognize that the present application is not limited thereto. Thus, while exemplary aspects of the present application have been described in detail herein, it should be understood that the inventive concepts may be embodied and utilized in different and various ways, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. The various features and aspects of the present application described above may be used individually or jointly. Moreover, the aspects may be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of the present specification. Thus, the present specification and drawings should be regarded as illustrative and not restrictive. For purposes of illustration, the methods have been described in a particular order. It should be understood that in alternative aspects, the methods may be performed in an order different from that described.
[0293] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with the less than or equal to ("≦") and greater than or equal to ("≧") symbols, respectively, without departing from the scope of the description.
[0294] When a component is described as being "configured to" perform some operation, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operation, by programming programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or any combination thereof.
[0295] The phrase "coupled to" refers to any component that is physically connected, either directly or indirectly, to another component and / or any component that is in communication, either directly or indirectly, with another component (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).
[0296] Claim language or other language in this disclosure reciting "at least one of" a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting "at least one of A and B" and "at least one of A or B" means A, B, or A and B. In another example, claim language reciting "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" and "at least one of A or B" can mean A, B, or A and B, and can, in addition, include unrecited items in the set of A and B.
[0297] The various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, the various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0298] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device, or separately as separate but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise a memory or data storage medium, such as a random access memory (RAM), such as a synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a FLASH memory, a magnetic or optical data storage medium, or the like. The techniques may additionally or alternatively be realized at least in part by a computer-readable communications medium, such as a propagated signal or wave, that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer.
[0299] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, the term "processor" as used herein may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated software or hardware modules configured for encoding and decoding, or may be incorporated into a composite video encoder-decoder (codec).
[0300] Illustrative aspects of the present disclosure include the following.
[0301] Aspect 1: An apparatus for processing video data, comprising: a memory; and one or more processors coupled to the memory. The one or more processors are configured to obtain a current picture of the video data, obtain a reference picture for the current picture from the video data, determine a merge mode candidate from the current picture, identify a first motion vector and a second motion vector for the merge mode candidate, select a motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, derive a refined motion vector from the first motion vector, the second motion vector, and the reference picture using the motion vector search strategy, and process the merge mode candidate using the refined motion vector.
[0302] Aspect 2: The apparatus of aspect 1, wherein the merge mode candidate is selected from a merge candidate list.
[0303] Aspect 3: The apparatus of aspect 2, wherein the merge candidate list is constructed from one or more of a spatial motion vector predictor from spatially adjacent blocks of the merge mode candidate, a temporal motion vector predictor from a co-located block of the merge mode candidate, a history-based motion vector predictor from a history table, a pairwise average motion vector predictor, and a zero-value motion vector.
[0304] Aspect 4: The apparatus of any of aspects 1 to 3, wherein the one or more processors are configured to generate a motion vector bi-prediction signal using a first motion vector and a second motion vector by averaging two prediction signals obtained from two different reference pictures.
[0305] Example 5: The apparatus of any one of examples 1 to 4, wherein the plurality of motion vector search strategies includes a fractional sample refinement strategy.
[0306] Example 6: The apparatus of example 5, wherein the plurality of motion vector search strategies includes a bi-predictive optical flow strategy.
[0307] Example 7: The apparatus of example 6, wherein the plurality of motion vector search strategies includes a sub-block based bilateral matching motion vector refinement strategy.
[0308] Example 8: The apparatus of any of examples 1 to 7, wherein the first motion vector and the second motion vector are associated with one or more constraints.
[0309] Aspect 9: The apparatus of aspect 8, wherein the one or more constraints include a mirroring constraint.
[0310] Example 10: The apparatus of any of examples 1 to 9, wherein the one or more constraints include a zero value constraint for the first motion vector differential.
[0311] Example 11: The apparatus of any of examples 1 to 9, wherein the one or more constraints include a zero value constraint for the second motion vector differential.
[0312] Example 12: The apparatus of any of examples 1 to 11, wherein the video data includes syntax indicating one or more constraints.
[0313] Example 13: The apparatus of any of examples 1 to 12, wherein the motion vector search strategy comprises a multi-pass decoder-side motion vector refinement strategy.
[0314] Example 14: The apparatus of example 13, wherein the multi-pass decoder-side motion vector refinement strategy includes two or more refinement passes of the same refinement type.
[0315] Example 15: The apparatus of example 14, wherein the multi-pass decoder-side motion vector refinement strategy includes one or more refinement passes of different types from the same refinement type.
[0316] Example 16: The apparatus of any of examples 14 or 15, wherein the two or more refinement passes of the same refinement type are block-based bilateral matching motion vector refinement, sub-block-based bilateral matching motion vector refinement, or sub-block-based bilateral optical flow motion vector refinement.
[0317] Example 17: The apparatus of any one of examples 1 to 16, wherein the plurality of motion vector search strategies comprises a plurality of subsets of multi-pass strategies.
[0318] Example 18: The apparatus of example 17, wherein a plurality of subsets of the multi-pass strategies are signaled in one or more syntax elements of the video data.
[0319] Example 19: The apparatus of any of examples 1 to 18, wherein deriving the refined motion vector comprises calculating matching costs for a plurality of motion vector candidate pairs according to a motion vector search strategy.
[0320] Example 20: The apparatus of any of examples 1 to 19, wherein the motion vector search strategy is adaptively selected based on a matching cost determined from the video data.
[0321] Example 21: The apparatus of any of examples 1 to 19, wherein the motion vector search strategy is selected to adaptively set a number of passes of the motion vector search strategy based on the video data.
[0322] Example 22: The apparatus of any of examples 1 to 19, wherein the motion vector search strategy is selected to adaptively set a search pattern for determining candidates for the refined motion vector based on the video data.
[0323] Example 23: The apparatus of any of examples 1 to 19, wherein a motion vector search strategy is selected to adaptively set a set of criteria for generating a list of candidates for refined motion vectors based on the video data.
[0324] Example 24: The apparatus of any one of examples 1 to 19, wherein the motion vector search strategy is adaptively performed based on a decoder-side motion vector refinement constraint from the video data.
[0325] Example 25: The apparatus of any of examples 1 to 19, wherein the motion vector search strategy is adaptively performed based on block sizes for merge mode candidates in the video data.
[0326]
[0031] Aspect 26: The apparatus of any one of aspects 1 to 25, wherein the one or more processors are configured to disable multiple hypothesis predictions.
[0327] Example 27: The apparatus of any of examples 1 to 26, wherein the one or more processors are configured to perform multiple hypothesis prediction in conjunction with a motion vector search strategy.
[0328]
[0031] Aspect 28: The apparatus of any of aspects 1 to 27, wherein the one or more processors are configured to generate a merge candidate list including the merge mode candidates.
[0329] Aspect 29: The apparatus of aspect 28, wherein to generate the merge candidate list, the one or more processors are configured to determine one or more default candidates to add to the merge candidate list based on the number of candidates in the merge candidate list being less than a maximum number of candidates based on one or more conditions related to the adaptive merge mode (e.g., conditions for adaptive_bm_mode).
[0330] Aspect 30: The apparatus of aspect 28, wherein to generate the merge candidate list, the one or more processors are configured to determine one or more default candidates to add to the merge candidate list based on the number of candidates in the merge candidate list being less than a maximum number of candidates based on one or more conditions associated with a constraint associated with the adaptive merge mode (e.g., a condition according to bm_dir).
[0331] Example 31: The apparatus of any one of examples 1 to 30, wherein the apparatus is a mobile device.
[0332] Example 32: The apparatus of any of examples 1 to 31, further comprising a camera configured to capture one or more frames.
[0333] Embodiment 33: The apparatus of any of embodiments 1 to 32, further comprising a display configured to display one or more frames.
[0334] Aspect 34: A method of processing video data according to any of the operations of aspects 1 to 33.
[0335] Aspect 35: A computer-readable storage medium comprising instructions that, when executed by one or more processors of a device, cause the device to perform any of the operations of aspects 1 to 33.
[0336] Example 36: An apparatus, comprising one or more means for performing any of the operations of examples 1 to 33.
[0337] Aspect 37: An apparatus for processing video data, comprising at least one memory and at least one processor coupled to the at least one memory, wherein the at least one processor is configured to obtain one or more reference pictures for a current picture, identify a first motion vector and a second motion vector for a merge mode candidate, determine a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, determine one or more refined motion vectors based on at least one of the first motion vector or the second motion vector and the one or more reference pictures using the selected motion vector search strategy, and process the merge mode candidate using the one or more refined motion vectors.
[0338] Example 38. The apparatus of example 37, wherein the selected motion vector search strategy is associated with one or more constraints based at least on at least one of the first motion vector or the second motion vector.
[0339] Example 39. The apparatus of example 38, wherein the one or more constraints are determined for a block of video data based on syntax elements signaled for the block.
[0340] Example 40. The apparatus of any of examples 38 or 39, wherein the one or more constraints are associated with at least one of a first motion vector differential associated with the first motion vector or a second motion vector differential associated with the second motion vector.
[0341] Embodiment 41. The apparatus of embodiment 40, wherein the one or more refined motion vectors include a first refined motion vector and a second refined motion vector, and wherein at least one processor is configured to determine a first motion vector differential as a difference between the first refined motion vector and the first motion vector, and to determine a second motion vector differential as a difference between the second refined motion vector and the second motion vector.
[0342] Example 42. The apparatus of any of examples 40 or 41, wherein the one or more constraints include a mirroring constraint for the first motion vector differential and the second motion vector differential, wherein the first motion vector differential and the second motion vector differential have the same magnitude and different signs.
[0343] Example 43. The apparatus of any of examples 40 to 42, wherein the one or more constraints include a zero value constraint for at least one of the first motion vector differential or the second motion vector differential.
[0344] Example 44. The apparatus of example 43, wherein the at least one processor is configured to determine one or more refined motion vectors using a selected motion vector search strategy by keeping a first one of the first motion vector differential or the second motion vector differential at a constant value and searching against a second one of the first motion vector differential or the second motion vector differential based on a zero value constraint.
[0345] Example 45. The apparatus of any of examples 37 to 44, wherein the selected motion vector search strategy is a bilateral matching (BM) motion vector search strategy.
[0346] Example 46. The apparatus of any of examples 37 to 45, wherein at least one processor is configured to determine one or more refined motion vectors based on one or more constraints associated with a selected motion vector search strategy, and wherein, to determine the one or more refined motion vectors based on the one or more constraints, the at least one processor is configured to determine a first refined motion vector by searching a first reference picture around the first motion vector based on the selected motion vector search strategy and determine a second refined motion vector by searching a second reference picture around the second motion vector based on the selected motion vector search strategy, and wherein the one or more constraints include a motion vector differential constraint.
[0347] Example 47. The apparatus of example 46, wherein to determine the first refined motion vector and the second refined motion vector, at least one processor is configured to minimize a difference between a first reference block associated with the first refined motion vector and a second reference block associated with the second refined motion vector.
[0348] Example 48. The apparatus of any of examples 37 to 47, wherein the plurality of motion vector search strategies includes at least two of a multi-pass decoder-side motion vector refinement strategy, a fractional sample refinement strategy, a bidirectional optical flow strategy, or a sub-block-based bilateral matching motion vector refinement strategy.
[0349] Example 49. The apparatus of any of examples 37 to 48, wherein the selected motion vector search strategy comprises a multi-pass decoder-side motion vector refinement strategy.
[0350] Example 50. The apparatus of example 49, wherein the multi-pass decoder-side motion vector refinement strategy includes at least one of one or more block-based bilateral matching motion vector refinement passes or one or more sub-block-based motion vector refinement passes.
[0351] Example 51. The apparatus of example 50, wherein at least one processor is configured to perform one or more block-based bilateral matching motion vector refinement passes using a first constraint associated with at least one of the first motion vector differential or the second motion vector differential, and to perform one or more sub-block-based motion vector refinement passes using a second constraint associated with at least one of the first motion vector differential or the second motion vector differential, wherein the first constraint differs from the second constraint.
[0352] Embodiment 52. The apparatus of any of embodiments 50 or 51, wherein the one or more sub-block-based motion vector refinement passes include at least one of a sub-block-based bilateral matching motion vector refinement pass or a sub-block-based bidirectional optical flow motion vector refinement pass.
[0353] Embodiment 53. The apparatus of any of embodiments 37 to 52, wherein the apparatus is a wireless communication device.
[0354] Example 54. The apparatus of any of examples 37 to 53, wherein at least one processor is configured to determine one or more refined motion vectors for a block of the video data, and the merge mode candidates include neighboring blocks of the block.
[0355] Aspect 55: A method for processing video data, comprising the steps of obtaining one or more reference pictures for a current picture; identifying a first motion vector and a second motion vector for a merge mode candidate; determining a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies; determining one or more refined motion vectors using the selected motion vector search strategy based on at least one of the first motion vector or the second motion vector and one or more reference pictures; and processing the merge mode candidate using the one or more refined motion vectors.
[0356] Embodiment 56. The method of embodiment 55, wherein the selected motion vector search strategy is associated with one or more constraints based at least on at least one of the first motion vector or the second motion vector.
[0357] Embodiment 57. The method of embodiment 56, wherein one or more constraints are determined for a block of video data based on syntax elements signaled for the block.
[0358] Embodiment 58. The method of any of embodiments 56 or 57, wherein one or more constraints are associated with at least one of a first motion vector differential associated with the first motion vector or a second motion vector differential associated with the second motion vector.
[0359] Example 59. The method of example 58, wherein the one or more refined motion vectors include a first refined motion vector and a second refined motion vector, further comprising a step of determining a first motion vector differential as a difference between the first refined motion vector and the first motion vector, and a step of determining a second motion vector differential as a difference between the second refined motion vector and the second motion vector.
[0360] Embodiment 60. The method of any of embodiments 58 or 59, wherein the one or more constraints include a mirroring constraint for the first motion vector differential and the second motion vector differential, wherein the first motion vector differential and the second motion vector differential have the same magnitude and different signs.
[0361] Embodiment 61. The method of any of embodiments 58 to 60, wherein the one or more constraints include a zero value constraint for at least one of the first motion vector differential or the second motion vector differential.
[0362] Aspect 62. The method of aspect 61, wherein one or more motion refined motion vectors are determined using a selected motion vector search strategy by keeping a first one of the first motion vector differential or the second motion vector differential at a constant value and searching against a second one of the first motion vector differential or the second motion vector differential based on a zero value constraint.
[0363] Embodiment 63. The method of any of embodiments 55 to 62, wherein the selected motion vector search strategy is a bilateral matching (BM) motion vector search strategy.
[0364] Embodiment 64. The method of any of embodiments 55 to 63, wherein the one or more refined motion vectors are determined based on one or more constraints associated with a selected motion vector search strategy, and the step of determining the one or more refined motion vectors based on the one or more constraints comprises: determining a first refined motion vector by searching a first reference picture around a first motion vector based on the selected motion vector search strategy; and determining a second refined motion vector by searching a second reference picture around a second motion vector based on the selected motion vector search strategy, wherein the one or more constraints include a motion vector differential constraint.
[0365] Example 65. The method of example 64, wherein the step of determining the first refined motion vector and the second refined motion vector comprises minimizing a difference between a first reference block associated with the first refined motion vector and a second reference block associated with the second refined motion vector.
[0366] Embodiment 66. The method of any of embodiments 55 to 65, wherein the multiple motion vector search strategies include at least two of a multi-pass decoder-side motion vector refinement strategy, a fractional sample refinement strategy, a bidirectional optical flow strategy, or a sub-block-based bilateral matching motion vector refinement strategy.
[0367] Example 67. The method of any of examples 55 to 66, wherein the selected motion vector search strategy comprises a multi-pass decoder-side motion vector refinement strategy.
[0368] Example 68. The method of example 67, wherein the multi-pass decoder-side motion vector refinement strategy includes at least one of one or more block-based bilateral matching motion vector refinement passes or one or more sub-block-based motion vector refinement passes.
[0369] Aspect 69. The method of aspect 68, further comprising: performing one or more block-based bilateral matching motion vector refinement passes using a first constraint associated with at least one of the first motion vector differential or the second motion vector differential; and performing one or more sub-block-based motion vector refinement passes using a second constraint associated with at least one of the first motion vector differential or the second motion vector differential, wherein the first constraint is different from the second constraint.
[0370] Embodiment 70. The method of any of embodiments 68 or 69, wherein the one or more sub-block-based motion vector refinement passes include at least one of a sub-block-based bilateral matching motion vector refinement pass or a sub-block-based bidirectional optical flow motion vector refinement pass.
[0371] Aspect 71: A method of processing video data according to any of the operations of aspects 37 to 70.
[0372] Aspect 72: A computer-readable storage medium comprising instructions that, when executed by one or more processors of a device, cause the device to perform any of the operations of aspects 37 to 70.
[0373] Example 73: An apparatus, comprising one or more means for performing any of the operations of examples 37 to 70. [Explanation of symbols]
[0374] 35 Division Units 41 Prediction Processing Unit 42 Motion Estimation Unit 44 Motion Compensation Unit 46 Intra Prediction Processing Units 50 Adder 52 Conversion Processing Unit 54 Quantization Units 56 Entropy Coding Units 57 After-treatment Devices 58 Inverse Quantization Unit 60 Inverse Transformation Processing Unit 62 Adder 63 Filter unit 64 Picture Memory 79 Network Entities 80 Entropy Decoding Unit 81 Prediction Processing Unit 82 Motion Compensation Unit 84 Intra Prediction Processing Unit 86 Inverse Quantization Unit 88 Inverse Transformation Processing Unit 90 Adder 91 Filter unit 92 Picture Memory 102 Video Sources 104 Encoding device, video encoding device 106 Encoder Engine 108 Storage 110 Output 112 Decryption Device 114 Input 116 Decoder Engine 118 Storage 120 Communication Links 122 Video Destination Device 402 Current Block 404 Reference Block 422 currently blocked 424 First Reference Block 426 Second Reference Block 500 processing blocks, blocks 610 Current Picture 612 Current CU 615 Current Reference Picture 630 Co-located Pictures 632 Same position CU 635 Co-located Reference Pictures 810 Current Picture 812 Bi-Predictive Merge Candidates 822 Second predictor, initial predictor 824 refined candidate blocks 830 1st reference picture 832 First predictor, initial predictor 834 refined candidate blocks 910 Subblock 970 lines 980 columns 1020 First Search Area 1030 Second Search Area 1040 Third Search Area 1050 Fourth Search Area
Claims
1. An apparatus for processing video data, comprising: at least one memory; at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain one or more reference pictures for a current picture; identify a first motion vector and a second motion vector for a merge mode candidate; determine a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, the selected motion vector search strategy being associated with one or more constraints based on at least one of the first motion vector or the second motion vector, the one or more constraints being associated with at least one of a first motion vector difference associated with the first motion vector or a second motion vector difference associated with the second motion vector, the one or more constraints including a zero value constraint for at least one of the first motion vector difference or the second motion vector difference; determine one or more refined motion vectors based on at least one of the first motion vector or the second motion vector and the one or more reference pictures using the selected motion vector search strategy, determining the one or more refined motion vectors using the selected motion vector search strategy being based on the zero value constraint, determining the one or more refined motion vectors using the selected motion vector search strategy including keeping a first one of the first motion vector difference or the second motion vector difference at a constant value and searching for a second one of the first motion vector difference or the second motion vector difference; process the merge mode candidate using the one or more refined motion vectors; and An apparatus configured to perform **Claim 2** The apparatus according to claim 1, wherein the one or more constraints are determined for the block of the video data based on a syntactic element signaled for the block. **Claim 3** The one or more refined motion vectors include a first refined motion vector and a second refined motion vector, and the at least one processor determines the first motion vector difference as a difference between the first refined motion vector and the first motion vector, The apparatus according to claim 1, configured to determine the second motion vector difference as a difference between the second refined motion vector and the second motion vector. **Claim 4** The apparatus according to claim 1, wherein the selected motion vector search strategy is a bilateral matching (BM) motion vector search strategy. **Claim 5** The at least one processor is configured to determine the one or more refined motion vectors based on one or more constraints associated with the selected motion vector search strategy, and to determine the one or more refined motion vectors based on the one or more constraints, the at least one processor determines a first refined motion vector by searching a first reference picture around the first motion vector based on the selected motion vector search strategy, configured to determine a second refined motion vector by searching a second reference picture around the second motion vector based on the selected motion vector search strategy, and the one or more constraints include a motion vector difference constraint, Optionally, to determine the first refined motion vector and the second refined motion vector, the at least one processor configured to minimize the difference between a first reference block associated with the first refined motion vector and a second reference block associated with the second refined motion vector The apparatus according to claim 4. **Claim 6** The apparatus according to claim 1, wherein the plurality of motion vector search strategies includes at least two of a multi-pass decoder-side motion vector refinement strategy, a fractional sample refinement strategy, a bidirectional optical flow strategy, or a sub-block-based bilateral matching motion vector refinement strategy. **Claim 7** The apparatus according to claim 1, wherein the selected motion vector search strategy comprises a multi-pass decoder-side motion vector refinement strategy. **Claim 8** The multi-pass decoder-side motion vector refinement strategy includes at least one of one or more block-based bilateral matching motion vector refinement passes or one or more sub-block-based motion vector refinement passes, Optionally, the at least one processor uses a first constraint associated with at least one of a first motion vector difference or a second motion vector difference to execute the one or more block-based bilateral matching motion vector refinement passes, and uses a second constraint associated with at least one of the first motion vector difference or the second motion vector difference to execute the one or more sub-block-based motion vector refinement passes, and / or Optionally, the one or more sub-block-based motion vector refinement passes include at least one of a sub-block-based bilateral matching motion vector refinement pass or a sub-block-based bidirectional optical flow motion vector refinement pass The apparatus according to claim 7. **Claim 9** A wireless communication device and / or the at least one processor is configured to determine the one or more refined motion vectors for blocks of the video data, and the merge mode candidate includes adjacent blocks of the block The apparatus according to claim 1.
10. A method for processing video data, comprising: obtaining one or more reference pictures for a current picture; identifying a first motion vector and a second motion vector for a merge mode candidate; determining a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, the selected motion vector search strategy being associated with one or more constraints based on at least one of the first motion vector or the second motion vector, the one or more constraints being associated with at least one of a first motion vector difference associated with the first motion vector or a second motion vector difference associated with the second motion vector, the one or more constraints including a zero value constraint for at least one of the first motion vector difference or the second motion vector difference; determining one or more refined motion vectors based on at least one of the first motion vector or the second motion vector and the one or more reference pictures using the selected motion vector search strategy, the step of determining the one or more refined motion vectors using the selected motion vector search strategy being based on the zero value constraint, the step of determining the one or more refined motion vectors using the selected motion vector search strategy including keeping a first one of the first motion vector difference or the second motion vector difference at a constant value and searching for a second one of the first motion vector difference or the second motion vector difference; A method comprising the step of processing the merge mode candidate using the one or more refined motion vectors.
11. The method according to claim 10, wherein the one or more constraints are determined for the block of the video data based on a syntax element signaled for the block.
12. The one or more refined motion vectors include a first refined motion vector and a second refined motion vector, and the method further comprises: Determining the first motion vector difference as a difference between the first refined motion vector and the first motion vector; The method according to claim 10, further comprising determining the second motion vector difference as a difference between the second refined motion vector and the second motion vector.
13. The selected motion vector search strategy is a bilateral matching (BM) motion vector search strategy, the one or more refined motion vectors are determined based on one or more constraints associated with the selected motion vector search strategy, and the step of determining the one or more refined motion vectors based on the one or more constraints comprises: Determining a first refined motion vector by searching a first reference picture around the first motion vector based on the selected motion vector search strategy; Determining a second refined motion vector by searching a second reference picture around the second motion vector based on the selected motion vector search strategy, and The one or more constraints include motion vector difference constraints, Optionally, the step of determining the first refined motion vector and the second refined motion vector is including the step of minimizing the difference between a first reference block associated with the first refined motion vector and a second reference block associated with the second refined motion vector The method according to claim 10.
14. the selected motion vector search strategy comprises a multi-pass decoder-side motion vector refinement strategy, the multi-pass decoder-side motion vector refinement strategy includes at least one of one or more block-based bilateral matching motion vector refinement paths or one or more sub-block-based motion vector refinement paths, and optionally, using a first constraint associated with at least one of a first motion vector difference or a second motion vector difference to execute the one or more block-based bilateral matching motion vector refinement paths; using a second constraint associated with at least one of the first motion vector difference or the second motion vector difference to execute the one or more sub-block-based motion vector refinement paths. The method according to claim 10. A computer-readable storage medium including instructions that, when executed by one or more processors of a device, cause the device to perform the operations of any one of claims 10 to 14.