Adaptive bilateral matching for decoder-side motion vector refinement

Adaptive bilateral matching for decoder-side motion vector refinement addresses the challenge of high-quality video compression by refining motion vectors, enhancing video quality and device performance through efficient bitrate usage.

JP7834783B2Active Publication Date: 2026-03-24QUALCOMM INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing video coding techniques face challenges in achieving high-quality video compression with efficient bitrate usage, particularly in decoder-side motion vector refinement, which affects video quality and device performance.

Method used

Implementing adaptive bilateral matching for decoder-side motion vector refinement (DMVR) using selected search strategies and constraints to refine motion vectors, improving accuracy and compression efficiency.

Benefits of technology

Enhances video quality and device performance by refining motion vectors through adaptive bilateral matching, leading to improved compression and reduced bitrate requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007834783000010
    Figure 0007834783000010
  • Figure 0007834783000011
    Figure 0007834783000011
  • Figure 0007834783000012
    Figure 0007834783000012
Patent Text Reader

Abstract

Systems and techniques are provided for processing video data. For example, the systems and techniques may include obtaining a current picture of the video data and obtaining a reference picture for the current picture from the video data. A merge mode candidate may be determined for the current picture. A first motion vector and a second motion vector may be identified for the merge mode candidate. A motion vector search strategy may be selected for the merge mode candidate from a plurality of motion vector search strategies. The selected motion vector search strategy may be associated with one or more constraints corresponding to at least one of the first motion vector or the second motion vector. The selected motion vector search strategy may be used to determine a refined motion vector based on the first motion vector, the second motion vector and the reference picture. The merge mode candidate may be processed using the refined motion vector.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure generally relates to video coding and decoding. For example, aspects of this disclosure include improving video coding techniques related to decoder-side motion vector refinement (DMVR) using bilateral matching. [Background technology]

[0002] Digital video capabilities can be incorporated into a wide range of devices, including digital television, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radio phones, so-called "smartphones," video teleconferencing devices, and video streaming devices. Such devices enable video data to be processed and output for consumption. Digital video data involves large amounts of data to meet the demands of consumers and video providers. For example, consumers of video data desire the highest quality video with high fidelity, resolution, frame rate, etc. As a result, the large amounts of video data required to meet these demands burden communication networks and devices that process and store video data. [Overview of the project] [Problems that the invention aims to solve]

[0003] A digital video device can implement video coding techniques for compressing video data. Video coding is performed according to one or more video coding standards or formats. For example, video coding standards or formats include, among others, versatile video coding (VVC), high-efficiency video coding (HEVC), advanced video coding (AVC), MPEG-2 Part 2 coding (MPEG represents moving picture experts group), and proprietary video codec / formats such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media. Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit redundancies present in the video image or sequence. The purpose of video coding techniques is to compress video data in a form that uses a lower bitrate while avoiding or minimizing degradation of video quality. As continuously evolving video services become available, more highly efficient coding techniques are required.

Means for Solving the Problem

[0004] In some examples, systems and techniques for decoder-side motion vector refinement (DMVR) using adaptive bilateral matching are described. According to an example for at least one description, an apparatus for processing video data is provided that includes at least one memory (e.g., configured to store data such as video data) and at least one processor (e.g., implemented in circuitry) coupled to the at least one memory. The at least one processor is configured to obtain one or more reference pictures for a current picture, identify a first motion vector and a second motion vector for a merge mode candidate, determine a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, determine one or more refined motion vectors based on at least one of the first motion vector or the second motion vector and the one or more reference pictures using the selected motion vector search strategy, and process the merge mode candidate using the one or more refined motion vectors, and can perform these actions.

[0005] In another example, a method for processing video data is provided. The method includes obtaining one or more reference pictures for a current picture, identifying a first motion vector and a second motion vector for a merge mode candidate, determining a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, determining one or more refined motion vectors based on at least one of the first motion vector or the second motion vector and the one or more reference pictures using the selected motion vector search strategy, and processing the merge mode candidate using the one or more refined motion vectors.

[0006] In another example, a non-temporary computer-readable medium is provided on which instructions are stored, and when the instructions are executed by one or more processors, the one or more processors cause one or more processors to obtain one or more reference pictures for the current picture, to identify a first motion vector and a second motion vector for merge mode candidates, to determine a selected motion vector search strategy for merge mode candidates from a plurality of motion vector search strategies, to determine one or more refined motion vectors using the selected motion vector search strategy based on at least one of the first or second motion vectors and one or more reference pictures, and to process the merge mode candidates using one or more refined motion vectors.

[0007] In another example, a device for processing video data is provided. The device includes means for obtaining one or more reference pictures for a current picture; means for identifying a first motion vector and a second motion vector for a merge mode candidate; means for determining a selected motion vector search strategy for a merge mode candidate from a plurality of motion vector search strategies; means for determining one or more refined motion vectors using the selected motion vector search strategy, based on at least one of the first or second motion vectors and one or more reference pictures; and means for processing the merge mode candidate using one or more refined motion vectors.

[0008] This summary is not intended to identify the main or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by referring to the entire specification of this patent, any or all of the drawings, and the appropriate portions of each claim.

[0009] The above, along with other features and embodiments, will become clearer with reference to the following specification, claims, and accompanying drawings.

[0010] Embodiments for the purpose of describing this application will be described in detail below with reference to the following figures. [Brief explanation of the drawing]

[0011] [Figure 1] This block shows some examples of encoding and decoding devices according to the present disclosure. [Figure 2A] This is a conceptual diagram illustrating exemplary spatially adjacent motion vector candidates for merge modes, based on several examples of the present disclosure. [Figure 2B] This is a conceptual diagram illustrating exemplary spatially adjacent motion vector candidates for an advanced motion vector prediction (AMVP) mode, based on several examples of the present disclosure. [Figure 3A] This is a conceptual diagram illustrating exemplary temporal motion vector predictor (TMVP) candidates, based on several examples of the present disclosure. [Figure 3B] This is a conceptual diagram illustrating examples of motion vector scaling, using several examples from the present disclosure. [Figure 4A] This is a conceptual diagram illustrating examples of adjacent samples of a current coding unit used to estimate motion compensation parameters for a current coding unit, based on several examples of the present disclosure. [Figure 4B] This is a conceptual diagram illustrating, by some examples of the present disclosure, an example of adjacent samples of a reference block currently used to estimate motion compensation parameters for a coding unit. [Figure 5] This figure shows, with some examples from the present disclosure, the locations of potential spatial merge candidates for use when processing blocks. [Figure 6] This figure shows some examples of motion vector scaling for time merge candidates for use when processing blocks, according to the present disclosure. [Figure 7]This figure shows some examples of time merge candidates for use when processing blocks, as shown in the present disclosure. [Figure 8] This figure shows some examples of bilateral matching in the present disclosure. [Figure 9] This figure shows some examples of bidirectional optical flow (BDOF) in the present disclosure. [Figure 10] This figure shows some examples of the search area regions in this disclosure. [Figure 11] This flowchart illustrates an exemplary process for decoder-side motion vector refinement using adaptive bilateral matching, based on several examples of the present disclosure. [Figure 12] This block diagram shows an exemplary video encoding device, illustrating some examples of the present disclosure. [Figure 13] An exemplary video decoding device, illustrated by some examples of the present disclosure, is shown in block A. [Modes for carrying out the invention]

[0012] Several aspects of this disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects may be applied independently, and some of them may be applied in combination. For illustrative purposes, specific details are provided in the following description to provide a complete understanding of the aspects of this application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and descriptions are not intended to be limiting.

[0013] The following description provides only illustrative embodiments and is not intended to limit the scope, applicability, or configuration of the Disclosure. Rather, the following description of the exemplary embodiments provides a description that enables the implementation of the exemplary embodiments for those skilled in the art. It should be understood that various modifications may be made to the function and configuration of the elements without departing from the spirit and scope of the application as set forth in the appended claims.

[0014] Video coding devices (e.g., encoding devices, decoding devices, or composite encoding-decoding devices) implement video compression techniques for efficiently encoding and / or decoding video data. Video compression techniques may include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques to reduce or eliminate redundancy inherent in the video sequence. A video encoder can divide each picture of the original video sequence into rectangular regions called video blocks or coding units (described in more detail below). These video blocks may be encoded using a specific prediction mode.

[0015] A video block can be divided in one or more ways into one or more groups of smaller blocks. A block may include a coding tree block, a prediction block, a transformation block, or other suitable blocks. Generally, a reference to a “block” may refer to such a video block (e.g., a coding tree block, a coding block, a prediction block, a transformation block, or any other suitable block or subblock as understood by those skilled in the art) unless otherwise specified. Furthermore, each of these blocks may also be interchangeably referred to herein as a “unit” (e.g., a coding tree unit (CTU), a coding unit, a prediction unit (PU), a transformation unit (TU), etc.). In some cases, a unit may refer to a coding logic unit encoded in a bitstream, while a block may refer to a portion of a video frame buffer that is the subject of the process.

[0016] In interprediction mode, the video encoder can search for a block similar to the one being encoded in a frame (or picture) located at a different temporal location, called a reference frame or reference picture. The video encoder may restrict this search to a certain spatial displacement from the block to be encoded. The best match may be identified using a two-dimensional (2D) motion vector that includes horizontal and vertical displacement components. In intraprediction mode, the video encoder may use spatial prediction techniques to form a predicted block based on data from previously encoded adjacent blocks within the same picture.

[0017] A video encoder can determine a prediction error. For example, the prediction may be determined as the difference between the pixel value in the encoded block and the predicted pixel value in the block. The prediction error may also be called the residual. The video encoder may also apply a transformation (e.g., a discrete cosine transform (DCT) or another suitable transformation) to the prediction error to generate transformation coefficients. After the transformation, the video encoder may quantize the transformation coefficients. The quantized transformation coefficients and motion vectors can be represented using syntax elements and, together with control information, form a coded representation of the video sequence. In some cases, the video encoder may entropy code the syntax elements, thereby further reducing the number of bits required for their representation.

[0018] A video decoder can use the syntax elements and control information discussed above to construct prediction data (e.g., a prediction block) for decoding the current frame. For example, a video decoder may add the prediction block to a compressed prediction error. The video decoder may determine the compressed prediction error by weighting the transformation basis function using quantized coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.

[0019] This specification describes systems, apparatus, processes (also called methods), and computer-readable media (collectively referred to herein as “systems and techniques”) for improving the accuracy of one or more motion vectors that may be used by a video coding device (e.g., a video decoder or decoding device) when performing prediction techniques (e.g., interpredictive modes). For example, systems and techniques can perform bilateral matching for decoder-side motion vector refinement (DMVR). Bilateral matching is a technique for refining a pair of two initial motion vectors. Such refinement may occur with a search around the pair of initial motion vectors to derive an updated motion vector that minimizes the block matching cost. The block matching cost can be generated in various ways, including using a sum of absolute difference (SAD) criterion, a sum of absolute transformed difference (SATD) criterion, a sum of square error (SSE) criterion, or other such criteria. The embodiments described herein can improve the accuracy of the motion vectors of the dual predictive merge candidates, resulting in improved video quality and associated device performance of devices operating according to the embodiments described herein.

[0020] In some embodiments, systems and techniques can be used to perform adaptive bilateral matching for DMVRs. For example, systems and techniques can perform bilateral matching using different search strategies and / or search parameters for different coded blocks. As will be described in more detail below, adaptive bilateral matching for DMVRs may be based on a selected search strategy determined or signaled for a given block. The selected search strategy may include one or more constraints for the bilateral matching search process. In some examples, the selected search strategy may include, as an addition or alternative, one or more constraints for a first motion vector difference and / or a second motion vector difference. In some examples, the selected search strategy may include one or more constraints between the first motion vector difference and the second motion vector difference.

[0021] In some embodiments, constraints are selected for the motion vector to be refined. These constraints may be mirroring constraints, zero constraints for a first vector, zero constraints for a second vector, or other types of constraints. In some cases, the constraints are applied to merge-mode coded blocks within merge candidates that satisfy one or more DMVR conditions. These constraints may then be used in conjunction with one or more search strategies to identify candidates and select the refined motion vector.

[0022] In some embodiments, different search strategies are used. Search strategies can be grouped into multiple subsets, each subset containing one or more search strategies. In some cases, the decoder can use syntax elements to determine the selected subset. For example, an encoder may include syntax elements in the bitstream. In such an example, the decoder may receive the bitstream and decode the syntax elements from the bitstream. The decoder can use syntax elements to determine a selected subset and any associated constraints for a given one or more blocks of video data contained in the bitstream. Using the selected subset and any associated constraints, the decoder can process motion vectors (e.g., two motion vectors of a bipredictive merge candidate) to identify an refined motion vector. In an embodiment for one description, an adaptive bilateral mode is provided, in which the coding device signals selected motion information candidates that satisfy the associated DMVR conditions (e.g., together with a signaling structure as part of a new adaptive bilateral mode).

[0023] Using the search strategies and associated constraints described above can provide improvements to decoder-side motion vector refinement, such as by providing adaptive bilateral motion vector refinement using selectable search algorithms and associated constraints. Such improvements to decoder-side motion vector refinement can be used with various video codecs, such as enhanced compression model (ECM) implementations. Examples described herein include implementation forms applied to multipath DMVRs to improve ECM systems operating according to one or more video coding standards. The techniques described herein can be implemented using one or more coding devices, including one or more coding devices, decoding devices, or composite coding-decoding devices. Coding devices can be implemented by one or more player devices, such as mobile devices, augmented reality (XR) devices, vehicles or vehicle computing systems, server devices or systems (e.g., distributed server systems including multiple servers, single server devices or systems), or other devices or systems.

[0024] The systems and techniques described herein may be applied to any existing video codec, any video codec under development, and / or any future video coding standard, including, but not limited to, High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Multipurpose Video Coding (VVC), VP9, ​​AOMedia Video1 (AV1) format / codec, and / or other existing, developing, or future video coding standards. The systems and techniques described herein can improve the operation of communication systems and devices in a system by improving the performance of video data transfer by devices, along with improved compression and associated video quality, based on improved motion vector selection from adaptive bilateral matching as described herein.

[0025] Figure 1 is a block diagram showing an example of a system 100 including an encoding device 104 and a decoding device 112. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device. The source device and / or receiving device may include electronic devices such as mobile or fixed telephone handsets (e.g., smartphones, mobile phones, etc.), desktop computers, laptop or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices. In some examples, the source device and receiving device may include one or more wireless transceivers for wireless communication. The coding techniques described herein are applicable to video coding in a variety of multimedia applications, including (e.g., streaming video transmission over the Internet), television broadcasting or transmission, coding of digital video for storage on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. The term coding as used herein may refer to coding and / or decoding. In some cases, system 100 can support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video calls.

[0026] The encoding device 104 (or encoder) may be used to encode video data using a video coding standard, format, codec, or protocol for generating an encoded video bitstream. Examples of video coding standards and formats / codecs include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, including its Scalable Video Coding (SVC) and Multiview Video Coding (MVC) extensions, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), High Efficiency Video Coding (HEVC) or ITU-T H.265, and Multipurpose Video Coding (VVC) or ITU-T H.266. Various extensions of HEVC exist that handle multi-layer video coding, including range and screen content coding extensions, 3D video coding (3D-HEVC), and multi-view extensions (MV-HEVC) and scalable extensions (SHVC). HEVC and its extensions are developed by the Joint Collaboration Team on Video Coding (JCT-VC), as well as the Joint Collaboration Team on 3D Video Coding Extension Development (JCT-3V) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). VP9, ​​AOMedia Video1 (AV1) developed by the Alliance for Open Media Alliance of Open Media (AOMedia), and Essential Video Coding (EVC) are other video coding standards to which the techniques described herein may be applied.

[0027] The techniques described herein may be applied to existing video codecs (e.g., High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), or other suitable existing video codecs) and / or may be efficient coding tools for any video coding standard, including video coding standards under development, and / or future video coding standards, such as VVC and / or other video coding standards under development or to be developed in the future. For example, the examples described herein may be performed using video codecs such as VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein may also be applicable to other coding standards, codecs, or formats, such as MPEG, JPEG (or other coding standards for still images), VP9, ​​AV1, their extensions, or other suitable coding standards that are already available or are not yet available or under development. For example, in some cases, the encoding device 104 and / or decoding device 112 may operate according to a proprietary video codec / format, such as AV1, an extension of AVI, and / or a successor version of AV1 (e.g., AV2), or other proprietary formats or industry standards. Therefore, while the techniques and systems described herein may be described in relation to a particular video coding standard, those skilled in the art will understand that the descriptions should not be construed as applying only to that particular standard.

[0028] Referring to Figure 1, the video source 102 may provide video data to the encoding device 104. The video source 102 may be part of a source device or part of a device other than a source device. The video source 102 may include a video capture device (e.g., a video camera, camera phone, video phone, etc.), a video archive containing stored video, a video server or content provider that provides video data, a video feed interface that receives video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.

[0029] The video data from video source 102 may include one or more input pictures or input frames. A picture or frame is a still image that is possibly part of the video. In some examples, the data from video source 102 may be a still image that is not part of the video. In HEVC, VVC, and other video coding specifications, a video sequence may include a series of pictures. A picture may include three sample arrays denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples, SCb is a two-dimensional array of Cb chrominance samples, and SCr is a two-dimensional array of Cr chrominance samples. Chrominance samples are sometimes referred to herein as “chroma” samples. A pixel can refer to all three components (luma samples and chroma samples) at a given location in the array of pictures. In other cases, a picture may be monochrome and may include only arrays of luma samples, in which case the terms pixel and sample may be used interchangeably. With regard to the exemplary techniques described herein, referring to individual samples for illustrative purposes, the same techniques may be applied to pixels (for example, all three sample components for a given location in an array of pictures). With regard to the exemplary techniques described herein, referring to pixels (for example, all three sample components for a given location in an array of pictures) for illustrative purposes, the same techniques may be applied to individual samples.

[0030] The encoder engine 106 (or encoder) of the encoding device 104 encodes video data to produce an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more coded video sequences. A coded video sequence (CVS) includes a series of access units (AUs), starting with an access unit (AU) having a random access point picture with certain characteristics in the base layer, up to immediately before the next AU having a random access point picture with certain characteristics in the base layer. For example, some characteristics of a random access point picture that initiates a CVS may include a RASL flag equal to 1 (e.g., NoRaslOutputFlag). Otherwise, a random access point picture (with a RASL flag equal to 0) will not initiate a CVS. An access unit (AU) includes one or more coded pictures and control information corresponding to coded pictures that share the same output time. Coded slices of pictures are encapsulated at the bitstream level into data units called network abstraction layer (NAL) units. For example, an HEVC video bitstream may contain one or more CVSs, each containing a NAL unit. Each NAL unit has a NAL unit header. In one example, the header is 1 byte for H.264 / AVC (except for multilayer extensions) and 2 bytes for HEVC. The syntax elements within the NAL unit header take the specified bits and are therefore recognizable to all kinds of systems and, in particular, to transport layers such as transport streams, Real-time Transport (RTP) protocols, and file formats.

[0031] Two classes of NAL units exist in the HEVC standard: video coding layer (VCL) NAL units and non-VCL NAL units. A VCL NAL unit contains coded picture data that forms a coded video bitstream. For example, a sequence of bits that forms a coded video bitstream resides in a VCL NAL unit. A VCL NAL unit may contain a slice or slice segment (described below) of coded picture data, while a non-VCL NAL unit contains control information about one or more coded pictures. In some cases, a NAL unit may be called a packet. An HEVC AU includes a VCL NAL unit containing coded picture data and a non-VCL NAL unit corresponding to the coded picture data (if any). A non-VCL NAL unit may contain a parameter set with high-level information about the coded video bitstream, in addition to other information. For example, the parameter set may include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). In some cases, each slice or other portion of the bitstream may refer to a single active PPS, SPS, and / or VPS to allow the decoding device 112 to access information that can be used to decode the slice or other portion of the bitstream.

[0032] A NAL unit may include a sequence of bits that form a coded representation of video data (e.g., an encoded video bitstream, a CVS of a bitstream, etc.), such as a coded representation of a picture in a video. The encoder engine 106 generates a coded representation of a picture by dividing each picture into multiple slices. Slices are independent of each other so that the information in a slice is coded without depending on data from other slices in the same picture. A slice includes one or more slice segments, each containing an independent slice segment and, if present, one or more dependent slice segments that depend on a previous slice segment.

[0033] In HEVC, a slice is then divided into coding tree blocks (CTBs) of luma samples and chroma samples. A luma sample CTB, and one or more chroma sample CTBs, along with the syntax for the samples, are called coding tree units (CTUs). A CTU is sometimes also called a “tree block” or “largest coding unit” (LCU). A CTU is the basic processing unit for HEVC coding. A CTU can be divided into multiple coding units (CUs) of varying sizes. A CU contains luma sample arrays and chroma sample arrays called coding blocks (CBs).

[0034] Luma CBs and chroma CBs can be further divided into prediction blocks (PBs). A PB is a block of samples of lumina or chroma components that uses the same motion parameters for inter-prediction or intra-block copy (IBC) prediction (when available or enabled for use). Luma PBs and one or more chroma PBs, along with their associated syntax, form a prediction unit (PU). For inter-prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU and used for inter-prediction of the lumina PBs and one or more chroma PBs. Motion parameters can also be called motion information. CBs can also be divided into one or more transformation blocks (TBs). A TB represents a rectangular block of color component samples to which a residual transformation (e.g., possibly the same 2D transformation) is applied to code the prediction residual signal. A transformation unit (TU) represents the TBs of lumina and chroma samples, as well as the corresponding syntax elements. Transform coding is described in more detail below.

[0035] The size of a CU corresponds to the size of the coding mode and can be square in shape. For example, the size of a CU can be 8×8 samples, 16×16 samples, 32×32 samples, 64×64 samples, or any other suitable size up to the size of the corresponding CTU. The phrase "N×N" is used herein to refer to the pixel dimensions of a video block with respect to the vertical and horizontal dimensions (e.g., 8 pixels × 8 pixels). Pixels within a block can be arranged in rows and columns. In some implementations, a block may not have the same number of pixels horizontally as it does vertically. Syntax data associated with a CU may, for example, describe the division of a CU into one or more PUs. The division mode may differ depending on whether the CU is intra-predictive mode coded or inter-predictive mode coded. PUs may be divided such that their shape is non-square. Syntax data associated with a CU may also describe the division of a CU into one or more TUs according to a CTU. TUs may be square or non-square in shape.

[0036] According to the HEVC standard, the conversion can be performed using a conversion unit (TU). A TU can be different for different CUs. A TU can be sized based on the size of the PU within a given CU. A TU may be the same size as the PU, or it may be smaller than the PU. In some examples, residual samples corresponding to a CU can be subdivided into smaller units using a quadtree structure known as a residual quadtree (RQT). The leaf nodes of the RQT may correspond to a TU. The pixel difference values ​​associated with the TU can be converted to generate conversion coefficients. These conversion coefficients can then be quantized by the encoder engine 106.

[0037] Once the video data picture is divided into CUs, the encoder engine 106 predicts each PU using a prediction mode. The prediction unit or prediction block is then subtracted from the original video data to obtain a residual (described below). For each CU, the prediction mode may be signaled within the bitstream using syntax data. The prediction mode may include intra-prediction (or intra-picture prediction) or inter-prediction (or inter-picture prediction). Intra-prediction utilizes correlations between spatially adjacent samples within a picture. For example, with intra-prediction, each PU is predicted from adjacent image data within the same picture, for example, using a DC prediction to find the mean for the PU, a plane prediction to fit a flat surface to the PU, a direction prediction to extrapolate from adjacent data, or any other suitable type of prediction. Inter-prediction uses temporal correlations between pictures to derive motion-compensated predictions for blocks of image samples. For example, with inter-prediction, each PU is predicted using motion-compensated predictions from image data in one or more reference pictures (before or after the current picture in the output order). The decision of whether to code a picture area using cross-picture prediction or intra-picture prediction can be made, for example, at the CU level.

[0038] The encoder engine 106 and decoder engine 116 (described in more detail below) may be configured to operate according to VVC. According to VVC, a video coder (such as the encoder engine 106 and / or decoder engine 116) divides a picture into multiple coding tree units (CTUs) (one or more CTBs for luma samples and chroma samples, along with the syntax for the samples, are called CTUs). The video coder may divide the CTUs according to a tree structure such as a quadwood-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple division types, such as the distinction between CU, PU, ​​and TU in HEVC. The QTBT structure includes two levels, a first level divided according to quadwood divisions and a second level divided according to binary tree divisions. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0039] In an MTT partitioned structure, blocks can be partitioned using quadtree partitions, binary tree partitions, and one or more types of ternary tree partitions. A ternary tree partition is a partition in which a block is divided into three subblocks. In some examples, a ternary tree partition divides a block into three subblocks without separating the original block through a central point. The partition types in an MTT (e.g., quadtrees, binary trees, and ternary trees) can be symmetric or asymmetric.

[0040] When operating according to the AV1 codec, the encoding device 104 and decoding device 112 may be configured to code video data in blocks. In AV1, the largest coding block that can be processed is called the superblock. In AV1, the superblock can be either 128×128 lumasamps or 64×64 lumasamps. However, in successor video coding formats (e.g., AV2), the superblock may be defined by a different (e.g., larger) lumasamp size. In some examples, the superblock is the top level of a block quadtree. The encoding device 104 may further subdivide the superblock into smaller coding blocks. The encoding device 104 may subdivide the superblock and other coding blocks into smaller blocks using rectangular or non-rectangular partitions. Non-rectangular blocks may include N / 2×N, N×N / 2, N / 4×N, and N×N / 4 blocks. The encoding device 104 and decoding device 112 may perform separate prediction and transformation processes for each coding block.

[0041] AV1 also defines tiles of video data. A tile is a rectangular array of superblocks that can be coded independently of other tiles. That is, the encoding device 104 and decoding device 112 can encode and decode coding blocks within a tile, respectively, without using video data from other tiles. However, the encoding device 104 and decoding device 112 can perform filtering across tile boundaries. Tiles may be uniform or non-uniform in size. Tile-based coding can enable parallel processing and / or multithreading for the implementation of encoders and decoders.

[0042] In some examples, a video coder may use a single QTBT or MTT structure to represent each of the luminance and saturation components, while in other examples, a video coder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luminance component and another QTBT or MTT structure for both saturation components (or two QTBT and / or MTT structures for each saturation component).

[0043] The video coder can be configured to use a quadtree partition, QTBT partition, MTT partition, superblock partition, or other partitioning structure.

[0044] In some examples, slice types are assigned to one or more slices of a picture. Slice types include intra-coded slices (I slices), intercoded P slices, and intercoded B slices. An I slice (independently decodeable intra-coded frame) is a slice of a picture that is coded using intra-prediction only, and is therefore independently decodeable because an I slice only requires data in the frame to predict any prediction unit or prediction block in the slice. A P slice (unidirectionally predicted frame) is a slice of a picture that can be coded using intra-prediction and unidirectional inter-prediction. Each prediction unit or prediction block in a P slice is coded using either intra-prediction or inter-prediction. When inter-prediction is applied, the prediction unit or prediction block is predicted by only one reference picture, and therefore the reference sample is from only one reference region in one frame. A B slice (bidirectionally predicted frame) is a slice of a picture that can be coded using intra-prediction and inter-prediction (e.g., either bidirectional or unidirectional). A prediction unit or prediction block of a B-slice may be predicted bidirectionally from two reference pictures, where each picture contributes to one reference region, and the sample sets of the two reference regions are weighted (e.g., with equal or different weights) to generate the prediction signal for the bidirectional prediction block. As described above, slices of one picture are coded independently. In some cases, a picture may be coded as just one slice.

[0045] As mentioned above, intra-picture prediction of a picture utilizes correlations between spatially adjacent samples within the picture. There are multiple intra-prediction modes (also called "intra-modes"). In some examples, Lumablock's intra-prediction includes 35 modes, including a planar mode, a DC mode, and 33 angular modes (e.g., a diagonal intra-prediction mode and angular modes adjacent to the diagonal intra-prediction mode). The 35 modes of intra-prediction are indexed as shown in Table 1 below. In other examples, more intra-modes may be defined, including predicted angles that may not yet be represented by the 33 angular modes. In other examples, the predicted angles associated with the angular modes may differ from those used in HEVC.

[0046] [Table 1]

[0047] Interpicture prediction uses time correlation between pictures to derive motion-compensated predictions for blocks of image samples. Using a translational motion model, the position of a block in a pre-decoded picture (reference picture) is indicated by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement and Δy specifies the vertical displacement of the reference block relative to the current block position. In some cases, the motion vector (Δx, Δy) may be integer sample precision (also called integer precision), in which case the motion vector points to an integer Pell grid (or integer pixel sampling grid) of the reference frame. In some cases, the motion vector (Δx, Δy) may be fractional sample precision (also called fractional Pell precision or non-integer precision) to more accurately capture the motion of the underlying object, without being limited to an integer Pell grid of the reference frame. The precision of the motion vector can be expressed by the quantization level of the motion vector. For example, the quantization level may be integer precision (e.g., 1 pixel) or fractional Pell precision (e.g., 1 / 4 pixel, 1 / 2 pixel, or other subpixel values). When the corresponding motion vector has fractional sample precision, interpolation is applied to the reference picture to derive the prediction signal. For example, to estimate values ​​at fractional positions, available samples at integer positions may be filtered (e.g., using one or more interpolation filters). Previously decoded reference pictures are indicated by their reference index (refIdx) in the reference picture list. The motion vector and reference index may be called motion parameters. Two types of picture-to-picture predictions can be performed, including single and bi-prediction.

[0048] In interpretation using biprediction (also called bidirectional interpretation), two sets of motion parameters (Δx0, y0, refIdx0 and Δx1, y1, refIdx1) are used to generate two motion-compensated predictions (from the same or possibly different reference pictures). For example, in biprediction, each prediction block uses two motion-compensated prediction signals to generate a B-prediction unit. The two motion-compensated predictions are then combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions may be combined by averaging. In another example, weighted prediction can be used, in which case different weights may be applied to each motion-compensated prediction. The reference pictures that may be used in biprediction are stored in two separate lists, denoted as List 0 and List 1. Motion parameters may be derived in the encoder using a motion estimation process.

[0049] In interpretation using single predictions (also called unidirectional interpretations), a single set of motion parameters (Δx0, y0, refIdx0) is used to generate motion-compensated predictions from a reference picture. For example, in single prediction, each prediction block uses at most one motion-compensated prediction signal to generate P prediction units.

[0050] A PU may contain data about the prediction process (e.g., motion parameters or other appropriate data). For example, when a PU is encoded using intra-prediction, the PU may contain data describing the intra-prediction mode of the PU. As another example, when a PU is encoded using inter-prediction, the PU may contain data defining the motion vector of the PU. The data defining the motion vector of the PU may describe, for example, the horizontal component (Δx) of the motion vector, the vertical component (Δy) of the motion vector, the resolution of the motion vector (e.g., integer precision, 1 / 4 pixel precision, or 1 / 8 pixel precision), the reference picture pointed to by the motion vector, the reference index, the reference picture list of the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.

[0051] AV1 includes two common techniques for encoding and decoding coding blocks of video data. The two common techniques are intra-prediction (e.g., intra-frame prediction or spatial prediction) and inter-prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when predicting a block of the current frame of video data using the intra-prediction mode, the encoding device 104 and decoding device 112 do not use video data from other frames of the video data. In most intra-prediction modes, the video encoding device 104 encodes the block of the current frame based on the difference between the sample value in the current block and the predicted value generated from a reference sample in the same frame. The video encoding device 104 determines the predicted value generated from the reference sample based on the intra-prediction mode.

[0052] After performing predictions using intra-prediction and / or inter-prediction, the encoding device 104 can perform transformations and quantization. For example, following predictions, the encoder engine 106 may compute residual values ​​corresponding to the PU. The residual values ​​may consist of pixel difference values ​​between the current block (PU) of the pixels being coded and the prediction block used to predict the current block (e.g., the predicted version of the current block). For example, after generating a prediction block (e.g., emitting an inter-prediction or intra-prediction), the encoder engine 106 may generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block contains a set of pixel difference values ​​that quantify the difference between the pixel values ​​of the current block and the pixel values ​​of the prediction block. In some examples, the residual block may be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of the pixel values.

[0053] Any residual data that may remain after the prediction has been performed is transformed using a block transform, which may be based on a discrete cosine transform, discrete sine transform, integer transform, wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., of sizes 32×32, 16×16, 8×8, 4×4, or other suitable sizes) may be applied to the residual data in each CU. In some embodiments, a TU may be used for the transformation and quantization processes implemented by the encoder engine 106. A given CU having one or more PUs may also include one or more TUs. As will be described in more detail below, the residual values ​​may be transformed into transformation coefficients using a block transform, and then quantized and scanned using a TU to generate serialized transformation coefficients for entropy coding.

[0054] In some embodiments, following intra-predictive coding or inter-predictive coding using the PU of the CU, the encoder engine 106 may compute residual data for the CU relative to the TU. The PU may comprise pixel data in the spatial domain (or pixel domain). The TU may comprise coefficients in the transformation domain after applying a block transformation. As previously stated, the residual data may correspond to the pixel difference between the pixels of the unencoded picture and the predicted values ​​corresponding to the PU. The encoder engine 106 may form a TU containing the residual data for the CU, and then transform the TU to generate transformation coefficients for the CU.

[0055] The encoder engine 106 may perform quantization of the conversion coefficients. Quantization achieves further compression by quantizing the conversion coefficients to reduce the amount of data used to represent the coefficients. For example, quantization may reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient with an n-bit value may be truncated to an m-bit value during quantization, where n is greater than m.

[0056] Once quantization is performed, the coded video bitstream contains quantized transformation coefficients, prediction information (e.g., prediction mode, motion vector, block vector, etc.), piecewise information, and any other appropriate data such as other syntax data. Various elements of the coded video bitstream can then be entropy coded by the encoder engine 106. In some examples, the encoder engine 106 may scan the quantized transformation coefficients using a predetermined scan order to generate a serialized vector that can be entropy coded. In some examples, the encoder engine 106 may perform an adaptive scan. After scanning the quantized transformation coefficients to form a vector (e.g., a one-dimensional vector), the encoder engine 106 may entropy coded the vector. For example, the encoder engine 106 may use context-adaptive variable-length coding, context-adaptive binary arithmetic coding, syntax-based context-adaptive binary arithmetic coding, probability interval piecewise entropy coding, or another appropriate entropy coding technique.

[0057] The output 110 of the encoding device 104 may transmit NAL units constituting the encoded video bitstream data to the decoding device 112 of the receiving device via the communication link 120. The input 114 of the decoding device 112 may receive the NAL units. The communication link 120 may include channels provided by a wireless network, a wired network, or a combination of a wired network and a wireless network. The wireless network may include any wireless interface or combination of wireless interfaces, and may include any suitable wireless network (e.g., the Internet or other wide area networks, packet-based networks, WiFi®, radio frequency (RF), ultra-wideband (UWB), WiFi-Direct, cellular, Long-Term Evolution (LTE), WiMax®, etc.). The wired network may include any wired interface (e.g., fiber, Ethernet, power line Ethernet, Ethernet over coaxial cable, digital signaling circuit (DSL), etc.). Wired and / or wireless networks can be implemented using various devices such as base stations, routers, access points, bridges, gateways, and switches. Encoded video bitstream data can be modulated according to communication standards such as wireless communication protocols and transmitted to a receiving device.

[0058] In some examples, the encoding device 104 may store the encoded video bitstream data in storage 108. Output 110 may retrieve the encoded video bitstream data from the encoder engine 106 or from storage 108. Storage 108 may include any of a variety of distributed or locally accessed data storage media. For example, storage 108 may include a hard drive, storage disk, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. Storage 108 may also include a decoded picture buffer (DPB) for storing a reference picture for use in interpretation. In further examples, storage 108 may correspond to a file server or another intermediate storage device capable of storing the encoded video generated by the source device. In such cases, a receiving device including a decoding device 112 can access the stored video data from the storage device via streaming or download. The file server may be any type of server capable of storing the encoded video data and transmitting that encoded video data to the receiving device. An exemplary file server may include a web server (for a website, for example), an FTP server, a network-attached storage (NAS) device, or a local disk drive. A receiving device may access the encoded video data through any standard data connection, including an internet connection. This may include wireless channels (e.g., Wi-Fi connection), wired connections (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. Transmission of the encoded video data from storage 108 may be streaming transmission, download transmission, or a combination thereof.

[0059] The input 114 of the decoding device 112 may receive encoded video bitstream data and provide the video bitstream data to the decoder engine 116 or to storage 118 for later use by the decoder engine 116. For example, storage 118 may include a decoded picture buffer (DPB) for storing reference pictures for use in interpretation. A receiving device including the decoding device 112 may receive encoded video data to be decoded via storage 108. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to a receiving device. The communication medium for transmitting the encoded video data may comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from a source device to a receiving device.

[0060] The decoder engine 116 may decode the encoded video bitstream data by entropy decoding (for example, using an entropy decoder) and by extracting elements of one or more coded video sequences that make up the encoded video data. The decoder engine 116 may then rescale the encoded video bitstream data and perform an inverse transform on the encoded video bitstream data. The residual data is then passed to the prediction stage of the decoder engine 116. The decoder engine 116 then predicts blocks of pixels (for example, PU). In some examples, the predictions are added to the output of the inverse transform (residual data).

[0061] The video decoding device 112 may output the decoded video to a video destination device 122, which may include a display or other output device for displaying the decoded video data to a consumer of the content. In some embodiments, the video destination device 122 may be part of a receiving device that includes the decoding device 112. In some embodiments, the video destination device 122 may be part of a separate device other than the receiving device.

[0062] In some embodiments, the video encoding device 104 and / or the video decoding device 112 may be integrated with an audio encoding device and an audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 may also include other hardware or software necessary to implement the coding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. The video encoding device 104 and the video decoding device 112 may be integrated in each device as part of a composite encoder / decoder (codec).

[0063] The exemplary system shown in Figure 1 is an example for one description that may be used herein. Techniques for processing video data using the techniques described herein can be performed by any digital video encoding and / or decoding device. Generally, the techniques of this disclosure are performed by video encoding devices or video decoding devices, but these techniques can also be performed by a composite video encoder / decoder, usually called a “codec”. Furthermore, the techniques of this disclosure can also be performed by a video preprocessor. Source devices and receiving devices are merely examples of coding devices such that the source device generates coded video data to send to the receiving device. In some examples, the source devices and receiving devices may operate substantially symmetrically such that each device includes video encoding and decoding components. Thus, exemplary systems may support one-way or two-way video transmission between video devices for, for example, video streaming, video playback, video broadcasting, or video phone calls.

[0064] This disclosure may refer in general to “signaling” certain information, such as syntax elements. The term “signaling” may generally refer to the communication of values ​​for syntax elements and / or other data used to decode encoded video data. For example, video encoding device 104 may signal values ​​for syntax elements in a bitstream. In general, signaling refers to generating values ​​in a bitstream. As stated above, video source 102 may transport the bitstream to video destination device 122 in substantially real time, or not in real time, such as when storing syntax elements in storage 108 for later retrieval by video destination device 122.

[0065] A video bitstream may also contain Supplemental Enhancement Information (SEI) messages. For example, an SEI NAL unit may be part of a video bitstream. In some cases, SEI messages may contain information that is not required by the decoding process. For example, information in an SEI message may not be essential for the decoder to decode the video picture in the bitstream, but the decoder may use that information to improve the display or processing of the picture (e.g., the decoded output). Information in an SEI message may be embedded metadata. In one example for explanation, information in an SEI message may be used by a decoder-side entity to improve the readability of the content. In some cases, certain application standards may require the presence of such SEI messages in the bitstream so that quality improvements can be brought to all devices conforming to the application standard (e.g., carrying frame-packing SEI messages for frame-compatible plano-stereoscopic 3DTV video formats where SEI messages are carried for every frame of video, processing recovery point SEI messages, and the use of pan-scan scan rectangle SEI messages in DVB, in addition to many other examples).

[0066] As described above, for each block, a set of motion information (also referred to herein as motion parameters) may be available. The set of motion information includes motion information for the forward and backward prediction directions. The forward and backward prediction directions are the two prediction directions in a bidirectional prediction mode, in which case the terms “forward” and “backward” do not necessarily have a geometric meaning. Instead, “forward” and “backward” correspond to the current picture’s reference picture list 0 (RefPicList0 or L0) and reference picture list 1 (RefPicList1 or L1). In some examples, when only one reference picture list is available for a picture or slice, only RefPicList0 is available, and the motion information for each block in the slice is always forward.

[0067] In some cases, motion vectors are used in the coding process (e.g., motion compensation) along with their reference index. Such motion vectors, with their associated reference index, are represented as a single predicted set of motion information. For each predicted direction, the motion information may include a reference index and a motion vector. In some cases, for simplicity, the motion vector itself may be referenced in such a way that it is assumed that the motion vector has an associated reference index. The reference index is used to identify a reference picture in the current reference picture list (RefPicList0 or RefPicList1). The motion vector has horizontal and vertical components that provide an offset from the coordinate position of the current picture to the coordinates in the reference picture identified by the reference index. For example, the reference index may indicate a particular reference picture to be used for a block in the current picture, and the motion vector may indicate where in the reference picture the best-matching block (the block that best matches the current block) is located.

[0068] In H.264 / AVC, each intermacroblock (MB) can be divided in four different ways, including one 16x16MB segment, two 16x8MB segments, two 8x16MB segments, and four 8x8MB segments. Different MB segments within a single MB may have different reference index values ​​(RefPicList0 or RefPicList1) for each direction. In some cases, when an MB is not divided into four 8x8MB segments, each MB segment may have only one motion vector in each direction. In some cases, when an MB is divided into four 8x8MB segments, each 8x8MB segment can be further divided into subblocks, in which case each subblock may have a different motion vector for each direction. In some examples, there are four different ways to obtain subblocks from an 8x8MB segment, including one 8x8 subblock, two 8x4 subblocks, two 4x8 subblocks, and four 4x4 subblocks. Each subblock may have a different motion vector for each direction. Therefore, motion vectors exist at a level higher than subblocks.

[0069] In AVC, a temporal direct mode may be enabled at either the MB level or the MB segment level for skipping and / or direct modes in B slices. For each MB segment, the motion vectors of the blocks at the same location as the current MB segment in the current block's RefPicList1[0] are used to derive the motion vector. Each motion vector within the same location block is scaled based on the POC distance.

[0070] Spatial direct modes can also be implemented in AVC. For example, in AVC, direct modes can also predict motion information from spatial neighbors.

[0071] As mentioned above, in HEVC, the largest coding unit in a slice is called a coding tree block (CTB). A CTB contains a quadtree, and the nodes of the quadtree are coding units. The size of a CTB can range from 16x16 to 64x64 in the HEVC main profile. In some cases, an 8x8 CTB size may be supported. Coding units (CUs) may be the same size as the CTB, or they may be as small as 8x8. In some cases, each coding unit is coded using one mode. When a CU is intercoded, it may be further divided into two or four predictive units (PUs), or there may be only one PU if no further division is applied. When there are two PUs in one CU, they may be a rectangle half the size, or two rectangles having 1 / 4 or 3 / 4 the size of the CU. When a CU is intercoded, there is one set of motion information for each PU. In addition, each PU is coded in a unique interpredictive mode to derive the set of motion information.

[0072] For example, in the case of motion prediction in HEVC, there are two interpretation modes, including a merge mode and an evolutionary motion vector prediction (AMVP) mode for the prediction unit (PU). Skipping is considered a special case of merging. In either AMVP mode or merge mode, a motion vector (MV) candidate list is maintained for multiple motion vector predictors. The motion vectors of the current PU, as well as the reference index in merge mode, are generated by taking one candidate from the MV candidate list. In some examples, one or more scaling window offsets may be included in the MV candidate list along with the stored motion vectors.

[0073] In examples where a list of motion compensators (MVs) is used for predicting block motion, the MVs may be constructed separately by the encoding device and the decoding device. For example, the MVs may be generated by the encoding device when encoding a block and by the decoding device when decoding a block. Information about motion compensator candidates in the MVs (e.g., information about one or more motion vectors, possibly information about one or more LIC flags that may be stored in the MVs) may be signaled between the encoding and decoding devices. For example, in merge mode, index values ​​for stored motion compensator candidates may be signaled from the encoding device to the decoding device (e.g., in the Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), slice header, video bitstream, or in syntax structures such as Supplementary Enhancement Information (SEI) messages transmitted separately from the video bitstream, and / or in other signaling). The decoding device can construct an MVs list and use the signaled references or indices to retrieve one or more motion compensator candidates from the constructed MVs list for use in motion compensation prediction. For example, the decoding device 112 may construct an MV candidate list and use motion vectors (and possibly LIC flags) from indexed locations for block motion prediction. In AMVP mode, in addition to references or indices, difference or residual values ​​may be signaled as deltas. For example, in AMVP mode, the decoding device may construct one or more MV candidate lists and apply delta values ​​to one or more motion information candidates obtained using signaled index values ​​when performing block motion compensation prediction.

[0074] In some examples, the MV candidate list contains up to five candidates for merge mode and two candidates for AMVP mode. In other examples, the MV candidate list may contain a different number of candidates for merge mode and / or AMVP mode. A merge candidate may contain a set of motion information. For example, a set of motion information may contain motion vectors corresponding to both the reference picture list (list 0 and list 1) and the reference index. When a merge candidate is identified by the merge index, the reference picture is used for the prediction of the current block, and the associated motion vector is determined. However, in AMVP mode, for each possible prediction direction from either list 0 or list 1, the reference index must be explicitly signaled along with the MVP index to the MV candidate list, since the AMVP candidate contains only motion vectors. In AMVP mode, the predicted motion vector may be further refined.

[0075] Thus, merge candidates correspond to a complete set of motion information, while AMVP candidates contain only one motion vector for a specific prediction direction and reference index. Candidates for both modes are similarly derived from the same spatial and temporal adjacency blocks.

[0076] In some cases, merge mode allows an inter-predicted PU to inherit one or more identical motion vectors, prediction directions, and one or more reference picture indices from an inter-predicted PU that includes motion data locations selected from one group of spatially adjacent motion data locations and two temporally identical motion data locations. In AMVP mode, one or more motion vectors of a PU can be predictively coded for one or more motion vector predictors (MVPs) from an AMVP candidate list constructed by the encoder and / or decoder. In some cases, for unidirectional inter-prediction of a PU, the encoder and / or decoder may generate a single AMVP candidate list. In some cases, for bidirectional prediction of a PU, the encoder and / or decoder may generate two AMVP candidates, one using motion data from spatially and temporally adjacent PUs from the forward prediction direction and the other using motion data from spatially and temporally adjacent PUs from the backward prediction direction.

[0077] Candidates for both modes are derived from spatial and / or temporal adjacent blocks. For example, Figures 2A and 2B include conceptual diagrams showing spatial adjacent candidates. Figure 2A shows spatial adjacent motion vector (MV) candidates for merge mode. Figure 2B shows spatial adjacent motion vector (MV) candidates for AMVP mode. Spatial MV candidates are derived from adjacent blocks for a given PU (PU0), but the method of generating candidates from blocks differs between merge mode and AMVP mode.

[0078] In merge mode, the encoder and / or decoder can form a merge candidate list by considering merge candidates from various motion data locations. For example, as shown in Figure 2A, up to four spatial MV candidates may be derived with respect to the spatially adjacent motion data locations shown in Figure 2A with numbers 0-4. The MV candidates may be ordered in the merge candidate list in the order indicated by numbers 0-4. For example, the locations and order may include the left location (0), the top location (1), the upper right location (2), the lower left location (3), and the upper left location (4).

[0079] In the AVMP mode shown in Figure 2B, adjacent blocks are divided into two groups: the left group containing blocks 0 and 1, and the upper group containing blocks 2, 3, and 4. For each group, possible candidates among adjacent blocks that reference the same reference picture as indicated by the signaled reference index have the highest priority when selected to form the final candidate for the group. It is possible that not all adjacent blocks contain motion vectors pointing to the same reference picture. Therefore, if no such candidate can be found, the difference in time distance can be compensated for by scaling the first available candidate to form the final candidate.

[0080] Figures 3A and 3B include conceptual diagrams illustrating time motion vector prediction. Time motion vector predictor (TMVP) candidates are added to the MV candidate list after spatial motion vector candidates if enabled and available. The motion vector derivation process for TMVP candidates is the same for both merge mode and AMVP mode. However, in some cases, the target reference index for TMVP candidates in merge mode may be set to 0 or derived from the target reference index of an adjacent block.

[0081] The primary block location for TMVP candidate derivation is the lower-right block outside the copositional PU, as shown as block "T" in Figure 3A, to compensate for biases to the upper and left blocks used to generate spatially adjacent candidates. However, if that block is currently located outside the CTB (or LCU) row or motion information is unavailable, the block is replaced with the central block of the PU. The motion vector for the TMVP candidate is derived from the copositional PU of the copositional picture, as shown at the slice level. Similar to the temporal direct mode in AVC, the motion vector of the TMVP candidate may undergo motion vector scaling, which is done to compensate for distance differences.

[0082] Other aspects of motion prediction are covered in the HEVC standard and / or other standards, formats, or codecs. For example, several other aspects of merge mode and AMVP mode are covered. Other aspects include motion vector scaling. With respect to motion vector scaling, it can be assumed that the value of the motion vector is proportional to the distance between pictures at presentation time. The motion vector associates two pictures, namely a reference picture and a picture containing the motion vector (i.e., a stored picture). When a motion vector is used to predict other motion vectors, the distance between the stored picture and the reference picture is calculated based on the picture order count (POC) value.

[0083] For a motion vector to be predicted, both its associated storage picture and reference picture may be different. Therefore, a new distance (based on POC) is calculated. The motion vector is then scaled based on these two POC distances. For spatially adjacent candidates, the storage pictures for two motion vectors may be the same, but the reference pictures may be different. In HEVC, motion vector scaling is applied to both TMVP and AMVP for both spatially adjacent candidates and temporally adjacent candidates.

[0084] Another aspect of motion prediction involves the generation of artificial motion vector candidates. For example, if the motion vector candidate list is incomplete, artificial motion vector candidates are generated and inserted at the end of the list until all candidates are obtained. In merge mode, there are two types of artificial MV candidates: composite candidates derived only for B slices, and zero candidates used only for AMVP when the first type does not provide enough artificial candidates. For each pair of candidates already in the candidate list that have the required motion information, a bidirectional composite motion vector candidate is derived by combining the motion vector of a first candidate that references a picture in list 0 and the motion vector of a second candidate that references a picture in list 1.

[0085] In some implementations, a pruning process may be performed when adding or inserting new candidates into the MV candidate list. For example, in some cases, MV candidates from different blocks may contain the same information. In such cases, storing duplicate motion information for multiple MV candidates in the MV candidate list can lead to redundancy and a decrease in the efficiency of the MV candidate list. In some examples, the pruning process can eliminate or minimize redundancy in the MV candidate list. For example, the pruning process may involve comparing MV candidates that may be added to the MV candidate list with MV candidates already stored in the MV candidate list. In an example for one explanation, the horizontal displacement (Δx) and vertical displacement (Δy) of a stored motion vector (indicating the position of the reference block relative to the current block's position) may be compared with the horizontal displacement (Δx) and vertical displacement (Δy) of a potential candidate's motion vector. If the comparison reveals that the potential candidate's motion vector does not match any of the one or more stored motion vectors, that potential candidate is not considered a candidate to be pruned and may be added to the MV candidate list. If a match is found based on this comparison, that potential MV candidate is not added to the MV candidate list, avoiding the insertion of identical candidates. In some cases, to reduce complexity, instead of comparing each potential MV candidate with all existing candidates, only a limited number of comparisons are performed during the pruning process.

[0086] Some coding schemes, such as HEVC, support weighted prediction (WP), in which case a scaling factor (denoted by a), a shift factor (denoted by s), and an offset (denoted by b) are used in motion compensation. If the pixel value at position (x,y) of the reference picture is p(x,y), then instead of p(x,y), p'(x,y) = ((a*p(x,y) + (1 << (s-1))) >> s) + b is used as the predicted value in motion compensation.

[0087] When WP is enabled, a flag is signaled for each reference picture in the slice to indicate whether WP is applied to that reference picture. If WP is applied to a single reference picture, a set of WP parameters (i.e., a, s, and b) is sent to the decoder and used for motion compensation from the reference picture. In some examples, WP flags and parameters are signaled separately for luma and chroma components to flexibly turn WP on and off for luma and chroma components. In WP, one identical set of WP parameters is used for all pixels in a single reference picture.

[0088] Figure 4A shows an example of reconstructed adjacent samples of block 402 and adjacent samples of reference block 404 used for unidirectional interpretation. A motion vector MV can be coded for block 402, and MV may include a reference index and / or other motion information to identify reference block 404 in the reference picture list. For example, MV may include horizontal and vertical components that provide an offset from the coordinate position in the current picture to the coordinate in the reference picture identified by the reference index. Figure 4B shows an example of reconstructed adjacent samples of block 422 and adjacent samples of a first reference block 424 and a second reference block 426 used for bidirectional interpretation. In this case, two motion vectors MV0 and MV1 may be coded for block 422 to identify the first reference block 424 and the second reference block 426, respectively.

[0089] Bilateral matching (BM) is a technique that can be used to refine a pair of initial motion vectors (e.g., a first motion vector MV0 and a second motion vector MV1). For example, BM can be performed by searching around the initial pair of motion vectors MV0 and MV1 to derive refined motion vectors (e.g., refined motion vectors MV0' and MV1'). The refined motion vectors MV0' and MV1' can then be used to replace the first motion vector MV0 and the second motion vector MV1, respectively. The refined motion vectors can be selected in the search as motion vectors identified in the search that minimize the block matching cost.

[0090] In some examples, a block matching cost may be generated based on the similarity of two motion-compensated predictors produced for two MVs. Exemplary criteria for block matching costs include, but are not limited to, the absolute difference sum (SAD), the absolute transform difference sum (SATD), and the sum of squared errors (SSE). The block matching cost criterion may also include a regularization term derived based on the MV difference between the current MV pair (e.g., the MV pair being considered for selection as refined motion vectors MV0' and MV1') and the initial MV pair (e.g., MV0 and MV1).

[0091] In some examples, one or more constraints may be applied to the MV difference terms MVD0 and MVD1 (e.g., MVD0 = MV0' - MV0 and MVD1 = MV1' - MV1). For example, in some cases, constraints may be applied based on the assumption that MVD0 and MVD1 are proportional to the time distance (TD) between the current picture (e.g., the current block) and the reference picture (e.g., the reference block) pointed to by the two MVs. In some examples, constraints may be applied based on the assumption that MVD0 = -MVD1 (e.g., MVD0 and MVD1 are equal in magnitude but opposite in sign).

[0092] In some examples, an inter-predicted CU may be associated with one or more motion parameters. For example, in the Versatile Video Coding Standard (VVC), each inter-predicted CU may be associated with one or more motion parameters, which may include, but are not limited to, motion vectors, reference picture indices, and reference picture list usage indices. The motion parameters may further include additional information related to the coding function of the VVC to be used for inter-predicted sample generation. Motion parameters may be signaled explicitly or implicitly. For example, when a CU is coded in skip mode, the CU is associated with one PU, has no significant residual coefficients, and does not have a coded motion vector delta or reference picture index.

[0093] In some embodiments, a merge mode may be defined such that motion parameters for the current CU are obtained from adjacent CUs, including spatial and temporal candidates. Additional or alternative merge modes may be defined based on additional schedules provided in the VVC standard. In some examples, a merge mode may be applicable to any inter-predicted CU (e.g., a merge mode may be applicable to modes other than skip modes). In some examples, an alternative to a merge mode may involve the explicit transmission of one or more motion parameters. For example, motion vectors, corresponding reference picture indices for each reference picture list, reference picture list usage flags, and other relevant information may be explicitly signaled for each CU.

[0094] In addition to the intercoding capabilities in HEVC, VVC includes several new, refined interpredictive coding techniques, including augmented merge prediction, merge mode with motion vector difference (MMVD), symmetric MVD (SMVD) signaling, affine motion-compensated prediction, sub-block based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage: 1 / 16 luma-sample MV storage and 8x8 motion field compression, bi-prediction with CU-level weight (BCW), bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), geometric partitioning mode (GPM), and combined inter and intra prediction (CIIP).

[0095] For extended merge prediction in VVC merge mode (for example, referred to as normal or default merge mode), the merge candidate list may be constructed by sequentially including five types of candidates: spatial motion vector predictions (MVPs) from spatially adjacent CUs, temporal MVPs from co-location CUs, history-based MVPs from first-in-first-out (FIFO) tables, pairwise average MVPs, and zero MVPs. The size of the merge candidate list may be signaled in the sequence parameter set header. The maximum allowable size of the merge candidate list may be 6 (e.g., 6 entries or 6 candidates). For each CU coded in merge mode, the index of the best merge candidate is encoded using truncated unary (TU). In some examples, VVC can also support parallel derivation of merge candidate lists for all CUs of a certain size or area (e.g., as done in HEVC). The five aforementioned types of merge candidates, as well as the relevant exemplary derivation processes for each category of merge candidates, are then described below.

[0096] Figure 5 shows the locations or positions of spatial merge candidates for use when processing a block, according to some examples of the present disclosure. For example, Figure 5 shows exemplary locations of spatial merge candidates (also called “spatial neighbors”) A0, A1, B0, B1, and B2 for use in processing block 500 according to some examples of the present disclosure. Spatial neighbors A0, A1, B0, B1, and B2 are shown in Figure 5 in relation to block 500. The derivation of spatial merge candidates in VVC may be the same as in HEVC, except that the positions of the first two merge candidates are swapped. In some examples, the largest of the four merge candidates may be selected from five spatial merge candidates (e.g., A0, A1, B0, B1, and B2) located at the positions shown in Figure 5.

[0097] The order of derivation may be B0, A0, B1, A1, and B2. For example, merge candidate position B2 may only be considered when one or more CUs associated with positions B0, A0, B1, or A1 are unavailable or are intra-coded. CUs associated with positions B0, A0, B1, or A1 may be unavailable because the CUs belong to different slices or tiles. In some embodiments, after a merge candidate at position A1 has been added, the addition of remaining merge candidates may undergo a redundancy check. A redundancy check may be performed so that merge candidates with the same motion information are excluded from the merge candidate list (for example, to improve coding efficiency).

[0098] Figure 6 shows some examples of motion vector scaling 600 for time merge candidates for use when processing blocks. In some examples, time merge candidates can be derived and one merge candidate (e.g., time merge candidate) is added to the merge candidate list. Time merge candidates can be derived based on scaled motion vectors. Scaled motion vectors can be derived based on colocation CUs contained in colocation reference pictures. A list of reference pictures to be used for deriving colocation CUs can be explicitly signaled in the slice header.

[0099] For example, Figure 6 shows current picture 610 and co-position picture 630, which may be associated with current reference picture 615 and co-position reference picture 635, respectively. Figure 6 also shows current CU 612 (e.g., associated with current picture 610) and co-position CU (e.g., associated with co-position picture 630). In some examples, scaled motion vectors for deriving time merge candidates may be derived or obtained as shown in Figure 6. For example, Figure 6 shows dotted line 611, which is scaled from the motion vector of co-position CU 632 using Picture Order Count (POC) distances tb and td. In some examples, tb is the POC difference between current reference picture 615 and current picture 610, and td is the POC difference between co-position reference picture 635 and co-position picture 630. The reference picture index of the time merge candidate may be set to equal 0.

[0100] Figure 7 shows some examples of time merge candidates 700 for use when processing blocks in this disclosure. In some examples, after a single time merge candidate is selected as discussed above in relation to Figure 6, the position of the time merge candidate may be selected from candidate positions C0 and C1, as shown in Figure 7. In some examples, candidate position C1 may be used if the CU at position C0 is not available, is intracoded, or is outside the current row of the CTU. Otherwise, position C0 is used in the derivation of the time merge candidate.

[0101] In some embodiments, history-based motion vector prediction (HMVP) merge candidates may be added to the merge candidate list after spatial MVP merge candidates (e.g., as described above with respect to Figure 5) and TMVP merge candidates (e.g., as described above with respect to Figures 6 and 7). HMVP merge candidates may be derived based on motion information of previously coded blocks. For example, motion information of previously coded blocks may be stored (e.g., in a table) and used as a motion vector prediction (MVP) for the current CU. In some examples, a table with multiple HMVP candidates may be maintained between the coding and / or decoding processes. Whenever a new CTU row is encountered, the table is reset (e.g., emptied). Whenever there is a non-subblock intercoded CU, the relevant motion information is added to the last entry in the table as a new HMVP candidate.

[0102] In some examples, the HMVP table size S can be set to a value of 6 (for example, up to 6 history-based MVP (HMVP) candidates can be added to the HMVP table). When inserting a new HMVP candidate into the HMVP table, a constrained first-in-first-out (FIFO) rule may be used. The constrained FIFO rule may include redundancy checks applied to determine whether the same HMVP already exists in the table (for example, to determine if the newly inserted HMVP candidate is the same as an existing HMVP candidate in the table). If the redundancy check for the newly inserted HMVP candidate finds that the same HMVP already exists in the table, the same HMVP can be removed from the table, and all subsequent HMVP candidates are moved forward.

[0103] In some embodiments, HMVP candidates (e.g., included in an HMVP list or HMVP table) may subsequently be used to construct or otherwise generate a merge candidate list. For example, to generate a merge candidate list, several recent HMVP candidates in the table may be reviewed in order and inserted into the merge candidate list after TMVP candidates. Redundancy checks may be applied to HMVP candidates added to the merge candidate list, and these redundancy checks are used to determine whether an HMVP candidate is the same as or identical to any spatial or temporal merge candidate that has been previously added to or is already included in the merge candidate list.

[0104] In some cases, the number of redundancy checks performed in connection with generating the merge candidate list and / or HMVP table can be reduced. For example, the number of HMVP candidates used for merge list generation can be set as (N <= 4) ? M : (8 - N), where N is the number of existing candidates in the merge candidate list and M is the number of available HMVP candidates in the HMVP table. In other words, if the condition N <= 4 evaluates to true (e.g., the merge candidate list contains 4 or fewer candidates), all M HMVP candidates in the HMVP table are used for merge list generation. If the condition N <= 4 evaluates to false (e.g., the merge candidate list contains more than 4 candidates), then 8-N HMVP candidates in the HMVP table are used for merge list generation. In some cases, the process of building the merge candidate list from HMVP can be terminated when the total number of available merge candidates reaches the maximum allowed number of merge candidates minus 1.

[0105] The average merge candidate per pair can be derived based on predetermined pairs of merge candidates in an existing list of merge candidates. For example, the average merge candidate per pair can be generated by averaging predetermined pairs of merge candidates in an existing list of merge candidates, given as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where these numbers represent the merge index for the list of merge candidates. In some cases, the averaged motion vector can be calculated separately for each reference list. If both motion vectors (for example, of a predetermined pair) are in the same list, these two motion vectors can be averaged even if the two motion vectors point to different reference pictures. In some cases, if only one motion vector (for example, of a predetermined pair) is available, that single available motion vector can be used as the averaged motion vector. If no motion vectors (for example, of a predetermined pair) are available, the list can be identified as invalid. In some examples, if the merge candidate list is not filled after the average merge candidates per pair have been added, as described above, one or more zero MVPs may be inserted at the end of the merge candidate list until the maximum number of merge candidates is reached.

[0106] In some embodiments, bi-prediction with CU-level weights (BCW) can be used. For example, a bi-prediction signal can be generated by averaging two prediction signals obtained from two different reference pictures. Additionally or alternatively, a bi-prediction signal can be generated using two different motion vectors. In some examples, a bi-prediction signal can be generated using the HEVC standard by averaging two prediction signals obtained from two different reference signals and / or by using two different motion vectors.

[0107] In other embodiments, the dual prediction mode can be extended from a simple average to include a weighted average of two prediction signals. For example, the dual prediction mode may include a weighted average of two prediction signals using the VVC standard. In some examples, the dual prediction mode may include a weighted average of two prediction signals as follows: P bi-pred = ((8-w)*P0+w*P1+4)≫3 Equation (1)

[0108] The weighted average biprediction given in equation (1) may include five weights w ∈ {-2, 3, 4, 5, 10}. For each bipredicted CU, the weights w may be determined according to one or more of the following: In one example, for non-merged CUs, the weight index may be signaled after motion vector difference (MVD). In another example, for merged CUs, the weight index may be inferred from adjacent blocks based on the merge candidate index. In some cases, BCW may be applied only to CUs with 256 or more luma samples (e.g., CUs where the product of the width and height of the CU is 256 or greater). In some cases, all five weights w may be used for low-latency pictures. For non-low-latency pictures, only three of the five weights w may be used (e.g., three weights w ∈ {3, 4, 5}).

[0109] On the encoder side, fast search algorithms can be applied to find the weight index without significantly increasing the complexity of the encoder, as will be explained in more detail below. In some examples, fast search algorithms can be applied to find the weight index, as will be explained in part in JVET-L0646. For example, when combined with adaptive motion vector resolution (AMVR), if the current picture is a low-latency picture, unequal weights are only conditionally confirmed for 1-Pel and 4-Pel motion vector precision. When combined with affine modes, affine motion estimation (ME) may be performed for unequal weights only if the affine mode is selected as the best mode at the moment. When the two reference pictures in the biprediction mode are the same, unequal weights may only be conditionally confirmed. In some cases, unequal weights are not searched when certain conditions are met (e.g., depending on the picture order count (POC) distance between the current picture and its reference picture, the coding quantization parameter (QP), and / or the time level, etc.).

[0110] In some examples, the BCW weight index may be coded using one context-coded bin followed by one or more bypass-coded bins. For example, the first context-coded bin may be used to indicate whether equal weights are used. If unequal weights are used, an additional bin may be signaled using bypass coding to indicate which unequal weights are used.

[0111] Weighted Prediction (WP) is a coding tool supported by the H.264 / AVC and HEVC standards for efficiently coding video content using fading. Support for WP has also been added to the VVC standard. WP can be used to enable the signaling of one or more weighting parameters (e.g., weight and offset) for each reference picture in each of the reference picture lists L0 and L1. Subsequently, the weight and offset of the corresponding reference picture are applied during motion compensation.

[0112] WP and BCW can be used with different types of video content. In some examples, when CU uses WP, the BCW weight index may not need to be signaled, and w may be inferred to have a value of 4 (e.g., equal weights are applied). For example, when CU uses WP, the BCW weight index may not need to be signaled to avoid interaction between WP and BCW (e.g., this can complicate the design of the VVC decoder).

[0113] For merged CUs, the weight index can be inferred from adjacent blocks based on the merge candidate index in both normal merge mode and inherited affine merge mode. In constructed affine merge mode, the affine motion information can be constructed based on the motion information of up to three blocks. For example, the BCW index for a CU using constructed affine merge mode can be set to equal the BCW index of the first control point MV. In some examples, using the VVC standard, composite inter and intra prediction (CIIP) and biprediction with CU-level weights (BCW) cannot be applied together to a CU. When a CU is coded using CIIP mode, the BCW index of the current CU can be set to a value of 2 (for example, set to equal weights).

[0114] Figure 8 is a figure illustrating some examples of bilateral matching in the present disclosure. As previously mentioned, bilateral matching (BM) may be used to refine a pair of two initial motion vectors MV0 and MV1. For example, BM may be performed by exploring around MV0 and MV1 to derive refined motion vectors MV0' and MV1', respectively, that minimize the block matching cost. The block matching cost may be calculated based on the similarity of two motion-compensated predictors generated using the pair of initial motion vectors (e.g., MV0 and MV1). For example, the block matching cost may be based on the absolute difference sum (SAD). Additionally or alternatively, the block matching cost may be based on, or include, a regularization term based on the motion vector difference (MVD) between the current MV pair (e.g., MV0' and MV1' currently being tested) and the initial MV pair (e.g., MV0 and MV1). As will be explained in more detail below, one or more constraints may be applied based on MVD0 (e.g., MVD0 = MV0' - MV0) and MVD1 (e.g., MVD1 = MV1' - MV1).

[0115] As previously mentioned, in the Multipurpose Video Coding Standard (VVC), bilateral matching-based decoder-side motion vector refinement (DMVR) may be applied to improve (e.g., refine) the accuracy of the MV of a bilateral predictive merge candidate. For example, as shown in the example in Figure 8, bilateral matching-based DMVR may be applied to improve, or otherwise refine, the accuracy of the MV of a bilateral predictive merge candidate 812. The bilateral predictive merge candidate 812 may now be contained in picture 810 and may be associated with an initial pair of motion vectors MV0 and MV1. Before performing the bilateral matching-based DMVR, the initial motion vectors MV0 and MV1 may be acquired, identified, or otherwise determined for the bilateral predictive merge candidate 812. Subsequently, as will be described in more detail below, the initial motion vectors MV0 and MV1 may be used to identify the refined motion vectors MV0' and MV1' for the bilateral predictive merge candidate 812.

[0116] As shown in the example in Figure 8, the first initial motion vector MV0 may point to the first reference picture 830. The first reference picture 830 may be associated with the backward direction (for example, with respect to the current picture 810) and / or may be included in the reference picture list L0. The second initial motion vector MV1 may point to the second reference picture 820. The second reference picture 820 may be associated with the forward direction (for example, with respect to the current picture 810) and / or may be included in the reference picture list L1.

[0117] A first initial motion vector MV0 may be used to determine or generate a first predictor 832, which may be a block contained in the first reference picture 830. The first predictor 832 may also be called a first candidate block (for example, the first predictor 832 is a candidate block in the first reference picture 830 and / or in the reference picture list L0). A second initial motion vector MV1 may be used to determine or generate a second predictor 822, which may be a block contained in the second reference picture 820. The second predictor 822 may also be called a second candidate block (for example, the second predictor 822 is a candidate block in the second reference picture 820 and / or in the reference picture list L1).

[0118] Next, a search may be performed around the first predictor 832 and the second predictor 822 to identify or determine the first refined motion vector MV0' and the second refined motion vector MV1', respectively. As shown, the surrounding areas associated with each of the first predictors 832 (e.g., areas in the first reference picture 830 and / or reference picture list L0) and the surrounding areas associated with the second predictor 822 (e.g., areas in reference picture 820 and / or reference picture list L1) may be searched. For example, the surrounding areas associated with the first predictor 832 may be searched to identify or examine one or more refined candidate blocks 834, and the surrounding areas associated with the second predictor 822 may be searched to identify or examine one or more refined candidate blocks 824. The search may be performed based on one or more distortions (e.g., SAD, SATD, SSE, etc.) and / or regularization terms calculated between one of the initial predictors 832 or 822 and the corresponding refined candidate block 834 or 824. In some examples, the distortions and / or regularizations may be calculated based on the distance traveled from the initial point (e.g., the distance between the initial point associated with the initial predictor 832 or 822 and the searched point associated with the refined candidate block 834 or 824, respectively).

[0119] As the search moves around the initial points associated with the initial predictors 832 and 822 respectively, new refined candidate blocks (e.g., refined candidate blocks 834 and 824) are obtained. Each new refined candidate block may be associated with a new cost (e.g., a calculated SAD value, one of the motion vector differences MVD0 or MVD1, etc.). The search may be associated with a search range and / or search interval. After searching each candidate block contained within the search range and / or search window for the initial predictors 832 and 822 (e.g., and determining the corresponding cost for each explored candidate block), the candidate block with the lowest determined cost is identified and may be used to generate refined motion vectors MV0' and MV1'.

[0120] In some examples, bilateral matching (BM) based DMVR can be performed by calculating the SAD between two candidate blocks in the reference picture lists L0 and L1. As shown in Figure 8, the SAD between blocks can be calculated based on each MV' candidate (e.g., blocks 834 and 824) around the initial MV (e.g., around predictors 832 and 822, respectively). The MV' candidate with the smallest SAD can be selected as the refined MV and used to generate the bipredicted signal. In some examples, the SAD of the initial MV becomes a regularization term by subtracting 1 / 4 of the SAD value. In some cases, the time distances from the two reference pictures to the current picture (e.g., picture order count (POC) difference) may be the same, and MVD0 and MVD1 may be the same magnitude but opposite signs (e.g., MVD0 = -MVD1).

[0121] In some cases, bilateral matching-based DMVR can be performed using a refined search range of two integer luma samples from the initial MV. For example, in the context of Figure 8, bilateral matching-based DMVR can be performed using a refined search range of two integer luma samples from the initial motion vectors MV0 and MV1 (e.g., from the initial predictors 832 and 822, respectively). This search may include an integer sample offset search stage and a fractional sample refinement stage.

[0122] In some cases, a 25-point full search may be applied for integer sample offset searches. A 25-point full refinement search can be performed by first calculating the SAD of an initial MV pair (e.g., the initial MV pair MV0 and MV1 and / or initial predictors 832 and 822). If the SAD of the initial MV pair is less than the threshold, the integer sample stage of the DMVR may be terminated. If the SAD of the initial MV pair is not less than the threshold, the SADs of the remaining 24 points may be calculated and checked in raster scan order. The point with the smallest SAD may then be selected as the output of the integer sample offset search stage.

[0123] As mentioned above, fractional sample refinement may follow integer sample offset search. In some cases, computational complexity can be reduced by deriving fractional sample refinement using one or more parametric error surface equations (for example, instead of performing additional searches using SAD comparisons). Fractional sample refinement may be conditionally invoked based on the output of the integer sample offset search stage. For example, if the integer sample offset search stage terminates with the center having the smallest SAD in either the first or second iteration, non-integer sample refinement may be further applied in response.

[0124] As mentioned above, the parametric error surface can be used to derive fractional sample refinement. For example, in parametric error surface-based sub-pixel offset estimation, the cost at the center position and the costs at four adjacent positions (e.g., relative to the center) can be used to fit a 2D parabolic error surface equation of the following form. E(x,y) = A(x - x min ) 2 + B(y - y min ) 2 + C Equation (2)

[0125] Here, (x min ,y min ) corresponds to the fractional position where the cost is minimum, and C corresponds to the minimum cost value. By solving Equation (2) using the cost values of five search points (e.g., the center position and four adjacent positions), (x min , y min ) can be calculated as follows. x min = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0))) Equation (3) y min = (E(0,-1) - E(0,1)) / (2((E(0,-1) + E(0,1) - 2E(0,0))) Equation (4)

[0126] x min and y min values may be automatically constrained to be between -8 and 8 (e.g., all cost values are positive and the minimum value is E(0,0)), and these may correspond to a half-pel offset at 1 / 16 pel MV accuracy in VVC. The calculated fractional (x min ,y min ) values can be added to the integer distance refinement MV (e.g., from the integer sample offset search described above) to obtain or determine the sub-pixel accuracy refinement delta MV.

[0127] In VVC, the motion vector resolution can be 1 / 16 of a sample. In some examples, samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the refined search points may enclose an initial fractional PelMV with integer sample offsets, and the DMVR search process may include interpolating samples at fractional positions. In some examples, a bilinear interpolation filter may be used to generate fractional samples for the DMVR search process. Using a bilinear interpolation filter to generate fractional samples for the DMVR search process can reduce computational complexity. In some cases, when a bilinear interpolation filter is used with a 2-sample search range, the DMVR search process may not need to access additional reference samples (compared to, for example, existing motion compensation processes).

[0128] After the refined MV is determined using the DMVR search process, an 8-tap interpolation filter may be applied to generate the final prediction. In some examples, additional reference samples may not be required to perform the interpolation process based on the original MV (e.g., as described above). In some examples, additional samples may be utilized to perform interpolation based on the refined MV. Rather than accessing additional reference samples (e.g., accessing additional reference samples compared to the existing motion compensation process), existing or available reference samples may be padded to generate additional reference samples for performing interpolation based on the refined MV. In some cases, when the width and / or height of the CU is greater than 16 luma samples, the DMVR process may further involve dividing the CU into subblocks, each having a width and / or height equal to 16 luma samples, as will be described in more detail below.

[0129] In VVC, DMVR may apply to CUs coded using one or more of the following modes or features. For the sake of explanation, the modes or features associated with a CU to which DMVR may apply may also be called “DMVR conditions.” In some examples, DMVR conditions may include, but are not limited to, CU-level merge modes with dual prediction MV, one reference picture in the past (e.g., with respect to the current picture) and another reference picture in the future (e.g., with respect to the current picture), the distances from the two reference pictures to the current picture (e.g., POC distance) are the same, both reference pictures are short-term reference pictures, a CU containing more than 64 luma samples, both the height and width of the CU being greater than or equal to 8 luma samples, BCW weight indices showing equal weights, weighted prediction (WP) not being enabled for the current block, and synthetic inter- and intra-prediction (CIIP) modes not being used for the current block.

[0130] Figure 9 shows an exemplary extended CU region 900 that may be used to perform bidirectional optical flow (BDOF) in some examples of the present disclosure. For example, BDOF may be used to refine the biprediction signals of luma samples in a CU at the 4×4 subblock level (e.g., 4×4 subblock 910). In some cases, BDOF may be used for subblocks of other sizes (e.g., 8×8 subblocks, 4×8 subblocks, 8×4 subblocks, 16×16 subblocks, and / or other sizes). The BDOF mode may be based on optical flow, assuming that the motion of the object is smooth. As shown in Figure 9, the BDOF mode may utilize one extended row and column (e.g., extended row 970 and extended column 980) around the boundary of the extended CU region 900.

[0131] For each 4x4 subblock (for example, 4x4 subblock 910), motion refinement is performed based on minimizing the difference between the L0 predicted sample and the L1 predicted sample. x ,v y ) can be calculated. Motion refinement (v x ,v y ) can then be used to adjust the sample values ​​that are predicted to be bipredicted within a 4x4 subblock (for example, a 4x4 subblock 910). In an example for one explanation, BDOF can be performed as described below.

[0132] First, the horizontal and vertical gradients of the two prediction signals.

[0133]

number

[0134] and

[0135]

number

[0136] (k=0,1) can be calculated by directly calculating the difference between two adjacent samples, respectively.

[0137]

number

[0138] Here I (k) (i,j) represents the sample value at coordinate (i,j) of the predicted signal in list k, where k=0,1. Since shift1 can be set to equal to 6, shift1 can be calculated based on the luma bit depth (e.g., bitDepth).

[0139] Next, the autocorrelation and cross-correlation S1, S2, S3, S5, and S6 of the gradient can be calculated as follows. S1 = Σ (i,j)∈Ω│ψ x (i,j)│ Equation (7) S2 = Σ (i,j)∈Ω ψ x (i,j)·sign(ψ y (i,j)) Equation (8) S3 = Σ (i,j)∈Ω θ(i,j)·(-sign(ψ x (i,j))) Equation (9) S5=Σ (i,j)∈Ω │ψ y (i,j)│ Equation (10) S6=Σ (i,j)∈Ω θ(i,j)·(-sign(ψ y (i,j))) Equation (11)

[0140] Here,

[0141]

number

[0142] θ(i,j)=(I (0) (i,j)≫shift2)-(I (1) (i,j)≫shift2) This is equation (14).

[0143] Here, Ω could be a 6x6 window (or a window of other size) around a 4x4 subblock 910 (or a subblock of other size). The value of shift2 may be set to equal 4 (or another suitable value), and the value of shift3 may be set to equal 1 (or another suitable value).

[0144] Motion refinement (v x ,v y ) can then be derived using the terms of cross-correlation and autocorrelation, as follows:

[0145]

number

[0146] Here, th' BIO= 1 << 4,

[0147]

number

[0148] This is the floor function.

[0149] Based on motion refinement and gradient, the following adjustments can be calculated for each sample within a 4x4 subblock 910 (or other subblock sizes).

[0150]

number

[0151] Finally, by adjusting the biprediction samples as follows, the BDOF samples for the extended CU region 900 shown in Figure 9 can be calculated. Nod BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset )≫shift5 Formula (18)

[0152] Here, shift5 may be set to be equal to Max(3,15 - BitDepth), and the variable ο offset It may be set to be equal to (1 << (shift5 - 1)).

[0153] In some cases, the above values ​​may be chosen such that the multiplier in the BDOF process (for example, as described above) does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process (for example, as described above) is kept within 32 bits.

[0154] In some examples, the gradient value is one or more predicted samples I in a list k where k=0,1. (k)The BDOF may be derived at least in part on the generation of (i,j), and one or more predicted samples are now outside the boundary of the CU. As shown in Figure 9 and described above, the BDOF can utilize one extended row and one extended column (e.g., extended row 970 and extended column 980) around the boundary of the extended CU region 900. In some cases, the computational complexity of generating out-of-bounds predicted samples can be controlled at least in part on the generation of predicted samples within the extended area (e.g., unshaded blocks contained in extended row 970 and extended column 980 along the perimeter of the extended CU region 900) by directly using reference samples at nearby integer positions without interpolation. For example, the floor() operation may be used on the coordinates of the reference sample at a nearby reference sample. Next, an 8-tap motion-compensated interpolation filter may be used to generate prediction samples within the shaded box of the extended CU region 900 (for example, the box inside the unshaded outer perimeter of the extended area prediction samples for the extended row 970 and extended column 980). In some examples, the extended sample values ​​may be used only in gradient calculations. The remaining steps of the BDOF process may be performed, if necessary, by padding (e.g., iterating) any sample and gradient values ​​outside the boundary of the extended CU region 900 based on their nearest neighbors.

[0155] In some cases, BDOF can be used to refine the biprediction signal of a CU at a 4x4 subblock level (or other subblock level), as mentioned above. In an example for one explanation, BDOF can be applied to a CU if the CU satisfies some or all of the following conditions: the CU is coded using a "true" biprediction mode (e.g., one of the two reference pictures is in display order before the current picture and the other reference picture is in display order after the current picture), the CU is not coded using affine mode or SbTMVP merge mode, the CU has more than 64 luma samples, both the height and width of the CU are greater than or equal to 8 luma samples, the BCW weight indices show equal weights, WP is not currently valid for the CU, and / or CIIP mode is not currently used for the CU.

[0156] In some embodiments, multipath decoder-side motion refinement can be used. For example, at the JVET-V meeting, the Enhanced Compression Model (ECM) was established to study compression techniques beyond VVC (https: / / vcgit.hhi.fraunhofer.de / ecm / VVCSoftware_VTM / - / tree / ECM). In the ECM, "multipath decoder-side motion refinement" [JVET U0100] is employed to replace DMVR in VVC. Multipath decoder-side motion refinement may involve multiple passes. In some examples, multipath decoder-side motion refinement can be performed by applying bilateral matching (BM) multiple times. For example, in each pass (e.g., of multipath decoder-side motion refinement), BM may be applied to a different block size.

[0157] In the first pass, bilateral matching (BM) may be applied to an entire coding block or coding unit (for example, as with those of DMVR in VVC and / or as described above). For example, BM in the first pass may be applied to a coding block or CU using sizes such as 64×64, 64×32, 32×32, 32×16, 16×16, and combinations thereof.

[0158] In the second pass, the BM can be applied again to each 16x16 subblock contained within (or potentially generated from) the entire coding block where the BM was executed in the first pass. For example, if the BM in the first pass was applied to a 64x64 block, the 64x64 block may be divided into a 4x4 grid of subblocks, each of which is 16x16. In the second pass, in the above example, the BM can be applied to each of the 16 subblocks. In another example, if the BM in the first pass was applied to a 32x32 block, the 32x32 block may be divided into a 2x2 grid of subblocks, each of which is 16x16. In this example, the second pass can be performed by applying the BM to each of the four 16x16 subblocks of the 32x32 block. The refined MV generated or obtained from the BM in the first pass can be used as the initial MV for each 16x16 subblock on which the BM in the second pass is executed.

[0159] In the third pass, multiple 8x8 subblocks can be obtained or generated either based on the original block (e.g., from the first BM pass) and / or based on a 16x16 subblock (e.g., from the second BM pass). In some examples, each 16x16 subblock from the second BM pass may be divided into a 2x2 grid of 8x8 subblocks. In the third BM pass, one or more MVs associated with each 8x8 subblock may be further refined by applying a bidirectional optical flow (BDOF).

[0160] In one example for explanation, multipath decoder-side motion refinement may be performed for a 64x64 block or CU. In the first pass, a BM may be applied to the 64x64 block to generate or obtain a pair of refined MVs. In the second pass, the 64x64 block can be divided into 16 subblocks, each subblock being 16x16 in size. In the second pass, the BM can be applied again to each of the 16 subblocks, this time using the refined MV from the first pass as the initial MV. In the third pass, each 16x16 subblock can be divided into four subblocks of size 8x8 (for example, a total of 16*4=64 8x8 subblocks for the original 64x64 block). In the third pass, as previously described, one or more MVs associated with each 8x8 subblock can be refined by applying a BDOF.

[0161] Examples of first, second, and third passes that may be included or used to perform multipath decoder-side motion refinement are described below.

[0162] In some examples, the first pass may involve performing block-based bilateral matching (BM) motion vector (MV) refinement. For example, in the first pass, the refined MV is derived by applying the BM to the coding blocks. Similar to decoder-side motion vector refinement (DMVR), in bi-predictive operations associated with BM, the refined MV is searched around two initial MVs (e.g., MV0 and MV1) in reference picture lists L0 and L1. pass1 and MV1 pass1 ) is derived around an initial pair of MVs (e.g., MV0 and MV1, respectively) based on the lowest bilateral matching cost between two reference blocks in L0 and L1.

[0163] The first pass of the BM may involve performing a local search to derive an integer sample-precision intDeltaMV. The local search may be performed by repeatedly applying a 3x3 square search pattern (or other search pattern) horizontally to the search range [-sHor, sHor] and vertically to the search range [-sVer, sVer]. The values ​​of sHor and sVer may be determined by the dimensions of the block. In some cases, the maximum value of sHor and / or sVer may be 8 (or other appropriate value).

[0164] The bilateral matching cost can be calculated as follows: bilCost = mvDistanceCost + sadCost Formula (19)

[0165] When the block size cbW*cbH is greater than 64 (or other block size thresholds), the mean-removed sum of absolute difference (MRSAD) cost function may be applied to remove the DC effect of distortion between reference blocks. The intDeltaMV local search terminates when bilCost at the center point of the 3x3 search pattern (or other search pattern) has the lowest cost. Otherwise, the current minimum cost search point becomes the new center point of the 3x3 search pattern (or other search pattern), and the local search continues, searching for the minimum cost until the end of the search range is reached (e.g., [-sHor, sHor] horizontally and [-sVer, sVer] vertically).

[0166] In some cases, the existing fractional sample refinement may be further applied to derive the final deltaMV. Then, the refined MV after the first pass is MV0 pass1 = MV0 + deltaMV Equation (20) MV1 pass1 = MV1 - deltaMV This can be derived as equation (21).

[0167] In some examples, the second pass may involve performing subblock-based bilateral matching (BM) motion vector (MV) refinement. For example, in the second pass, the refined MV may be derived by applying BM to a 16x16 (or other size) grid subblock. For each subblock, the refined MV is derived from the two MVs obtained in the first pass (e.g., MV0) in the reference picture lists L0 and L1. pass1 and MV1 pass1 It will be explored around ).

[0168] Based on the exploration, there are two refined MVs, MV0 pass2 (sbIdx2) and MV1 pass2 (sbIdx2) can be derived based on the lowest bilateral matching cost between two reference subblocks in L0 and L1, where sbIdx2=0,...,N-1 is the index of the subblock (for example, the BM of the second pass may be applied to each subblock generated from the original block used in the BM of the first pass). For example, as previously described, a total of 16 subblocks, each having dimensions of 16x16, may be generated or obtained for an input block of size 64x64. In this example, sbIdx2 could be the index for one of the individual 16 subblocks.

[0169] For each subblock, the second pass of the BM may include performing a complete search to derive an integer sample-precision intDeltaMV. The complete search may have a horizontal search range [-sHor, sHor] and a vertical search range [-sVer, sVer]. The values ​​of sHor and sVer may be determined by the dimensions of the block, and the maximum values ​​of sHor and sVer may be 8 (or other appropriate values).

[0170] Figure 10 shows exemplary search area regions within a coding unit (CU) 1000, according to some examples of the present disclosure. For example, Figure 10 shows four different search regions within the coding unit 1000 (e.g., a first search region 1020, a second search region 1030, a third search region 1040, a fourth search region 1050, etc.). In some cases, the bilateral matching cost may be calculated by applying a cost coefficient to the SATD cost (or other cost function) between two reference blocks, as follows: bilCost = satdCost * costFactor Formula (22)

[0171] As shown in Figure 10, the search area (2*sHor+1)*(2*sVer+1) can be divided into five diamond-shaped search regions. In other embodiments, other search regions may be used. Each search region is assigned a costFactor value, which is determined by the distance (intDeltaMV) between each search point and the starting MV. Each diamond-shaped search region (e.g., search regions 1020, 1030, 1040, and 1050) can be processed in an order starting from the center of the search region. Within each search region, search points are processed in a raster scan order starting from the top left of the region and moving toward the bottom right corner.

[0172] In some examples, the first search region 1020 is explored first, followed by the second search region 1030, then the third search region 1040, and then the fourth search region 1050. The integer Pell complete search terminates when the best bilCost in the current search region is less than the threshold equal to sbW*sbH. Otherwise, the integer Pell complete search continues to the next search region until all search points have been examined.

[0173] In some cases, further VVC DMVR fractional sample refinement may be applied to derive the final deltaMV(sbIdx2). The refined MV in the second pass can then be derived as follows: MV0 pass2 (sbIdx2) = MV0 pass1+ deltaMV(sbIdx2) Equation (23) MV1 pass2 (sbIdx2) = MV1 pass1 - deltaMV(sbIdx2) Equation (24)

[0174] In some examples, the third pass may involve performing subblock-based bidirectional optical flow (BDOF) motion vector (MV) refinement. For example, in the third pass, the refined MV may be derived by applying the BDOF to an 8x8 (or other size) grid subblock. For each 8x8 subblock, the BDOF refinement may be applied to derive scaled Vx and Vy without clipping, starting from the refined MV of the parent subblock in the second pass. For example, each parent subblock in the second pass may be 16x16 in size (for example, each parent subblock in the second pass may be associated with four 8x8 subblocks used in the third pass). The BDOF refinement in the third pass is the refined MV, MV0 pass2 (sbIdx2) and MV1 pass2 It may be applied or executed based on (sbIdx2), where sbIdx2 is an index of one of the parent subblocks of the second path.

[0175] Next, the derived bioMv(Vx, Vy) can be rounded to 1 / 16 sample precision and clipped between -32 and 32 (or other sample precisions and / or other clipping values ​​or ranges).

[0176] The refined MV(MV0) in the third pass pass3 (sbIdx3) and MV1 pass3 (sbIdx3)) can be derived as follows: MV0 pass3 (sbIdx3) = MV0 pass2 (sbIdx2) + bioMv formula (25) MV1 pass3 (sbIdx3) = MV0 pass2(sbIdx2) - bioMv formula (26)

[0177] Here, sbIdx3 is an index for a specific 8x8 subblock used in the third pass, and sbIdx2 is an index for a specific 16x16 subblock used in the second pass. In some examples, sbIdx3 may have a range such that each 8x8 subblock is uniquely identifiable by its sbIdx3 index. For example, a 64x64 block can be divided into 64 8x8 subblocks, and sbIdx3 can contain 64 unique index values ​​for different 8x8 subblocks. In some examples, sbIdx3, in combination with the sbIdx2 value of the corresponding 16x16 parent block, may have a range such that each 8x8 subblock is uniquely identifiable by its sbIdx3. For example, a 64x64 block can be divided into a total of 16 subblocks, each with a size of 16x16, and each 16x16 subblock can be further divided into four subblocks, each with a size of 8x8. In this example, sbIdx2 can take one of 16 unique values ​​and sbIdx3 can take one of 4 unique values, so each of the 64 8x8 subblocks can be identified based on the corresponding sbIdx2 and sbIdx3 indices.

[0178] Embodiments for improving the search strategies described above are described herein. Embodiments described herein may be applied to one or more coding techniques (e.g., coding, decoding, or composite coding-decoding), such as one or more coding techniques in which blocks of a picture are predicted from one or more reference pictures using motion vectors (e.g., two reference blocks from two respective reference pictures), and the motion vectors are refined by a refinement technique. These improvements may be applied to any suitable video coding standard or format as described above (e.g., HEVC, VVC, AV1), other existing standards or formats that apply coding for blocks based on reference blocks from two respective reference pictures, and any future standards that use such techniques. Generally, when two reference blocks from two reference pictures are used, such techniques are generally called bi-predicted merge mode and bilateral matching techniques.

[0179] Rather than strictly adhering to the search strategies described above, the various embodiments described herein can allow different coded blocks to have different search strategies (or methods) for bilateral matching. A selected search strategy for a block may be signaled in one or more syntax elements coded in the bitstream. The search strategy includes constraints / relationships between MVD0 and MVD1 imposed during the bilateral matching search process and may also be associated with any combination of search patterns, search range or maximum number of searches, cost criteria, etc. In some cases, constraints may also be called constraints.

[0180] The systems and techniques described herein may be used to perform adaptive bilateral matching (BM) for decoder-side motion vector refinement (DMVR). For example, the systems and techniques can perform adaptive bilateral matching for DMVR by applying different search strategies and / or search methods for different coded blocks. In some embodiments, a selected search strategy for a block may be signaled using one or more syntax elements coded in the bitstream. In some examples, a selected search strategy for a block may be signaled using explicitly, implicitly, or a combination thereof. As will be described in more detail below, a selected search strategy (for example, one that will be used to perform adaptive BM for DMVR for a given block) may include constraints or relationships between MVD0 and MVD1, and the constraints of the search strategy may be used or applied during the bilateral matching search process. In some examples, additionally or alternatively, the search strategy may include one or more combinations of search patterns, search ranges, maximum number of searches, cost criteria, etc.

[0181] As an example for one explanation, the systems and techniques described herein can perform adaptive bilateral matching for DMVR using one or more motion vector difference (MVD) constraints. As previously described, motion vector differences can be used to represent the difference between the initial motion vector and the refined motion vector (e.g., MVD0 = MV0' - MV0 and MVD1 = MV1' - MV1).

[0182] In some cases, an MVD constraint may be selected for a given bilateral matching block (for example, it may be included in the selected search strategy). For example, the MVD constraint may be a mirroring constraint, in which case MVD0 and MVD1 are of the same magnitude but opposite signs (for example, MVD0 = -MVD1). The MVD mirroring constraint is sometimes referred to herein as the “first constraint”.

[0183] In another example, the MVD constraint can be set to MVD0=0 (for example, both the x and y components of MVD0 are 0). For example, the MVD0=0 constraint may be used by keeping MV0 fixed while exploring around MV1 to derive the refined MV1', such that MV0'=MV0. The MVD0=0 constraint may also be referred to as the “second constraint” in this specification.

[0184] In another example, the MVD constraint can be set to MVD1=0 (for example, both the x and y components of MVD1 are 0). For example, the constraint MVD1=0 may be used by keeping MV1 fixed while exploring around MV0 to derive the refined MV0', such that MV1'=MV1. The constraint MVD1=0 is sometimes referred to as the “third constraint” in this specification.

[0185] In another example, the MVD constraint may be used to independently explore around MV0 to derive MV0', and independently explore around MV1 to derive MV1'. The constraint for independent MVD exploration is sometimes referred to as the "fourth constraint" in this specification.

[0186] As will be explained in more detail below, in some cases only the first and second constraints may be offered as options for each bilateral matching block (e.g., included in the search strategy being selected or signaled). In some cases only the first and third constraints may be offered as options for each bilateral matching block (e.g., included in the search strategy being selected or signaled). In some cases only the first and fourth constraints may be offered as options for each bilateral matching block (e.g., included in the search strategy being selected or signaled). In some cases only the second and third constraints may be offered as options for each bilateral matching block (e.g., included in the search strategy being selected or signaled). In some cases the first, second, and third constraints may be offered as options for each bilateral matching block (e.g., included in the search strategy being selected or signaled). Any other combination of constraints may be offered as options for each bilateral matching block (e.g., included in the search strategy being selected or signaled).

[0187] In some embodiments, one or more syntax elements may be signaled (for example, in a bitstream), and one or more syntax elements may contain a value indicating whether one or more constraints apply. In some examples, the syntax elements may be used to indicate or determine which particular one of the above MVD constraints should apply to a given bilateral matching block.

[0188] In one example for explanation, the first syntax element may be used to indicate whether a first constraint applies. For example, the first syntax element may be used to indicate whether a mirroring constraint (e.g., MVD0 = -MVD1) should apply to a given bilateral matching block for which the first syntax element signals. In some cases, the first syntax element may have a first value when the mirroring constraint should apply and a second value when the mirroring constraint should not apply. In some cases, the presence of the first syntax element may be used to infer (e.g., implicitly signal) that the mirroring constraint should apply to a given or current bilateral matching block, while the absence of the first syntax element may be used to infer (e.g., implicitly signal) that the mirroring constraint should not apply.

[0189] Continuing the above example, if the first syntax element indicates that the first constraint (e.g., the MVD mirroring constraint, MVD0=-MVD1) does not apply, a second syntax element may be used to indicate which of the remaining constraints should apply. For example, a second syntax element may be used to indicate whether the second constraint (e.g., MVD0=0) or the third constraint (e.g., MVD1=0) should apply to a given or current bilateral matching block for which the syntax element signals. In some examples, the second syntax element may have a first value when the second constraint should apply, and a second value when the third constraint should apply. In some cases, the presence of a second syntax element can be used to infer (e.g., implicitly signal) that one of the second or third constraints should apply, while the absence of a second syntax element can be used to infer (e.g., implicitly signal) that the remaining of the second or third constraints should apply instead.

[0190] In some examples, one or more syntax elements (e.g., a first syntax element and / or a second syntax element) may contain mode information, and the selected constraint from the aforementioned MVD constraints may be determined based on the mode indicated by the mode information (e.g., merge mode). For example, one or more syntax elements may contain mode information indicating a different merge mode for the current bilateral matching block. In an example for one explanation, a first constraint (e.g., a mirroring constraint, MVD0=-MVD1) may be applied to a coded block of a normal merge mode (e.g., a standard or default merge mode such as an extended merge prediction in a VVC merge mode) when the normal merge candidate satisfies the DMVR conditions described above.

[0191] As previously described, a DMVR condition can indicate a CU coded using one or more of the following modes or features. In some examples, a DMVR condition may include, but is not limited to, a CU level merge mode with bipredictive MV, one reference picture in the past (e.g., with respect to the current picture) and another reference picture in the future (e.g., with respect to the current picture), the distance from the two reference pictures to the current picture (e.g., POC distance) is the same, both reference pictures are short-term reference pictures, a CU containing more than 64 luma samples, both the height and width of the CU being greater than or equal to 8 luma samples, BCW weight indices showing equal weights, weighted prediction (WP) not being enabled for the current block, synthetic inter and intra prediction (CIIP) modes not being utilized for the current block, etc.

[0192] In an example for further explanation, when a coded block uses a specified new merge mode (e.g., an adaptive bilateral matching mode as described herein), one of a second constraint (e.g., MVD0=0) or a third constraint (e.g., MVD1=0) may be applied, in which case all merge candidates satisfy the DMVR condition. In some cases, the second and / or third constraints are indicated by a mode flag or merge index, either as an addition or an alternative. For example, the indication of a choice between the second and third constraints may be determined based on a mode flag or merge index.

[0193] In some examples, one or more syntax elements described herein may include an index of a merge candidate list. The selected constraint may then be determined by the index indicating the selected merge candidate from the merge candidate list (for example, the constraint depends on the selected merge candidate). In yet another example, the syntax element may include a combination of a mode flag and an index.

[0194] In some embodiments, the systems and techniques described herein can signal (e.g., explicitly and / or implicitly) a selected search strategy that can be used to perform the multi-level (e.g., multi-path) DMVR described above. In some examples, the selected search strategy may be applied in one path of a multi-path DMVR. In other examples, the selected search strategy may be applied in multiple levels or paths of a multi-path DMVR (e.g., but not in all levels or paths of a process in some cases).

[0195] In the example of the 3-pass DMVR described above, in the example for one explanation, the selected strategy may be applied only in the first pass (e.g., PU-level bilateral matching). The second and third passes (e.g., performing bilateral matching for a first subblock size and BDOF for a second subblock size smaller than the first subblock size, respectively) may be performed using the default strategy, e.g., the one described above with respect to Figure 10 and a standardized 3-pass structure. In the example for one explanation, the multi-pass DMVR (e.g., the 3-pass DMVR described above) may utilize the second search strategy (e.g., second constraint MVD0=0) and / or the third search strategy (e.g., third constraint MVD1=0) in the first pass to perform PU-level bilateral matching. Subsequent passes (e.g., the second and third passes) may use the default search strategy including the first constraint (e.g., mirroring constraint MVD0=-MVD1).

[0196] In some embodiments, search strategies can be grouped into multiple subsets. In some examples, one or more syntax elements may be used to determine the selected subset. In some cases, the selected strategy within a given subset may be implicitly determined. In an example for one explanation, a first constraint (e.g., mirroring constraint MVD0=-MVD1) may be included in the first subset, and both a second constraint (e.g., MVD0=0) and a third constraint (e.g., MVD1=0) may be included in the second subset. Syntax elements may be used to indicate whether the second subset is used. If the second subset is used (e.g., based on the corresponding syntax element), the choice or decision between applying the second constraint and applying the third constraint may be implicitly determined. For example, the implicit decision between using the second constraint for bilateral matching and using the third constraint for bilateral matching may be based on the minimum matching cost. The second constraint can be chosen if it is determined that bilateral matching using the second constraint results in a lower matching cost than bilateral matching using the third constraint; otherwise, the third constraint is chosen.

[0197] In some examples, one or more aspects of the systems and techniques described herein may be used in conjunction with or applied on an Enhanced Compression Model (ECM). For example, in an ECM, a multipath DMVR may be applied to a standard (or default) merge mode candidate, such as an enhanced merge prediction in a VVC merge mode, along with one or more features, some or similar to those described above. For example, one or more features may be the same or similar to some (or all) of the DMVR conditions described above. As stated above, a merge candidate refers to a candidate block from which information (e.g., one or more motion vectors, prediction modes, etc.) is inherited for use when coding (e.g., encoding and / or decoding) the current block, and the candidate block may be a block adjacent to the current block. For example, a merge candidate may be an interpredicted PU containing a group of spatially adjacent motion data locations and a motion data location selected from one of two temporally identical motion data locations.

[0198] Multipath DMVR in ECM may be performed on the basis of applying a first constraint (e.g., mirroring constraint MVD0=-MVD1) by default. For illustrative purposes, the systems and techniques described herein may utilize adaptive bilateral matching mode as a new mode for multipath DMVR. In some examples, adaptive bilateral matching mode may also be referred to as "adaptive_bm_mode". In adaptive_bm_mode, a merge index may be signaled to indicate selected motion information candidates. However, all candidates in the candidate list may satisfy the DMVR conditions. In some cases, a flag (e.g., bm_merge_flag) may be used to signal or indicate the use of adaptive bilateral matching mode. For example, if the flag is true (e.g., bm_merge_flag is equal to 1), adaptive_bm_mode may be used or applied (e.g., as described below). In some embodiments, when a flag is true (for example, bm_merge_flag is equal to 1), an additional flag (for example, bm_dir_flag) may be used to signal or indicate the bm_dir value that should be used in adaptive_bm_mode.

[0199] In some embodiments, adaptive bilateral matching can be performed when adaptive_bm_mode=1, and not when adaptive_bm_mode=0. For example, an adaptive bilateral matching process (e.g., associated with adaptive_bm_mode=1) may be performed on at least part basis of applying either a second or third constraint to a selected candidate (e.g., fixing either MVD0 or MVD1 to equal to 0, respectively). In some cases, the variable bm_dir may be used to indicate which constraints apply and / or to indicate selected constraints. For example, when adaptive bilateral matching is performed (e.g., signaled or determined based on adaptive_bm_mode), bm_dir=1 may be used to indicate or signal that adaptive bilateral matching should be performed by fixing MVD1 to 0 (e.g., the third constraint should apply). In some examples, bm_dir=2 may be used to indicate or signal that adaptive bilateral matching should be performed by fixing MVD0 to 0 (for example, that a second constraint should be applied).

[0200] In some cases, if adaptive bilateral matching is not performed, the normal merge mode may be used. As previously mentioned, the normal merge mode may be determined or signaled based on adaptive_bm_mode=0, which indicates that adaptive bilateral matching is not performed. In some examples, when the normal merge mode is used (e.g., because adaptive bilateral matching is not performed), the system and technique may use the inference bm_dir=3, which indicates or signals that MVD0 and MVD1 are not fixed and MVD0=-MVD1 (e.g., the first mirroring constraint should apply). In some examples, bm_dir=3 may be explicitly signaled or used to indicate the normal merge mode.

[0201] In some examples, the system and technique can perform bilateral matching with one or more modified bilateral matching operations. In some embodiments, when bm_dir=3, the bilateral matching process may be the same as the 3-pass bilateral matching process described above. For example, assuming an initial pair of MVs, the first predictor is generated by the first MV referencing the first reference picture, and the second predictor is generated by the second MV referencing the second reference picture. The refined pair of MVs is then derived by minimizing the BM cost between the refined first predictor (e.g., generated using the first refined MV) and the refined second predictor (e.g., generated using the second refined MV), where the motion vector difference between the refined MV and the initial MV is MVD0 and MVD1, and MVD0=-MVD1.

[0202] In some embodiments, when bm_dir=1, only the first MV is refined, while the second MV remains fixed. For example, the refined first MV can be derived by minimizing the BM cost between the second predictor generated by the second MV and the refined first predictor generated by the refined first MV. For example, when bm_dir=1, MV0 can be refined, but MV1 remains fixed. The refined motion vector MV0' can be derived by minimizing the BM cost between the predictor generated based on MV1 and the refined predictor generated based on MV0'.

[0203] In some cases, when bm_dir=2, only the second MV is refined, while the first MV remains fixed. For example, the refined second MV can be derived by minimizing the BM cost between the first predictor generated by the first MV and the refined second predictor generated by the refined second MV. For example, when bm_dir=2, MV1 can be refined, but MV0 remains fixed. The refined motion vector MV1' can be derived by minimizing the BM cost between the predictor generated based on MV1 and the refined predictor generated based on MV1'.

[0204] In some embodiments, the BM cost (e.g., as described above) may include, as an addition or alternative, a regularization term based on or derived from the motion vector difference (MVD). In an example for one explanation, a multipath DMVR search process may include a regularization term determined based on the MV cost, which depends on the refined MV position.

[0205] In some embodiments, one or more of the bilateral matching modifications described above may be applied only in the first pass of a multi-pass bilateral matching process (e.g., a PU-level DMVR). In other embodiments, one or more of the bilateral matching modifications described above may be applied in both the first and second passes (e.g., in the PU-level DMVR pass and the sub-PU-level DMVR pass).

[0206] In some examples, the systems and techniques described herein can perform multipath bilateral matching DMVR using more or fewer paths than those described in the examples above. For example, fewer than three paths may be used, and / or more than three paths may be used. In some examples, any number of paths may be used, and the paths may be constructed in any manner. In some embodiments, multipath design may be applied similarly in adaptive_bm_mode. In some embodiments, the second path may be skipped in adaptive_bm_mode. In some embodiments, both the second and third paths may be skipped in adaptive_bm_mode. In other embodiments, any combination of paths (e.g., path behavior from the three-path system described above) may be combined with iterative paths or other path types based on specific adaptive bilateral matching criteria.

[0207] In some aspects of multi-pass DMVR, different search patterns may be used for different search levels and / or search precisions. For example, square search may be used for both integer search and half-PEL search, which may be applied to perform a PU-level DMVR (e.g., the first pass). In some examples, in sub-PU-level DMVRs (e.g., the second and / or third passes), a complete search may be used for integer search and square search for half-PEL search. In some examples, one or more (or all) of the search patterns described above may be used based on the determination that bm_dir is equal to a first value (a value of 1), a second value (a value of 2), or a third value (a value of 3). In one example, one or more (or all) of the search patterns described above may be used based on the determination that bm_dir = 3.

[0208] The following embodiments describe exemplary search patterns and / or search processes when bm_dir is equal to 1 or 2. In one embodiment, the same search pattern as when bm_dir=3 may be used when bm_dir=1 and / or bm_dir=2. In another embodiment, bm_dir=1 and bm_dir=2 may use different search patterns than when bm_dir=3. For example, in one example, a complete search may be used for integer search in a PU-level DMVR, and a square search may be used for semi-Perl search in a PU-level DMVR.

[0209] In some embodiments, the same search range and / or maximum number of searches may be used for different values ​​of bm_dir. In other embodiments, different search ranges and / or different maximum number of searches may be used when bm_dir = 1 or 2. For example, in the case of a complete search where different cost coefficients are assigned to different MVDs, one or more MVD regions may be skipped. For example, as shown in Figure 10, the search area for CU1000 is divided into multiple search regions (e.g., a first search region 1020, a second search region 1030, a third search region 1040, and a fourth search region 1050). In some cases, regions far from the center of the search area (e.g., far from the first search region 1020) may be skipped.

[0210] In some cases, all search regions can be explored in normal merge mode. For example, with respect to Figure 10, in normal merge mode, the four search regions 1020, 1030, 1040, and 1050 can be explored along with a fifth search region containing the remaining blocks of CU1000 that are not yet included in one of the four search regions 1020-1050. As mentioned above, in some cases, search regions far from the center of the search area can be skipped. For example, in adaptive bilateral matching mode (for example, when bm_dir=1 or 2), only the first three search regions related to CU1000 in Figure 10 can be explored (for example, in adaptive bilateral matching mode, the first search region 1020, the second search region 1030, and the third search region 1040 can be explored).

[0211] In some embodiments of multipath DMVRs, SAD or mean-removed SAD (e.g., depending on PU size) may be used for integer search and half-PEL search associated with PU-level DMVR paths. In some cases, SATD may be used for sub-PU-level DMVRs. In some embodiments, the same cost criterion may be used for different values ​​of bm_dir. For example, the current cost criterion selection in the ECM may be used for all values ​​of bm_dir. In other embodiments, the cost criterion selection may differ for different values ​​of bm_dir. For example, when bm_dir=3, the current cost criterion selection in the ECM may be applied. SAD or mean-removed SAD may be used for integer search or half-PEL search in PU-level DMVRs, depending on PU size. When bm_dir=1 or 2, the PU-level DMVR process may use SATD for integer search and SAD for half-PEL search.

[0212] As previously mentioned, candidates in the candidate list for some embodiments satisfy the DMVR conditions. In one additional embodiment, bm_dir may be set to equal to 1 or 2, as indicated by the additional bm_dir_flag included in the adaptive_bm_mode mode. In some cases, the candidate list for adaptive_bm_mode is generated in addition to the normal merge candidate list. For example, one or more candidates in the normal merge candidate list that are determined to satisfy the DMVR conditions may be inserted into the candidate list for adaptive_bm_mode.

[0213] In another additional aspect, whether bm_dir is set to equal 1 or 2 may be indicated by the merge index within the adaptive_bm_mode mode. A candidate list for adaptive_bm_mode may be generated in addition to the normal merge candidate list. For each candidate in the normal merge candidate list that satisfies the DMVR condition, a pair of two candidates can be inserted into the adaptive_bm_mode candidate list, where one candidate has bm_dir=1 and the other has bm_dir=2, and the two candidates in the pair have identical motion information. In some examples, bm_dir may be determined by determining whether the merge index is even or odd.

[0214] For example, the candidate list for adaptive_bm_mode may be generated independently of the normal merge candidate list. In some cases, generating the candidate list for adaptive_bm_mode may follow the same or similar process as generating the candidate list for the normal merge mode (e.g., checking the same spatial, temporal, history-based candidates, pairwise candidates, etc.). In some cases, pruning may be applied during the list building process.

[0215] In yet another embodiment, one or more candidates related to a biprediction (BCW) weight index using CU-level weights that exhibit unequal weights may also be added to the candidate list (for example, compared to some systems with DMVR conditions where the BCW weight index may exhibit equal weights).

[0216] In some cases, padding may be applied if the number of candidates in the candidate list for adaptive_bm_mode and / or the candidate list for normal merge mode is less than a predetermined maximum. For example, when the candidate list for adaptive_bm_mode is generated, the number of candidates in the list may be less than a predetermined maximum number of candidates. In such cases, padding may be applied to generate a certain number of padded candidates for the candidate list so that the padded candidate list contains a predetermined number of candidates. In one example for explanation, when adaptive bilateral matching is enabled (e.g., adaptive_bm_mode=1), one or more default candidates may be used for padding in merge list construction. The default candidates may be derived to satisfy the DMVR conditions.

[0217] In some examples, MV may be set to 0 for the default candidate. For example, zero MV candidates may be added during padding. In some examples, a reference picture may be selected according to a DMVR condition. In some cases, the reference index can be iterated over all possible values ​​until the number of candidates in the candidate list reaches the maximum number of candidates (e.g., a predetermined maximum value). In another embodiment, a BCW weight index can be used to indicate equal weight for normal candidates, and one or more unequal weighted BCW candidates may be added thereafter, but before zero candidates are added.

[0218] In some embodiments, the reference picture assigned to the default candidate (for example, the reference picture assigned to the padded zero MV candidate described above) may be selected to satisfy one or more conditions related to adaptive_bm_mode. Examples for describing such conditions include, but may also include, one or more of the following conditions: that at least one pair of reference pictures is selected, each containing one past reference picture and one future reference picture relative to the current picture; that the distances from both reference pictures to the current picture are equal; that neither of the reference pictures is a long-term reference picture; that both reference pictures have the same resolution as the current picture; that no weighted prediction (WP) is applied to either of the reference pictures; any combination of these; and / or other conditions, and / or other conditions not listed herein.

[0219] In some embodiments, one or more reference pictures may be assigned to a default candidate (e.g., a padded zero MV candidate as described above) based on the selection of reference pictures that satisfy one or more conditions related to or based on a specified or selected constraint (e.g., a second constraint MVD0=0, or a third constraint MVD1=0). In some examples, reference pictures may be selected to satisfy one or more conditions related to a given constraint, which is determined based on bm_dir (e.g., as described earlier). For example, one or more conditions based on bm_dir may only apply to MVs for which refinement is performed (e.g., bm_dir=1). Examples of such conditions include one or more of the following: a reference picture in a given list X is not a long-term reference picture; a reference picture in list X has the same resolution as the current picture; weighted predictions (WP) are not applied to the reference pictures in list X; the distance from each first reference picture in list X to the current picture is not less than the distance from each other reference picture (e.g., a second reference picture in list X) to the current picture; any combination of these; and / or other conditions, and / or other conditions not listed herein.

[0220] In some cases, if bm_dir indicates that the MVs in List 0 are refined, then List X may be the same as List 0 (e.g., List L0). In some cases, if bm_dir indicates that the MVs in List 1 are refined, then List X may be the same as List 1 (e.g., List L1). In some embodiments, all potential zero MV candidates can be discovered by iterating over all possible combinations of reference pictures and identifying reference pictures that satisfy predetermined conditions in a certain order. In an example for one explanation, a first loop may be performed for List 0 and a second loop for List 1. In another example, a first loop may be performed for List 1 and a second loop for List 0. Other orders are also possible and should be considered within the scope of this disclosure. The process of determining possible zero MV candidates (e.g., by iterating over combinations of reference pictures and identifying reference pictures that satisfy predetermined conditions in a certain order) may be performed at the slice level, picture level, or other levels. A list of identified default MV candidates (e.g., zero MV candidates) may be stored as default candidates. In some cases, when determining potential zero MV candidates at the block level, if the number of candidates is less than a predetermined maximum number of candidates, the systems and techniques described herein may iterate through the default candidates and add one or more default candidates to the candidate list until the number of candidates reaches the predetermined maximum.

[0221] In some embodiments, one or more size constraints may be included in and / or utilized by the adaptive_bm_mode described herein. In one embodiment, the same size constraints as those for a normal DMVR may apply in adaptive_bm_mode. In another embodiment, adaptive_bm_mode does not apply if neither the width nor the height of the current block is greater than the minimum block size for the DMVR.

[0222] In some embodiments, adaptive_bm_mode may be signaled as an additional merge mode to the normal merge mode. In some examples, various signaling methods may be applied or utilized to signal adaptive_bm_mode as an additional merge mode. For example, adaptive_bm_mode may be considered or signaled as a variation of the normal merge mode. In an example for one explanation, one or more syntax elements may first be signaled to indicate the normal merge mode, and one or more additional flags and / or syntax elements may be signaled to indicate adaptive_bm_mode and / or to indicate a particular one of the constraints that may be applied in connection with the use of adaptive_bm_mode (e.g., a second constraint MDV0=0 or a third constraint MDV1=0).

[0223] In another embodiment, adaptive_bm_mode may be indicated by one or more flags before the indication of the normal merge mode. For example, if the syntax (e.g., one or more syntax elements) indicates that the current block is not using adaptive_bm_mode, one or more additional syntax elements may be signaled to indicate whether the current block is using the normal merge mode or another merge mode. For example, if the current block is not using adaptive_bm_mode or the normal merge mode, one or more additional syntax elements may be signaled to indicate that the current block is using another merge mode, such as synthetic inter- and intra-prediction (CIIP) or geometric partition mode (GPM).

[0224] In yet another embodiment, adaptive_bm_mode may be signaled in other merge mode branches. For example, adaptive_bm_mode may be signaled in template matching merge mode branches. In some cases, one or more syntax elements may be initially signaled to indicate whether adaptive_bm_mode or one of the template matching merge modes is used. If one or more syntax elements indicate that adaptive_bm_mode or the template matching merge mode is used, one or more additional flags or syntax elements may be signaled to indicate whether the template matching merge mode or adaptive_bm_mode is used.

[0225] In some embodiments, the merge index in adaptive_bm_mode can use the same signaling method as in normal merge mode. In one embodiment, adaptive_bm_mode can use the same (or similar) context model as in normal merge mode. In another embodiment, a separate context model may be used for adaptive_bm_mode. In some examples, the maximum number of merge candidates for adaptive_bm_mode may differ from the maximum number of merge candidates for normal merge mode.

[0226] In one example for explanation, one or more high-level syntax elements may be used to indicate whether adaptive_bm_mode may or will be applied. In one embodiment, the same high-level syntax used to indicate whether the normal DMVR will be applied may also be used to indicate whether adaptive_bm_mode will be applied. In another embodiment, one or more separate (e.g., additional) high-level syntax elements may be used to indicate whether adaptive_bm_mode should be applied. In yet another embodiment, a separate high-level syntax element may be used to indicate whether adaptive_bm_mode is utilized, and a separate high-level syntax element for adaptive_bm_mode exists only when the normal DMVR is enabled. For example, if a separate high-level syntax element associated with the normal DMVR is determined to indicate that the normal DMVR is not enabled or utilized, then a separate or additional high-level syntax element associated with adaptive_bm_mode will not be signaled, and adaptive_bm_mode will be inferred to be off (e.g., not enabled or utilized).

[0227] In some embodiments, in addition to one or more high-level syntax elements described above, adaptive_bm_mode may be disabled for coded pictures or slices depending on the available reference pictures. In some cases, if it is determined that no combination of reference pictures satisfies or can not satisfy the reference picture conditions, adaptive_bm_mode may be disabled and the corresponding syntax element (e.g., at the block level) will not be signaled. In some cases, in order to utilize adaptive_bm_mode, there must be at least one pair of reference pictures that satisfy the reference picture conditions. Examples for describing such conditions may include one or more of the following conditions and / or other conditions not listed herein: one past reference picture and one future reference picture relative to the current picture; the respective distances from both reference pictures to the current picture are equal; neither reference picture is a long-term reference picture; both reference pictures have the same resolution as the current picture; no weighted prediction (WP) is applied to either reference picture; any combination of these; and / or other conditions.

[0228] The conditions listed above can be used separately or in combination.

[0229] In some embodiments, only a subset of the adaptive bilateral matching modes described herein (for example, a first adaptive bilateral matching mode associated with a second condition MDV0=0 and a second adaptive bilateral matching mode associated with a third condition MDV1=0) may be enabled depending on the reference picture. Examples of conditions that allow only bm_dir=1 (for example, related to the second condition MDV0=0) or bm_dir=2 (for example, related to the third condition MDV1=0) include one or more of the following conditions, and / or other conditions not listed herein: one of the reference pictures is a long-term reference picture and the other reference picture is not a long-term reference picture; one of the reference pictures has the same resolution as the current picture but the other reference picture has a different resolution than the current picture; a weighted prediction (WP) is applied to one of the reference pictures; the distance from one of the reference pictures to the current picture is not shorter than the distance from the other reference picture to the current picture; any combination of these, and / or other conditions.

[0230] In some cases, syntax elements that identify the bilateral matching mode (e.g., at the block level) do not need to be signaled and can be inferred accordingly. In some examples, syntax elements may be implicitly signaled (e.g., inferred) rather than explicitly signaled (e.g., in a bitstream as part of a particular syntax table). In some cases, syntax elements may not need to be explicitly signaled or implicitly signaled, and can be inferred. For example, if a first value of bm_dir is enabled (e.g., bm_dir=1 is enabled) but a second value of bm_dir is disabled (e.g., bm_dir=2 is disabled), a syntax element is used to indicate that bm_dir at the block level does not need to be signaled. In the absence of this syntax element, bm_dir can be inferred to be the first value (e.g., bm_dir is inferred to be 1). In another example, if a second value for bm_dir is enabled (e.g., bm_dir=2) but the first value for bm_dir is disabled (e.g., bm_dir=1 is disabled), a syntax element is used to indicate that bm_dir does not need to be signaled at the block level. Without this syntax element, bm_dir could be inferred to be the second value (e.g., bm_dir could be inferred to be 2).

[0231] In another example, slice-level and / or picture-level flags may be used for adaptive_bm_mode. For example, slice-level and / or picture-level flags may be used as bitstream conformance constraints, where if at least one of the above conditions is not met, the flag is set to 0 (for example, DMVR mode is disabled).

[0232] In yet another example, bitstream conformance constraints can be introduced into existing signaling, and if at least one of the above conditions is not met, the bitstream conformance constraint indicates that adaptive_bm_mode should not be applied and the corresponding overhead should be set to 0 (for example, indicating that adaptive_bm_mode will not be used).

[0233] In some aspects of ECM, multiple hypothesis prediction (MHP) may be used. In MHP, inter-prediction techniques may be used to obtain or determine a weighted superposition of more than two motion-compensated prediction signals. Based on performing a sample-by-sample weighted superposition, the resulting overall prediction signal may be obtained. For example, single / dual prediction signal p uni / bi Based on the first additional interpretation signal / hypothesis h3, the resulting prediction signal p3 can be obtained as follows: p3 = (1-α)p uni / bi +αh3 formula (27)

[0234] Here, the weighting coefficient α can be specified by the syntax element add_hyp_weight_idx according to the following mapping.

[0235] [Table 2]

[0236] In some examples, more than one additional predictor signal may be used. In some cases, more than one additional predictor signal may be used in the same or similar manner as above. For example, when using multiple additional predictor signals, the resulting overall predictor signal may be iteratively accumulated for each additional predictor signal as follows: p n+1 =(1-α n+1 ) p n +α n+1 h n+1 Formula (28)

[0237] Here, the resulting overall prediction signal is the last p n (For example, p with the largest index n) n It can be obtained as follows:

[0238] In some embodiments, MHP may not apply to any adaptive_bm_mode (e.g., it may be disabled). In some embodiments, MHP may apply in addition to adaptive_bm_mode in the same or similar manner as MHP applies to standardized (e.g., normal) merge modes.

[0239] Figure 11 is a flowchart illustrating an example of process 1100 for processing video data. In some examples, process 100 may be used to perform decoder-side motion vector refinement (DMVR) with adaptive bilateral matching, as described in some examples of this disclosure. In some embodiments, process 1100 may be implemented in a device for processing video data, comprising memory and one or more processors coupled to the memory configured to perform the operation of process 1100. In other embodiments, process 1100 may be implemented in a non-temporary computer-readable medium, comprising instructions that cause the device to perform the operation of process 1100 when executed by one or more processors of the device.

[0240] In block 1102, process 1100 includes obtaining one or more reference pictures for the current picture (for example, the current block of the current picture). For example, one or more reference pictures may be obtained based on one or more inputs 114 to the decoding device 112 shown in Figure 1. In some examples, one or more reference pictures and the current picture may be obtained from video data obtained by or provided to the decoding device 112 shown in Figure 1.

[0241] In block 1104, process 1100 includes identifying a first motion vector and a second motion vector for merge mode candidates. For example, the first motion vector and / or the second motion vector may be identified by the decoding device 112 shown in Figure 1. In some examples, the first motion vector and / or the second motion vector may be identified using the decoder engine 116 of the decoding device 112 shown in Figure 1. In some cases, one or more (or both) of the first and second motion vectors may be identified using signaled information. For example, the encoding device 104 shown in Figure 1 may include signaling information that can be used by the decoding device 112 and / or the decoding engine 116 to identify one or more (or both) of the first and second motion vectors. In some cases, process 1100 may include determining merge mode candidates for the current picture. As described herein, merge mode candidates may include adjacent blocks of blocks from which prediction data can be inherited for the current picture blocks. For example, merge mode candidates may be determined by the decoding device 112 shown in Figure 1. In some examples, merge mode candidates may be determined using the decoder engine 116 of the decoding device 112 shown in Figure 1. In some examples, signaling information may be used by the decoding device 112 and / or the decoding engine 116 to determine merge mode candidates for the current picture. In some cases, merge mode candidates may be determined using the same signaling information used to identify one or more (or both) first and second motion vectors associated with the merge mode candidate. In some cases, merge mode candidates and the first and second motion vectors may be determined using separate signaling information.

[0242] In block 1106, process 1100 includes determining a selected motion vector search strategy for merge mode candidates from a plurality of motion vector search strategies. In some aspects, the selected motion vector search strategy is associated with one or more constraints based on or corresponding to the first motion vector and / or the second motion vector. In one example for illustration, the selected motion vector search strategy can be a bilateral matching (BM) motion vector search strategy. In some cases, the motion vector search strategy for merge mode candidates can be selected from a plurality of motion vector search strategies including at least two of a multi-pass decoder-side motion vector refinement strategy, a fractional sample refinement strategy, a bidirectional optical flow strategy, or a sub-block-based bilateral matching motion vector refinement strategy. In some examples, the selected motion vector search strategy can be determined before the merge mode candidates are determined. For example, the selected motion vector search strategy can be determined (e.g., as described above) and used to generate a merge candidate list. The merge mode candidates can be determined based on a selection from the generated merge candidate list. In some examples, the selected merge mode candidates can be determined before the selected motion vector search strategy is determined. For example, in some cases, a merge candidate list can be generated without using the selected search strategy (e.g., the generated merge candidate list can be the same for each of the plurality of search strategies), and the selected merge candidates can be determined before the selected search strategy.

[0243] In some examples, the selected motion vector search strategy may be a multipath decoder-side motion vector refinement (DMVR) search strategy. For example, a multipath DMVR search strategy may include one or more block-based bilateral matching motion vector refinement paths and one or more subblock-based motion vector refinement paths. In some examples, one or more block-based bilateral matching motion vector refinement paths may be performed using a first constraint associated with a first motion vector difference and / or a second motion vector difference. The first motion vector difference may be the difference determined between the first motion vector and the refined first motion vector. The second motion vector difference may be the difference determined between the second motion vector and the refined second motion vector. In some examples, one or more subblock-based motion vector refinement paths may be performed using a second constraint different from the first constraint. As described above, the second constraint may be associated with at least one of the first motion vector difference and / or the second motion vector difference. In some cases, one or more subblock-based motion vector refinement paths may include at least one of the following: a subblock-based bilateral matching (BM) motion vector refinement path and / or a subblock-based bidirectional optical flow (BDOF) motion vector refinement path.

[0244] In some examples, the selected motion vector search strategy is associated with one or more constraints corresponding to at least one of the first motion vector or the second motion vector (e.g., as described above). The one or more constraints may be determined based on one or more signaled syntax elements. For example, the one or more constraints may be determined for a block of the current picture based on the syntax elements signaled for the block. In some aspects, the one or more constraints are associated with at least one of a first motion vector difference associated with the first motion vector (e.g., the difference between the first motion vector and the refined first motion vector) and a second motion vector difference associated with the second motion vector (e.g., the difference between the second motion vector and the refined second motion vector). In some examples, the one or more constraints may include mirroring constraints for the first motion vector difference and the second motion vector difference. The mirroring constraints can set the first motion vector difference and the second motion vector difference to have equal magnitudes (e.g., absolute values) but opposite signs. In some cases, the one or more constraints may include a zero-value constraint for the first motion vector difference (e.g., setting the first motion vector difference equal to 0). In some examples, the one or more constraints may include a zero-value constraint for the second motion vector difference (e.g., setting the second motion vector difference equal to 0). In some aspects, the zero-value constraint may indicate keeping the motion vector difference constant. For example, based on the zero-value constraint, process 1100 may include keeping the first of the first motion vector difference or the second motion vector difference at a constant value and searching around the second of the first motion vector difference or the second motion vector difference to determine one or more refined motion vectors using the selected motion vector search strategy. For example, the first motion vector difference may be fixed, and the search may be performed around the second motion vector difference to derive the refined motion vector.

[0245] In block 1108, process 1100 includes determining one or more refined motion vectors based on a first motion vector, a second motion vector, and / or one or more reference pictures (for example, based on the first motion vector and one or more reference pictures, based on the second motion vector and one or more reference pictures, or based on the first motion vector, the second motion vector, and one or more reference pictures) using a selected motion vector search strategy. In some cases, determining one or more refined motion vectors may include determining one or more refined motion vectors for a block of video data. In some examples, one or more refined motion vectors may include a first refined motion vector and a second refined motion vector, which are determined for the first and second motion vectors, respectively. In some examples, a first motion vector difference is determined as the difference between the first refined motion vector and the first motion vector, and a second motion vector difference is determined as the difference between the second refined motion vector and the second motion vector.

[0246] In some examples, the selected motion vector search strategy is a bilateral matching (BM) motion vector search strategy, as previously mentioned. When the selected motion vector search strategy is a BM motion vector search strategy, determining one or more refined motion vectors may involve determining a first refined motion vector by searching for a first reference picture around a first motion vector. The first reference picture may be searched around the first motion vector based on the selected motion vector search strategy. A second refined motion vector may be determined by searching for a second reference picture around a second motion vector based on the selected motion vector search strategy. The selected motion vector search strategy may include motion vector difference constraints (e.g., a mirroring constraint that the magnitudes of the first motion vector difference and the second motion vector difference are equal but their signs are opposite, a constraint that sets the first motion vector difference equal to 0, a constraint that sets the second motion vector difference equal to 0, etc.). In some examples, the first refined motion vector and the second refined motion vector may be determined by minimizing the difference between the first reference block associated with the first refined motion vector and the second reference block associated with the second refined motion vector.

[0247] In block 1110, process 1100 includes processing merge mode candidates using one or more refined motion vectors. For example, the decoding device 112 shown in Figure 1 can process merge mode candidates using one or more refined motion vectors. In some examples, the decoder engine 116 of the decoding device 112 shown in Figure 1 can process merge mode candidates using one or more refined motion vectors.

[0248] In some implementations, the processes (or methods) described herein may be performed by a computing device or apparatus, such as the system 100 shown in Figure 1. For example, the process may be performed by an encoding device 104 shown in Figures 1 and 12, by another video source-side device or video transmission device, by a decoding device 112 shown in Figures 1 and 13, and / or by another client-side device, such as a player device, display, or any other client-side device. In some cases, the computing device or apparatus may include one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components configured to perform the steps of process 1100.

[0249] In some examples, computing devices may include mobile devices, tablet computers, extended reality (XR) devices (e.g., virtual reality (VR) devices such as head-mounted displays (HMDs), AR devices such as HMDs or augmented reality (AR) glasses, MR devices such as HMDs or mixed reality (MR) glasses), desktop computers, server computers and / or server systems, computing systems for vehicles or vehicle components, or other types of computing devices, as well as wireless communication devices. Components of a computing device (e.g., one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) may be implemented in circuitry. For example, a component may include, and / or be implemented using, one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuits), electronic circuits or other electronic hardware, and / or may include, and / or be implemented using, computer software, firmware, or any combination thereof to perform the various operations described herein. In some examples, a computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) containing video frames. In some examples, the camera or other capture device that captures the video data is separate from the computing device, in which case the computing device receives or acquires the captured video data.A computing device may include a network interface configured to communicate video data. The network interface may be configured to communicate Internet Protocol (IP) based data or other types of data. In some examples, the computing device or apparatus may include a display for showing output video content, such as a sample picture of a video bitstream.

[0250] A process is described in relation to a logical flow diagram, and its operation represents a sequence of operations that can be implemented by hardware, computer instructions, or a combination thereof. In the context of computer instructions, operation represents a computer-executable instruction stored in one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which operations are described is not intended to be interpreted as restrictive, and any number of operations described may be combined in any order and / or in parallel to implement a process.

[0251] In addition, the process may be executed under the control of one or more computer systems consisting of executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together by hardware or a combination thereof on one or more processors. As stated above, the code may be stored on a computer-readable or machine-readable storage medium in the form of a computer program comprising multiple instructions that can be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-temporary.

[0252] The coding techniques discussed herein may be implemented in exemplary video coding and decoding systems (e.g., System 100). In some examples, the system includes a source device that provides coded video data to be later decoded by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source and destination devices may comprise any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, and the like. In some cases, the source and destination devices may be compatible with wireless communication.

[0253] The destination device may receive encoded video data to be decoded via a computer-readable medium. The computer-readable medium may comprise any type of medium or device capable of moving the encoded video data from the source device to the destination device. For example, the computer-readable medium may comprise a communication medium that enables the source device to transmit the encoded video data directly to the destination device in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device. The communication medium may comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from the source device to the destination device.

[0254] In some examples, encoded data may be output to a storage device via an output interface. Similarly, encoded data may be accessed from a storage device via an input interface. The storage device may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In further examples, the storage device may correspond to a file server or another intermediate storage device capable of storing encoded video generated by the source device. The destination device may access the stored video data from the storage device via streaming or download. The file server may be any type of server capable of storing encoded video data and sending that encoded video data to the destination device. Exemplary file servers include web servers (for example, for websites), FTP servers, network-attached storage (NAS) devices, or local disk drives. The destination device may access the encoded video data through any standard data connection, including an internet connection. This may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on a file server. Transmission of encoded video data from the storage device may be via streaming, download, or a combination thereof.

[0255] The techniques of this disclosure are not necessarily limited to wireless applications or configurations. The techniques may be applied to video coding supporting any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system may be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video phone calls.

[0256] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device may include an input interface, a video decoder, and a display device. The video encoder of the source device may be configured to apply the techniques disclosed herein. In other examples, the source and destination devices may include other components or configurations. For example, the source device may receive video data from an external video source, such as an external camera. Similarly, the destination device may interface with an external display device rather than including an integrated display device.

[0257] The exemplary system described above is merely an example. Techniques for processing video data in parallel can be performed by any digital video encoding and / or decoding device. Generally, the techniques of this disclosure are performed by video encoding devices, but the techniques can also be performed by video encoders / decoders, commonly referred to as “codecs.” Furthermore, the techniques of this disclosure can also be performed by video processors. Source and destination devices are merely examples of encoding devices such that the source device generates encoded video data to send to the destination device. In some examples, the source and destination devices may operate substantially symmetrically, such that each device includes video encoding and decoding components. Thus, the exemplary system may support one-way or two-way video transmission between video devices for, for example, video streaming, video playback, video broadcasting, or video phone calls.

[0258] A video source may include a video capture device such as a camcorder, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. Alternatively, a video source may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In some cases, if the video source is a camcorder, the source and destination devices may form a so-called camera phone or videophone. However, as stated above, the techniques described in this disclosure may be applicable to video coding in general and to wireless and / or wired applications. In each case, captured, pre-captured, or computer-generated video may be encoded by a video encoder. The encoded video information may then be output onto a computer-readable medium by an output interface.

[0259] As stated, computer-readable media may include temporary media such as wireless broadcasting or wired network transmission, or storage media (i.e., non-temporary storage media) such as hard disks, flash drives, compact discs, digital video discs, and Blu-ray discs, or other computer-readable media. In some examples, a network server (not shown) may receive encoded video data from a source device via network transmission, for example, and provide the encoded video data to a destination device. Similarly, a computing device in a media manufacturing facility, such as a disc stamping facility, may receive encoded video data from a source device and manufacture a disc containing the encoded video data. Thus, computer-readable media may be understood to include one or more computer-readable media in various forms in various examples.

[0260] The input interface of the destination device receives information from a computer-readable medium. The information on the computer-readable medium may include syntax information defined by a video encoder, which includes syntax elements describing the characteristics and / or processing of blocks and other coded units, such as picture groups (GOPs), and this syntax information is also used by a video decoder. A display device displays the decoded video data to the user and may comprise any of various display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device. Various aspects of this application have been described.

[0261] Specific details of the encoding device 104 and the decoding device 112 are shown in Figures 12 and 13, respectively. Figure 12 is a block diagram illustrating an exemplary encoding device 104 capable of performing one or more of the techniques described herein. The encoding device 104 may, for example, generate syntax structures described herein (e.g., VPS, SPS, PPS, or syntax structures of other syntax elements). The encoding device 104 may perform intra-predictive and inter-predictive coding of video blocks within a video slice. As previously described, intra-coding relies at least partially on spatial prediction to reduce or eliminate spatial redundancy within a given video frame or picture. Inter-coding relies at least partially on temporal prediction to reduce or eliminate temporal redundancy within adjacent or surrounding frames of a video sequence. Intra-mode (I-mode) may refer to any of several spatial-based compression modes. Inter-mode, such as unidirectional prediction (P-mode) or bidirectional prediction (B-mode), may refer to any of several temporal-based compression modes.

[0262] The encoding device 104 includes a partitioning unit 35, a prediction processing unit 41, a filter unit 63, a picture memory 64, an adder 50, a transformation processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-prediction processing unit 46. For video block reconstruction, the encoding device 104 also includes an inverse quantization unit 58, an inverse transformation processing unit 60, and an adder 62. The filter unit 63 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample-adaptive offset (SAO) filter. The filter unit 63 is shown in Figure 12 as an in-loop filter, but in other configurations, the filter unit 63 may be implemented as a post-loop filter. The post-processing device 57 may perform additional processing on the encoded video data generated by the encoding device 104. The techniques of this disclosure may be performed by the encoding device 104 in some cases. However, in other cases, one or more of the techniques of this disclosure may be implemented by the post-processing device 57.

[0263] As shown in Figure 12, the encoding device 104 receives video data, and the partitioning unit 35 partitions the data into video blocks. Partitioning may also include partitioning into slices, slice segments, tiles, or other larger units, as well as video block partitioning according to a quadtree structure of LCUs and CUs, for example. The encoding device 104 generally represents the components that encode the video blocks within the video slice to be encoded. A slice may be divided into multiple video blocks (and possibly into sets of video blocks called tiles). The prediction processing unit 41 may select one of several possible coding modes for the current video block, such as one of several intra-predictive coding modes or one of several inter-predictive coding modes, based on the error results (e.g., coding rate and level of distortion). The prediction processing unit 41 may provide the resulting intra- or inter-coded blocks to the adder 50 to generate residual block data and to the adder 62 to reconstruct the encoded blocks for use as reference pictures.

[0264] An intra-prediction processing unit 46 within the prediction processing unit 41 may perform intra-prediction coding of the current video block for one or more adjacent blocks in the same frame or slice as the current block to be coded, in order to perform spatial compression. A motion estimation unit 42 and a motion compensation unit 44 within the prediction processing unit 41 perform inter-prediction coding of the current video block for one or more predicted blocks in one or more reference pictures, in order to perform temporal compression.

[0265] The motion estimation unit 42 may be configured to determine an inter-prediction mode for a video slice according to a predetermined pattern for a video sequence. The predetermined pattern may specify video slices in the sequence as P slices, B slices, or GPB slices. The motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion of a video block. The motion vector may indicate, for example, the displacement of a prediction unit (PU) of a video block in a current video frame or picture with respect to a prediction block in a reference picture.

[0266] The prediction block is a block that has been found to exactly match the PU of the video block to be coded with respect to pixel differences that may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some examples, the encoding device 104 may calculate values for sub-integer pixel positions of a reference picture stored in the picture memory 64. For example, the encoding device 104 may interpolate values at 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference picture. Thus, the motion estimation unit 42 may perform motion search for full pixel positions and fractional pixel positions and output a motion vector having fractional pixel positions.

[0267] The motion estimation unit 42 calculates a motion vector for the PU of a video block in an inter-coded slice by comparing the position of the PU with the position of a prediction block in a reference picture. The reference picture may be selected from a first reference picture list (list 0) or a second reference picture list (list 1) that each identify one or more reference pictures stored in the picture memory 64. The motion estimation unit 42 transmits the calculated motion vector to the entropy encoding unit 56 and the motion compensation unit 44.

[0268] Motion compensation performed by the motion compensation unit 44 may, in some cases, involve fetching or generating predicted blocks based on motion vectors determined by motion estimation, which perform interpolation to sub-pixel precision. Upon receiving the motion vector of the current video block from the PU, the motion compensation unit 44 may locate the position of the predicted block pointed to by the motion vector in the reference picture list. The encoding device 104 forms the residual video block by subtracting the pixel values ​​of the predicted block from the pixel values ​​of the current video block being coded to form a pixel difference value. The pixel difference value forms the residual data for the block and may include both lumens and chromens components. The adder 50 represents one or more components that perform this subtraction operation. The motion compensation unit 44 may also generate syntax elements related to video blocks and video slices for use by the decoding device 112 when decoding video blocks of video slices.

[0269] The intra-prediction processing unit 46 may intra-predict the current block as an alternative to inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44, as described above. Specifically, the intra-prediction processing unit 46 may determine which intra-prediction mode to use to encode the current block. In some examples, the intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, for example, between separate encoding passes, and the intra-prediction processing unit 46 may select an appropriate intra-prediction mode to use from the tested modes. For example, the intra-prediction processing unit 46 may use rate distortion analysis to calculate rate distortion values ​​for various tested intra-prediction modes and select an intra-prediction mode with the best rate distortion characteristics from the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to produce the encoded block, as well as the bit rate (i.e., number of bits) used to produce the encoded block. The intra-prediction processing unit 46 can calculate a ratio from the distortion and rate for various encoded blocks in order to determine which intra-prediction mode shows the best rate distortion value for the block.

[0270] In either case, after selecting an intra-prediction mode for a block, the intra-prediction processing unit 46 may provide the entropy coding unit 56 with information indicating the selected intra-prediction mode for the block. The entropy coding unit 56 may encode the information indicating the selected intra-prediction mode. The coding device 104 may include in the transmitted bitstream configuration data definitions for various blocks, as well as instructions for the most probable intra-prediction mode to be used for each context, an intra-prediction mode index table, and a modified intra-prediction mode index table. The bitstream configuration data may include multiple intra-prediction mode index tables and multiple modified intra-prediction mode index tables (also called codeword mapping tables).

[0271] After the prediction processing unit 41 generates a prediction block for the current video block via either interpretation or intraprediction, the encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block can be contained in one or more TUs and applied to the transformation processing unit 52. The transformation processing unit 52 transforms the residual video data into residual transformation coefficients using a transformation such as a discrete cosine transform (DCT) or a conceptually similar transformation. The transformation processing unit 52 may transform the residual video data from the pixel domain to a transformation domain such as the frequency domain.

[0272] The conversion processing unit 52 may transmit the obtained conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 may then perform a scan of the matrix containing the quantized conversion coefficients. Alternatively, the entropy coding unit 56 may perform the scan.

[0273] Following quantization, the entropy coding unit 56 entropy-codes the quantized transformation coefficients. For example, the entropy coding unit 56 can perform context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval piece entropy (PIPE) coding, or another entropy coding technique. Following entropy coding by the entropy coding unit 56, the encoded bitstream may be transmitted to the decoding device 112 or archived for later transmission or retrieval by the decoding device 112. The entropy coding unit 56 may also entropy-code the motion vector and other syntax elements of the currently coded video slice.

[0274] The inverse quantization unit 58 and the inverse transformation processing unit 60 apply inverse quantization and inverse transformation, respectively, to reconstruct residual blocks in the pixel region for later use as reference blocks of the reference picture. The motion compensation unit 44 may compute a reference block by adding the residual block to a predicted block of one of the reference pictures in the reference picture list. The motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to compute sub-integer pixel values ​​for use in motion estimation. The adder 62 adds the reconstructed residual block to the motion-compensated predicted block generated by the motion compensation unit 44 to generate a reference block for storage in the picture memory 64. The reference block may be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for interpreting blocks in subsequent video frames or pictures.

[0275] Thus, the encoding device 104 in Figure 12 represents an example of a video encoder configured to perform any of the techniques described herein, including the process described above with respect to Figure 11. In some cases, some of the techniques of this disclosure may also be performed by a post-processing device 57.

[0276] Figure 13 is a block diagram illustrating an exemplary decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, a filter unit 91, and a picture memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-prediction processing unit 84. In some examples, the decoding device 112 may perform a decoding path that is generally the reverse of the encoding path described with respect to the encoding device 104 from Figure 12.

[0277] During the decoding process, the decoding device 112 receives an encoded video bitstream representing video blocks and associated syntax elements of an encoded video slice transmitted by the encoding device 104. In some embodiments, the decoding device 112 may receive an encoded video bitstream from the encoding device 104. In some embodiments, the decoding device 112 may receive an encoded video bitstream from a network entity 79, such as a server, a media-aware network element (MANE), a video editor / splitter, or another such device configured to implement one or more of the techniques described above. The network entity 79 may or may not include the encoding device 104. Some of the techniques described herein may be performed by the network entity 79 before it transmits the encoded video bitstream to the decoding device 112. In some video decoding systems, the network entity 79 and the decoding device 112 may be part of a separate device, but in other instances, the functions described with respect to the network entity 79 may be performed by the same device comprising the decoding device 112.

[0278] The entropy decoding unit 80 of the decoding device 112 entropy-decodes the bitstream to generate quantized coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 transfers the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 may receive syntax elements at the video slice level and / or video block level. The entropy decoding unit 80 may process and parse both fixed-length and variable-length syntax elements in one or more parameter sets such as VPS, SPS, and PPS.

[0279] When a video slice is coded as an intra-coded (I) slice, the intra-prediction processing unit 84 of the prediction processing unit 81 may generate prediction data for the video blocks of the current video slice based on the signaled intra-prediction mode and data from previously decoded blocks of the current frame or picture. When a video frame is coded as an intercoded (i.e., B, P, or GPB) slice, the motion compensation unit 82 of the prediction processing unit 81 generates prediction blocks for the video blocks of the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit 80. The prediction blocks may be generated from one of the reference pictures in the reference picture list. The decoding device 112 may construct reference frame lists, i.e., List 0 and List 1, using default construction techniques based on the reference pictures stored in the picture memory 92.

[0280] The motion compensation unit 82 determines prediction information for the video blocks of the current video slice by parsing motion vectors and other syntax elements, and uses the prediction information to generate prediction blocks for the current video blocks being decoded. For example, the motion compensation unit 82 may use one or more syntax elements in the parameter set to determine the prediction mode used to code the video blocks of the video slice (e.g., intra-prediction or inter-prediction), the inter-prediction slice type (e.g., B-slice, P-slice, or GPB-slice), construction information about one or more reference picture lists for the slice, motion vectors for each intercoded video block of the slice, the inter-prediction status for each intercoded video block of the slice, and other information for decoding the video blocks in the current video slice.

[0281] The motion compensation unit 82 may perform interpolation based on an interpolation filter. The motion compensation unit 82 may calculate interpolated values ​​for sub-integer pixels of a reference block using an interpolation filter, such as the one used by the encoding device 104 during the encoding of the video block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the encoding device 104 from the received syntax elements and use that interpolation filter to generate the predicted block.

[0282] The inverse quantization unit 86 inversely quantizes or dequantizes the quantized transformation coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include determining the degree of quantization using quantization parameters calculated by the encoding device 104 for each video block in the video slice, and similarly determining the degree of inverse quantization to be applied. The inverse transformation processing unit 88 applies an inverse transformation (e.g., inverse DCT or other suitable inverse transformation), an inverse integer transformation, or a conceptually similar inverse transformation process to the transformation coefficients to generate residual blocks in the pixel region.

[0283] After the motion compensation unit 82 generates predicted blocks for the current video block based on motion vectors and other syntax elements, the decoding device 112 forms the decoded video block by adding the residual block from the inverse processing unit 88 with the corresponding predicted block generated by the motion compensation unit 82. The adder 90 represents one or more components that perform this addition operation. Loop filters (either in or after the coding loop), if desired, may also be used to smooth pixel transitions or to improve video quality in other ways. The filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample-adaptive offset (SAO) filter. The filter unit 91 is shown in Figure 8 as an in-loop filter, but in other configurations, the filter unit 91 may be implemented as a post-loop filter. The decoded video block in a given frame or picture is then stored in the picture memory 92, which stores a reference picture used for subsequent motion compensation. The picture memory 92 also stores the decoded video so that it can be later presented on a display device such as the video destination device 122 shown in Figure 1.

[0284] Thus, the decoding device 112 in Figure 13 represents an example of a video decoder configured to perform any of the techniques described herein, including the process described above with respect to Figure 11.

[0285] As used herein, the term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, storing, or transporting instructions and / or data. Computer-readable medium may also include non-transient media capable of storing data and that do not contain carrier waves and / or transient electronic signals that propagate wirelessly or via wired connections. Examples of non-transient media include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital multipurpose discs (DVDs), flash memory, memory, or memory devices. Computer-readable medium may store code and / or machine-executable instructions that may represent procedures, functions, subprograms, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuits by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, transferred, or transmitted via any appropriate means, including memory sharing, message passing, token passing, network transmission, etc.

[0286] In some embodiments, computer-readable storage devices, media, and memory may include cables or wireless signals, such as bitstreams. However, as stated, non-transient computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0287] Specific details are provided in the above description to give a complete understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that these embodiments may be practiced without these specific details. For clarity of description, in some cases the art may be presented as including individual functional blocks, which include a device, device components, and software, or steps or routines in a manner embodied in a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, or other components may be shown as components in the form of block diagrams, so as not to obscure the embodiments with unnecessary details. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details, so as not to obscure the embodiments.

[0288] Individual aspects may be described above as processes or methods represented as flowcharts, flow diagrams, data flow diagrams, structural diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many operations can be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when its operations are complete, but it may have additional steps not shown in the diagram. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to a calling function or main function.

[0289] The processes and methods described above may be carried out using computer-executable instructions stored or otherwise available from computer-readable media. Such instructions may include instructions and data that cause a general-purpose computer, a dedicated computer, or a processing device to perform a particular function or group of functions, or to otherwise configure a general-purpose computer, a dedicated computer, or a processing device to perform a particular function or group of functions. The portion of computer resources used may be accessible over a network. Computer-executable instructions may be binary or intermediate format instructions, such as assembly language, firmware, or source code. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods described above include magnetic or optical disks, flash memory, USB devices with non-volatile memory, and network-connected storage devices.

[0290] Devices that implement processes and methods in accordance with these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing the required tasks may be stored in computer-readable or machine-readable media. The processor may perform the required tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mount devices, and standalone devices. The functions described herein may also be embodied in peripheral devices or add-in cards. Such functions may also, as a further example, be implemented on circuit boards of different chips or on different processes running within a single device.

[0291] Instructions, a medium for transmitting such instructions, computing resources for executing such instructions, and other structures for supporting such computing resources are exemplary means for providing the functionality described herein.

[0292] In the above description, embodiments of this application are described with reference to those specific embodiments, but those skilled in the art will recognize that this application is not limited thereto. Therefore, while exemplary embodiments of this application are described in detail herein, it should be understood that the concepts of the invention can be embodied and utilized in various other ways, and that, unless limited by the prior art, the appended claims are intended to be interpreted as including such variations. The various features and embodiments of this application described above can be used individually or in combination. Furthermore, embodiments can be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of this specification. Therefore, this specification and the drawings should be considered illustrative, not restrictive. For illustrative purposes, the methods have been described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in a different order than described.

[0293] Those skilled in the art will understand that the symbols or terms less than ("<") and greater than (">") used herein may be replaced by the symbols less than or equal to ("≦") and greater than or equal to ("≧"), respectively, without departing from the scope of this description.

[0294] When a component is described as "configured to perform certain actions," such configuration can be achieved, for example, by designing electronics or other hardware to perform the actions, by programming programmable electronics (e.g., a microprocessor or other suitable electronics) to perform the actions, or by any combination thereof.

[0295] The phrase "connected" refers to any component that is physically connected to another component, either directly or indirectly, and / or any component that communicates with another component, either directly or indirectly (for example, connected to another component via a wired or wireless connection and / or other appropriate communication interface).

[0296] The claim language or other language in this disclosure that describes a set "at least one of" and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claim language that describes "at least one of A and B" and "at least one of A or B" means A, B, or A and B. In another example, the claim language that describes "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The phrase "at least one of" and / or "one or more" of a set does not limit the set to items enumerated in the set. For example, the wording of a claim that states "at least one of A and B" and "at least one of A or B" could mean A, B, or A and B, and in addition, could include items that are not listed in the set of A and B.

[0297] Various exemplary logic blocks, modules, circuits, and algorithmic steps described in relation to the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly demonstrate this hardware and software compatibility, various exemplary components, blocks, modules, circuits, and steps are described above in general terms of their function. Whether such functions are implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A person skilled in the art may implement the described functions in various ways for each specific application, but such a decision on implementation should not be construed as causing a departure from the scope of this application.

[0298] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication device handsets, or integrated circuit devices having multiple applications, including applications in wireless communication device handsets and other devices. Any feature described as a module or component may be implemented together in an integrated logic device, or separately as separate but interoperable logic devices. When implemented in software, the techniques may be at least partially implemented by a computer-readable data storage medium comprising program code that, when executed, performs one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise memory or data storage media such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, etc. The technique may, as an addition or alternative, be at least partially implemented by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures, such as propagating signals or waves, and which can be accessed, read, and / or executed by a computer.

[0299] The program code may be executed by a processor which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other uniform integrated or discrete logic circuits. Such processors may be configured to perform any of the techniques described herein. The general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors working with a DSP core, or any other such configuration. Accordingly, the term “processor” as used herein may refer to any of the above structures, any combination thereof, or any other structure or device suitable for implementing the techniques described herein. In addition, in some embodiments, the functions described herein may be provided in a dedicated software or hardware module configured for encoding and decoding, or incorporated into a composite video encoder / decoder (codec).

[0300] The embodiments for explaining this disclosure include the following:

[0301] Embodiment 1: An apparatus for processing video data, comprising a memory and one or more processors coupled to the memory. The one or more processors are configured to acquire a current picture of video data, acquire a reference picture for the current picture from the video data, determine a merge mode candidate from the current picture, identify a first motion vector and a second motion vector for the merge mode candidate, select a motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, use the motion vector search strategy to derive an refined motion vector from the first motion vector, the second motion vector and the reference picture, and process the merge mode candidate using the refined motion vector.

[0302] Embodiment 2: The apparatus of Embodiment 1, wherein a merge mode candidate is selected from a merge candidate list.

[0303] Embodiment 3: The apparatus of Embodiment 2, wherein the merge candidate list is constructed from one or more of the following: spatial motion vector predictors from spatially adjacent blocks of the merge mode candidate, time motion vector predictors from co-position blocks of the merge mode candidate, history-based motion vector predictors from a history table, pairwise average motion vector predictors, and zero-value motion vectors.

[0304] Embodiment 4: An apparatus according to any one of Embodiments 1 to 3, wherein one or more processors are configured to generate a motion vector dual prediction signal using a first motion vector and a second motion vector by averaging two prediction signals obtained from two different reference pictures.

[0305] Embodiment 5: An apparatus according to any one of Embodiments 1 to 4, wherein multiple motion vector search strategies include a fractional sample refinement strategy.

[0306] Embodiment 6: The apparatus of Embodiment 5, wherein the multiple motion vector search strategies include a dual-prediction optical flow strategy.

[0307] Embodiment 7: The apparatus of Embodiment 6, wherein multiple motion vector search strategies include a subblock-based bilateral matching motion vector refinement strategy.

[0308] Embodiment 8: An apparatus according to any of Embodiments 1 to 7, wherein the first motion vector and the second motion vector are associated with one or more constraints.

[0309] Embodiment 9: The apparatus of Embodiment 8, wherein one or more constraints include a mirroring constraint.

[0310] Embodiment 10: An apparatus according to any of Embodiments 1 to 9, wherein one or more constraints include a zero-value constraint for a first motion vector difference.

[0311] Embodiment 11: An apparatus according to any of Embodiments 1 to 9, wherein one or more constraints include a zero-value constraint for a second motion vector difference.

[0312] Embodiment 12: An apparatus according to any of Embodiments 1 to 11, wherein video data includes syntax indicating one or more constraints.

[0313] Embodiment 13: An apparatus according to any of Embodiments 1 to 12, wherein the motion vector search strategy comprises a multipath decoder-side motion vector refinement strategy.

[0314] Embodiment 14: The apparatus of Embodiment 13, wherein the multipath decoder-side motion vector refinement strategy includes two or more refinement paths of the same refinement type.

[0315] Embodiment 15: The apparatus of Embodiment 14, wherein the multipath decoder-side motion vector refinement strategy includes one or more refinement paths of different types from the same refinement type.

[0316] Embodiment 16: The apparatus of any embodiment 14 or 15, wherein two or more refinement paths of the same refinement type are block-based bilateral matching motion vector refinement, subblock-based bilateral matching motion vector refinement, or subblock-based bidirectional optical flow motion vector refinement.

[0317] Embodiment 17: An apparatus according to any of Embodiments 1 to 16, wherein the multiple motion vector search strategies comprise multiple subsets of multipath strategies.

[0318] Embodiment 18: The apparatus of Embodiment 17, wherein multiple subsets of a multipath strategy are signaled in one or more syntax elements of video data.

[0319] Embodiment 19: An apparatus according to any one of embodiments 1 to 18, comprising deriving refined motion vectors and calculating matching costs for multiple candidate motion vector pairs according to a motion vector search strategy.

[0320] Embodiment 20: An apparatus according to any of Embodiments 1 to 19, wherein the motion vector search strategy is adaptively selected based on the matching cost determined from the video data.

[0321] Embodiment 21: An apparatus of any embodiment 1 to 19, wherein a motion vector search strategy is selected to adaptively set the number of paths of the motion vector search strategy based on video data.

[0322] Embodiment 22: An apparatus of any embodiment 1 to 19, wherein the motion vector search strategy is selected to adaptively set a search pattern for determining candidates for refined motion vectors based on video data.

[0323] Embodiment 23: An apparatus of any embodiment 1 to 19, wherein the motion vector search strategy is selected to adaptively set a set of criteria for generating a list of candidates for refined motion vectors based on video data.

[0324] Embodiment 24: An apparatus according to any of Embodiments 1 to 19, wherein the motion vector search strategy is adaptively performed from video data based on decoder-side motion vector refinement constraints.

[0325] Embodiment 25: An apparatus according to any embodiment 1 to 19, wherein the motion vector search strategy is performed adaptively based on the block size for merge mode candidates in the video data.

[0326] Embodiment 26: An apparatus of any embodiment 1 to 25, wherein one or more processors are configured to disable multiple hypothetical predictions.

[0327] Embodiment 27: An apparatus according to any one of Embodiments 1 to 26, wherein one or more processors are configured to perform a plurality of hypothetical predictions together with a motion vector search strategy.

[0328] Embodiment 28: An apparatus according to any one of Embodiments 1 to 27, wherein one or more processors are configured to generate a merge candidate list that includes merge mode candidates.

[0329] Embodiment 29: The apparatus of Embodiment 28, wherein one or more processors are configured to determine one or more default candidates to add to the merge candidate list based on one or more conditions related to an adaptive merge mode (for example, a condition for adaptive_bm_mode) and on the basis that the number of candidates in the merge candidate list is less than the maximum number of candidates.

[0330] Embodiment 30: The apparatus of Embodiment 28, wherein one or more processors are configured to determine one or more default candidates to add to the merge candidate list based on one or more conditions (e.g., a condition according to bm_dir) relating to constraints relating to an adaptive merge mode, on the basis that the number of candidates in the merge candidate list is less than the maximum number of candidates.

[0331] Embodiment 31: An apparatus according to any of Embodiments 1 to 30, wherein the apparatus is a mobile device.

[0332] Embodiment 32: An apparatus according to any embodiment 1 to 31, further comprising a camera configured to capture one or more frames.

[0333] Embodiment 33: An apparatus of any embodiment 1 to 32, further comprising a display configured to display one or more frames.

[0334] Embodiment 34: A method for processing video data according to any of the operations described in Embodiments 1 to 33.

[0335] Embodiment 35: A computer-readable storage medium comprising instructions, when executed by one or more processors of the device, that cause the device to perform any of the operations described in Embodiments 1 to 33.

[0336] Embodiment 36: An apparatus comprising one or more means for performing any of the operations of Embodiments 1 to 33.

[0337] Embodiment 37: An apparatus for processing video data, comprising at least one memory and at least one processor coupled to the at least one memory, wherein the at least one processor is configured to acquire one or more reference pictures for a current picture, identify a first motion vector and a second motion vector for a merge mode candidate, determine a selected motion vector search strategy for a merge mode candidate from a plurality of motion vector search strategies, determine one or more refined motion vectors using the selected motion vector search strategy based on at least one of the first motion vectors or the second motion vector and one or more reference pictures, and process the merge mode candidate using one or more refined motion vectors.

[0338] Embodiment 38. The apparatus of Embodiment 37, wherein the selected motion vector search strategy is associated with one or more constraints based on at least one of the first motion vectors or the second motion vectors.

[0339] Embodiment 39. Apparatus of Embodiment 38, wherein one or more constraints are determined for a block of video data based on syntax elements signaled for the block.

[0340] Embodiment 40. An apparatus according to any one of Embodiments 38 or 39, wherein one or more constraints are associated with at least one of a first motion vector difference associated with a first motion vector or a second motion vector difference associated with a second motion vector.

[0341] Embodiment 41. An apparatus of Embodiment 40, wherein one or more refined motion vectors include a first refined motion vector and a second refined motion vector, and at least one processor is configured to determine a first motion vector difference as the difference between the first refined motion vector and the first motion vector, and to determine a second motion vector difference as the difference between the second refined motion vector and the second motion vector.

[0342] Embodiment 42. An apparatus according to either Embodiment 40 or 41, wherein one or more constraints include mirroring constraints for a first motion vector difference and a second motion vector difference, and the first motion vector difference and the second motion vector difference have the same magnitude and different signs.

[0343] Embodiment 43. An apparatus according to any of embodiments 40 to 42, wherein one or more constraints include a zero-value constraint for at least one of a first motion vector difference or a second motion vector difference.

[0344] Embodiment 44. The apparatus of Embodiment 43, configured such that, based on a zero-value constraint, at least one processor determines one or more refined motion vectors using a selected motion vector search strategy by keeping a first of the first or second motion vector differences constant and searching on the first or second of the second motion vector differences.

[0345] Embodiment 45. An apparatus according to any of Embodiments 37 to 44, wherein the selected motion vector search strategy is a bilateral matching (BM) motion vector search strategy.

[0346] Embodiment 46. An apparatus of any embodiment 37 to 45, wherein at least one processor is configured to determine one or more refined motion vectors based on one or more constraints relating to a selected motion vector search strategy, and in order to determine one or more refined motion vectors based on one or more constraints, at least one processor is configured to determine a first refined motion vector by searching a first reference picture around a first motion vector based on a selected motion vector search strategy, and to determine a second refined motion vector by searching a second reference picture around a second motion vector based on a selected motion vector search strategy, wherein one or more constraints include motion vector difference constraints.

[0347] Embodiment 47. The apparatus of Embodiment 46, wherein at least one processor is configured to minimize the difference between a first reference block associated with the first refined motion vector and a second reference block associated with the second refined motion vector in order to determine a first refined motion vector and a second refined motion vector.

[0348] Embodiment 48. An apparatus according to any one of Embodiments 37 to 47, wherein the multiple motion vector search strategies include at least two of the following: a multipath decoder-side motion vector refinement strategy, a fractional sample refinement strategy, a bidirectional optical flow strategy, or a subblock-based bilateral matching motion vector refinement strategy.

[0349] Embodiment 49. An apparatus according to any of Embodiments 37 to 48, wherein the selected motion vector search strategy comprises a multipath decoder-side motion vector refinement strategy.

[0350] Embodiment 50. Apparatus of Embodiment 49, wherein the multipath decoder-side motion vector refinement strategy includes at least one of one or more block-based bilateral matching motion vector refinement paths or one or more subblock-based motion vector refinement paths.

[0351] Embodiment 51. The apparatus of Embodiment 50, wherein at least one processor is configured to perform one or more block-based bilateral matching motion vector refinement paths using a first constraint associated with at least one of a first motion vector difference or a second motion vector difference, and to perform one or more subblock-based motion vector refinement paths using a second constraint associated with at least one of the first motion vector difference or a second motion vector difference, wherein the first constraint is different from the second constraint.

[0352] Embodiment 52. An apparatus according to any one of Embodiments 50 or 51, wherein one or more subblock-based motion vector refinement paths include at least one of a subblock-based bilateral matching motion vector refinement path or a subblock-based bidirectional optical flow motion vector refinement path.

[0353] Embodiment 53. An apparatus according to any of Embodiments 37 to 52, wherein the apparatus is a wireless communication device.

[0354] Embodiment 54. An apparatus according to any of embodiments 37 to 53, wherein at least one processor is configured to determine one or more refined motion vectors for a block of video data, and the merge mode candidates include adjacent blocks of the block.

[0355] Embodiment 55: A method for processing video data, comprising: obtaining one or more reference pictures for a current picture; identifying a first motion vector and a second motion vector for a merge mode candidate; determining a selected motion vector search strategy for a merge mode candidate from a plurality of motion vector search strategies; determining one or more refined motion vectors using the selected motion vector search strategy, based on at least one of the first motion vectors or the second motion vector and one or more reference pictures; and processing the merge mode candidate using the one or more refined motion vectors.

[0356] Embodiment 56. The method of Embodiment 55, wherein the selected motion vector search strategy is associated with one or more constraints based on at least one of the first motion vectors or the second motion vectors.

[0357] Embodiment 57. The method of Embodiment 56, wherein one or more constraints are determined for a block of video data based on syntax elements signaled for the block.

[0358] Embodiment 58. A method according to either Embodiment 56 or 57, wherein one or more constraints are associated with at least one of a first motion vector difference related to a first motion vector or a second motion vector difference related to a second motion vector.

[0359] Embodiment 59. The method of Embodiment 58, wherein one or more refined motion vectors include a first refined motion vector and a second refined motion vector, and further comprising the steps of determining a first motion vector difference as the difference between the first refined motion vector and the first motion vector, and determining a second motion vector difference as the difference between the second refined motion vector and the second motion vector.

[0360] Embodiment 60.1 or more constraints include mirroring constraints for a first motion vector difference and a second motion vector difference, wherein the first motion vector difference and the second motion vector difference have the same magnitude and different signs, according to either Embodiment 58 or 59.

[0361] Embodiment 61. Any method of Embodiments 58 to 60, wherein one or more constraints include a zero-value constraint for at least one of a first motion vector difference or a second motion vector difference.

[0362] Embodiment 62. The method of Embodiment 61, wherein, based on a zero-value constraint, one or more motion-refined motion vectors are determined using a selected motion vector search strategy by keeping a first of the first or second motion vector differences constant and searching against the second of the first or second motion vector differences.

[0363] Embodiment 63. Any method of Embodiments 55 to 62, wherein the selected motion vector search strategy is a bilateral matching (BM) motion vector search strategy.

[0364] Embodiment 64. A method of any embodiment 55 to 63, wherein one or more refined motion vectors are determined based on one or more constraints relating to a selected motion vector search strategy, and the step of determining one or more refined motion vectors based on one or more constraints comprises the steps of determining a first refined motion vector by searching for a first reference picture around a first motion vector based on a selected motion vector search strategy, and determining a second refined motion vector by searching for a second reference picture around a second motion vector based on a selected motion vector search strategy, wherein one or more constraints include motion vector difference constraints.

[0365] Embodiment 65. The method of Embodiment 64, wherein the step of determining a first refined motion vector and a second refined motion vector comprises the step of minimizing the difference between a first reference block associated with the first refined motion vector and a second reference block associated with the second refined motion vector.

[0366] Embodiment 66. Any method of Embodiments 55 to 65, wherein the multiple motion vector search strategies include at least two of the following: a multipath decoder-side motion vector refinement strategy, a fractional sample refinement strategy, a bidirectional optical flow strategy, or a subblock-based bilateral matching motion vector refinement strategy.

[0367] Embodiment 67. Any method from Embodiments 55 to 66, wherein the selected motion vector search strategy comprises a multipath decoder-side motion vector refinement strategy.

[0368] Embodiment 68. The method of Embodiment 67, wherein the multipath decoder-side motion vector refinement strategy includes at least one of one or more block-based bilateral matching motion vector refinement paths or one or more subblock-based motion vector refinement paths.

[0369] Embodiment 69. The method of Embodiment 68, further comprising the steps of: performing one or more block-based bilateral matching motion vector refinement paths using a first constraint associated with at least one of a first motion vector difference or a second motion vector difference; and performing one or more subblock-based motion vector refinement paths using a second constraint associated with at least one of a first motion vector difference or a second motion vector difference, wherein the first constraint is different from the second constraint.

[0370] Embodiment 70. The method according to any one embodiment 68 or 69, wherein one or more subblock-based motion vector refinement paths include at least one of a subblock-based bilateral matching motion vector refinement path or a subblock-based bidirectional optical flow motion vector refinement path.

[0371] Embodiment 71: A method for processing video data according to any of the operations described in Embodiments 37 to 70.

[0372] Embodiment 72: A computer-readable storage medium comprising instructions, when executed by one or more processors of the device, that cause the device to perform any of the operations of Embodiments 37 to 70.

[0373] Embodiment 73: An apparatus comprising one or more means for performing any of the operations of Embodiments 37 to 70. [Explanation of Symbols]

[0374] 35 division units 41 Prediction Processing Unit 42 Motion Estimation Unit 44 Motion compensation unit 46 Intra Prediction Processing Unit 50 Adder 52 Conversion Processing Unit 54 Quantization Units 56 Entropy coding unit 57 Post-processing devices 58 Inverse Quantization Unit 60 Inverse Transform Processing Unit 62 Adder 63 filter units 64 Picture Memory 79 Network Entities 80 Entropy Decoding Unit 81 Prediction Processing Unit 82 Motion Compensation Unit 84 Intra Prediction Processing Unit 86 Inverse Quantization Unit 88 Inverse Transformation Processing Unit 90 Adder 91 Filter Unit 92 Picture Memory 102 Video Sources 104 Encoding devices, video encoding devices 106 Encoder Engine 108 storage 110 Output 112 Decryption Device 114 inputs 116 Decoder Engine 118 storage 120 Communication Links 122 Video Destination Devices 402 Currently blocked 404 Reference Block 422 Currently blocked 424 First reference block 426 Second reference block 500 processing blocks, blocks 610 Current Picture 612 Current CU 615 Current reference picture 630 Same-position picture 632 Same position CU 635 Same-location reference picture 810 Current Picture 812 Dual Prediction Merge Candidates 822 Second predictor, initial predictor 824 Refined candidate blocks 830 First reference picture 832 First predictor, initial predictor 834 Refined candidate blocks 910 Subblock 970 lines 980 columns 1020 First search area 1030 Second search area 1040 Third Exploration Area 1050 Fourth Exploration Area

Claims

1. A device for processing video data, At least one memory, The system comprises at least one processor coupled to the at least one memory, and the at least one processor is Currently, we are obtaining one or more reference pictures for the picture, Identifying a first motion vector and a second motion vector for merge mode candidates, Determining a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, wherein the selected motion vector search strategy is associated with one or more constraints based on at least one of the first motion vector or the second motion vector, and the one or more constraints are associated with at least one of the first motion vector difference related to the first motion vector or the second motion vector difference related to the second motion vector, and the one or more constraints include a zero-value constraint for at least one of the first motion vector difference or the second motion vector difference. Determining one or more refined motion vectors based on at least one of the first motion vectors or the second motion vectors and one or more reference pictures using the selected motion vector search strategy, wherein determining the one or more refined motion vectors using the selected motion vector search strategy is based on the zero-value constraint, and determining the one or more refined motion vectors using the selected motion vector search strategy includes keeping the first of the first motion vector differences or the second motion vector differences at a constant value and searching for the second of the first motion vector differences or the second motion vector differences. Processing the merge mode candidates using one or more refined motion vectors. A device configured to perform the following actions.

2. The apparatus according to claim 1, wherein the one or more constraints are determined for the block of video data based on syntax elements signaled for the block.

3. The one or more refined motion vectors include a first refined motion vector and a second refined motion vector, and the at least one processor, The first motion vector difference is determined as the difference between the first refined motion vector and the first motion vector. The apparatus according to claim 1, configured to determine the second motion vector difference as the difference between the second refined motion vector and the second motion vector.

4. The apparatus according to claim 1, wherein the selected motion vector search strategy is a bilateral matching (BM) motion vector search strategy.

5. The at least one processor is configured to determine the one or more refined motion vectors based on one or more constraints related to the selected motion vector search strategy, and in order to determine the one or more refined motion vectors based on one or more constraints, the at least one processor is configured The first refined motion vector is determined by searching for the first reference picture around the first motion vector based on the selected motion vector search strategy. The system is configured to determine a second refined motion vector by searching for a second reference picture around the second motion vector based on the selected motion vector search strategy, wherein the one or more constraints include motion vector difference constraints. Optionally, in order to determine the first refined motion vector and the second refined motion vector, the at least one processor may It is configured to minimize the difference between the first reference block associated with the first refined motion vector and the second reference block associated with the second refined motion vector. The apparatus according to claim 4.

6. The apparatus according to claim 1, wherein the plurality of motion vector search strategies include at least two of the following: a multipath decoder-side motion vector refinement strategy, a fractional sample refinement strategy, a bidirectional optical flow strategy, or a subblock-based bilateral matching motion vector refinement strategy.

7. The apparatus according to claim 1, wherein the selected motion vector search strategy comprises a multipath decoder-side motion vector refinement strategy.

8. The multipath decoder-side motion vector refinement strategy includes at least one of one or more block-based bilateral matching motion vector refinement paths or one or more subblock-based motion vector refinement paths. Optionally, at least one of the processors, Performing one or more block-based bilateral matching motion vector refinement paths using a first constraint associated with at least one of the first motion vector difference or the second motion vector difference, The system is configured to perform the one or more subblock-based motion vector refinement paths using a second constraint associated with at least one of the first motion vector difference or the second motion vector difference, and / or Optionally, the one or more subblock-based motion vector refinement paths include at least one of the following: a subblock-based bilateral matching motion vector refinement path or a subblock-based bidirectional optical flow motion vector refinement path. The apparatus according to claim 7.

9. A wireless communication device, and / or the at least one processor is configured to determine the one or more refined motion vectors for the block of video data, wherein the merge mode candidate includes adjacent blocks of the block. The apparatus according to claim 1.

10. A method for processing video data, The steps include obtaining one or more reference pictures for the current picture, Steps include identifying a first motion vector and a second motion vector for merge mode candidates, A step of determining a selected motion vector search strategy for the merge mode candidate from a plurality of motion vector search strategies, wherein the selected motion vector search strategy is associated with one or more constraints based on at least one of the first motion vector or the second motion vector, and the one or more constraints are associated with at least one of the first motion vector difference related to the first motion vector or the second motion vector difference related to the second motion vector, and the one or more constraints include a zero-value constraint for at least one of the first motion vector difference or the second motion vector difference, A step of determining one or more refined motion vectors based on at least one of the first motion vectors or the second motion vectors and one or more reference pictures, wherein the step of determining the one or more refined motion vectors using the selected motion vector search strategy includes, based on the zero-value constraint, a step of keeping the first of the first motion vector differences or the second motion vector differences constant, and a step of searching for the second of the first motion vector differences or the second motion vector differences, A method comprising the step of processing the merge mode candidates using one or more refined motion vectors.

11. The method according to claim 10, wherein the one or more constraints are determined for the block of video data based on syntax elements signaled for the block.

12. The one or more refined motion vectors include a first refined motion vector and a second refined motion vector, and the method further The steps include determining the first motion vector difference as the difference between the first refined motion vector and the first motion vector, The method according to claim 10, comprising the step of determining the second motion vector difference as the difference between the second refined motion vector and the second motion vector.

13. The selected motion vector search strategy is a bilateral matching (BM) motion vector search strategy, and the one or more refined motion vectors are determined based on one or more constraints related to the selected motion vector search strategy, and the step of determining the one or more refined motion vectors based on one or more constraints is: The steps include determining a first refined motion vector by searching for a first reference picture around the first motion vector based on the selected motion vector search strategy, The process comprises the steps of determining a second refined motion vector by searching for a second reference picture around the second motion vector based on the selected motion vector search strategy, The one or more constraints include motion vector difference constraints, Optionally, the step of determining the first refined motion vector and the second refined motion vector is: The process includes the step of minimizing the difference between a first reference block associated with the first refined motion vector and a second reference block associated with the second refined motion vector. The method according to claim 10.

14. The selected motion vector search strategy comprises a multipath decoder-side motion vector refinement strategy, wherein the multipath decoder-side motion vector refinement strategy includes at least one of one or more block-based bilateral matching motion vector refinement paths or one or more subblock-based motion vector refinement paths, optionally, The steps of performing one or more block-based bilateral matching motion vector refinement paths using a first constraint associated with at least one of a first motion vector difference or a second motion vector difference, The steps include: performing the one or more subblock-based motion vector refinement paths using a second constraint associated with at least one of the first motion vector difference or the second motion vector difference. The method according to claim 10.

15. A computer-readable storage medium including instructions, wherein when the instructions are executed by one or more processors of the device, the device causes the device to perform the operation described in any one of claims 10 to 14.

Citation Information

Patent Citations

  • Image decoding device, image encoding device, image processing system, and program

    JP2020053723A

  • Symmetric motion vector differential coding

    JP2022531554A

  • Inter prediction methods for coding video data

    US20200296416A1

  • Symmetric motion vector difference coding

    WO2020221258A1