Improved motion vector on the multipath decoder side

A multi-pass decoder-side motion vector refinement technique addresses the narrow scope issue in existing standards, enhancing motion prediction and decoding accuracy in video coding.

JP7832938B2Active Publication Date: 2026-03-18QUALCOMM INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2026-03-18

AI Technical Summary

Technical Problem

Existing video coding standards have a narrow scope for motion vector improvement, leading to inaccurate motion prediction and decoding of encoded video data.

Method used

A multi-pass decoder-side motion vector refinement technique is applied, involving a block-based, subblock-based, and further subblock-based passes to improve motion vectors, enhancing accuracy in motion prediction and decoding.

Benefits of technology

This approach results in more accurate motion prediction and decoding of encoded video data, improving the reproduction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007832938000020
    Figure 0007832938000020
  • Figure 0007832938000021
    Figure 0007832938000021
  • Figure 0007832938000022
    Figure 0007832938000022
Patent Text Reader

Abstract

[0003] Exemplary devices and techniques for multi-pass decoder-side motion vector refinement (DMVR) are disclosed. The exemplary device includes a memory configured to store video data and one or more processors coupled to the memory. The one or more processors are configured to apply multi-pass DMVR to motion vectors for blocks of the video data to determine at least one refined motion vector, and decode the blocks based on the at least one refined motion vector. The multi-pass DMVR includes a block-based first pass, a sub-block-based second pass, and a sub-block-based third pass.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Patent Application No. 17 / 556,142, filed on 20 December 2021, and U.S. Provisional Application No. 63 / 129,221, titled "MULTI-PASS DECODER-SIDE MOTION VECTOR REFINEMENT," filed on 22 December 2020, the entire contents of these applications being incorporated herein by reference. U.S. Patent Application No. 17 / 556,142, filed on 20 December 2021, claims the benefit of U.S. Provisional Application No. 63 / 129,221, filed on 22 December 2020.

[0002] This disclosure relates to video coding and video decoding. [Background technology]

[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital television, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radio phones, so-called "smartphones," video teleconferencing devices, and video streaming devices. Digital video devices implement video coding techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), and extensions of such standards. By implementing such video coding techniques, video devices can transmit, receive, encode, decode, and / or store digital video information more efficiently.

[0004] Video coding techniques include spatial (intra-picture) and / or temporal (inter-picture) predictions to reduce or eliminate redundancy inherent in video sequences. In block-based video coding, a video slice (e.g., a video picture, or a portion of a video picture) may be divided into video blocks, which may also be called coding tree units (CTUs), coding units (CUs), and / or coding nodes. A video block in an intra-coded (I) slice of a picture is coded using spatial predictions for reference samples in adjacent blocks within the same picture. A video block in an inter-coded (P or B) slice of a picture may use spatial predictions for reference samples in adjacent blocks within the same picture or temporal predictions for reference samples in other reference pictures. A picture may be called a frame, and a reference picture may be called a reference frame. [Overview of the project] [Problems that the invention aims to solve]

[0005] In general, this disclosure describes techniques for decoder-side motion vector derivation techniques. More specifically, this disclosure describes a multi-pass decoder-side motion vector improvement technique for use in video coding. In some draft video standards, the scope of motion vector improvement may be too narrow for all cases. The technique of this disclosure addresses this problem, which can result in more accurate motion prediction and therefore more accurate decoding and reproduction of encoded video data. [Means for solving the problem]

[0006] In one example, the method includes the steps of applying a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector, and decoding the block based on at least one improved motion vector, wherein the multipath DMVR comprises a first pass that is block-based and applied to a block of video data; a second pass that is subblock-based and applied to at least one second pass subblock of the block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass that is subblock-based and applied to at least one third pass subblock of the block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0007] In another example, the device includes a memory configured to store video data and one or more processors implemented by circuitry and communicatively coupled to the memory, the one or more processors configured to apply a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector, and to decode the block based on the at least one improved motion vector, the multipath DMVR comprising: a first pass which is block-based and applied to a block of video data; a second pass which is subblock-based and applied to at least one second pass subblock of a block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass which is subblock-based and applied to at least one third pass subblock of a block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0008] In another example, a non-temporary computer-readable medium stores instructions that, when executed, cause one or more processors to apply a multipath decoder-side motion vector improvement (DMVR) to the motion vectors for a block of video data to determine at least one improved motion vector, and to decode the block based on the at least one improved motion vector, wherein the multipath DMVR comprises a first pass that is block-based and applied to a block of video data; a second pass that is subblock-based and applied to at least one second pass subblock of a block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass that is subblock-based and applied to at least one third pass subblock of a block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0009] In another example, the device includes means for applying a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector, and means for decoding the block based on at least one improved motion vector, wherein the multipath DMVR comprises a first pass that is block-based and applied to a block of video data; a second pass that is subblock-based and applied to at least one second pass subblock of the block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass that is subblock-based and applied to at least one third pass subblock of the block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0010] In one example, the method includes the steps of applying a multipath decoder-side motion vector improvement (DMVR) to the motion vector for a block of video data to determine an improved motion vector, and coding the block based on the improved motion vector.

[0011] In another example, the device includes memory configured to store video data and one or more processors that are circuit-implemented and communicatively coupled to the memory, the one or more processors configured to perform any of the techniques of the present disclosure.

[0012] In another example, the device includes at least one means for performing any of the techniques of the present disclosure.

[0013] In another example, a computer-readable storage medium encodes instructions that, when executed, cause a programmable processor to perform one of the techniques of this disclosure.

[0014] Details of one or more examples are described in the accompanying drawings and the following description. Other features, purposes, and advantages will become apparent from the description, drawings, and claims. [Brief explanation of the drawing]

[0015] [Figure 1] Block diagram shows an exemplary video encoding and decoding system capable of performing the techniques of this disclosure. [Figure 2A] This is a conceptual diagram illustrating an exemplary quadrow-binary tree (QTBT) structure. [Figure 2B] This is a conceptual diagram showing the corresponding coding tree unit (CTU). [Figure 3] A block diagram illustrating a video encoder capable of performing the techniques of this disclosure. [Figure 4] A block diagram illustrating an exemplary video decoder capable of performing the techniques of this disclosure. [Figure 5A] This is a conceptual diagram showing exemplary spatially adjacent MV candidates for merge mode. [Figure 5B] This is a conceptual diagram showing exemplary spatially adjacent MV candidates for AMVP mode. [Figure 6A] This is a conceptual diagram showing an example of a TMVP candidate. [Figure 6B] This is a conceptual diagram illustrating exemplary MV scaling. [Figure 7] This is a conceptual diagram illustrating exemplary template matching in the search area around the initial MV. [Figure 8A] This is a conceptual diagram illustrating an example where MVD0 and MVD1 are proportional based on time distance. [Figure 8B] This is a conceptual diagram illustrating an example where MVD0 and MVD1 are mirror images of each other regardless of time and distance. [Figure 9] This is a conceptual diagram showing an example of a 3x3 square search pattern within the search range [-8,8]. [Figure 10] This is a conceptual diagram illustrating an example of decoder-side motion vector improvement. [Figure 11] This is a conceptual diagram illustrating an example of an extended CU region used in BDOF. [Figure 12] This is a conceptual diagram illustrating an exemplary 3-pass DMVR technique. [Figure 13] This is a conceptual diagram illustrating an example of BDOF motion vector improvement. [Figure 14] This flowchart shows an exemplary multi-pass DMVR technique in this disclosure. [Figure 15] This flowchart shows an exemplary method for encoding a current block using the technique of the present disclosure. [Figure 16] This flowchart shows an exemplary method for decrypting a current block using the technique of the present disclosure. [Modes for carrying out the invention]

[0016] In some draft video standards, the scope of motion vector improvement may be too narrow for all cases. This can lead to erroneous motion prediction and therefore less accurate decoding. The technique of this disclosure addresses this problem and can result in more accurate motion prediction and therefore more accurate decoding and reproduction of encoded video data.

[0017] Figure 1 is a block diagram showing an exemplary video coding and decoding system 100 capable of performing the techniques of the present disclosure. The techniques of the present disclosure generally concern coding (encoding and / or decoding) video data. Generally, video data includes any data for processing video. Thus, video data may include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata such as signaling data.

[0018] As shown in Figure 1, system 100 includes, in this example, a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. Specifically, the source device 102 provides video data to the destination device 116 via a computer-readable medium 110. The source device 102 and destination device 116 may comprise any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, mobile devices, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, broadcast receiver devices, and the like. In some cases, the source device 102 and destination device 116 may be compatible with wireless communication and may therefore be referred to as wireless communication devices.

[0019] In the example in Figure 1, the source device 102 includes a video source 104, memory 106, a video encoder 200, and an output interface 108. The destination device 116 includes an input interface 122, a video decoder 300, memory 120, and a display device 118. According to this disclosure, the video encoder 200 of the source device 102 and the video decoder 300 of the destination device 116 may be configured to apply techniques for decoder-side motion vector derivation. Thus, the source device 102 represents an example of a video encoding device, while the destination device 116 represents an example of a video decoding device. In other examples, the source device and destination device may include other components or configurations. For example, the source device 102 may receive video data from an external video source, such as an external camera. Similarly, the destination device 116 may interface with an external display device instead of including an integrated display device.

[0020] System 100, as shown in Figure 1, is merely an example. In general, any digital video decoding device may perform techniques for decoder-side motion vector derivation. Source device 102 and destination device 116 are merely examples of coding devices, such that source device 102 generates video data coded for transmission to destination device 116. This disclosure refers to a device that performs coding (encoding and / or decoding) of data as a “coding” device. Thus, video encoder 200 and video decoder 300 represent examples of coding devices, specifically video encoder and video decoder, respectively. In some examples, source device 102 and destination device 116 may operate substantially symmetrically, such that each of source device 102 and destination device 116 includes video encoding and decoding components. Thus, system 100 may support one-way or two-way video transmission between source device 102 and destination device 116 for, for example, video streaming, video playback, video broadcasting, or video phone.

[0021] Generally, the video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides the video encoder 200 with a sequence of pictures (also called "frames") of video data, which the video encoder 200 encodes the data for the pictures. The video source 104 of source device 102 may include video capture devices such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 encodes the captured, pre-captured, or computer-generated video data. The video encoder 200 may rearrange the pictures from the order in which they were received (sometimes called the "display order") to the coding order for encoding. The video encoder 200 may generate a bitstream containing the encoded video data. The source device 102 may then output encoded video data to a computer-readable medium 110 via the output interface 108 for reception and / or retrieval, for example, via the input interface 122 of the destination device 116.

[0022] Memory 106 of source device 102 and memory 120 of destination device 116 represent general-purpose memory. In some examples, memories 106 and 120 may store raw video data, for example, raw video from video source 104 and raw decoded video data from video decoder 300. Additionally or alternatively, memories 106 and 120 may store, for example, software instructions executable by video encoder 200 and video decoder 300, respectively. Although memories 106 and 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memories 106 and 120 may store encoded video data, for example, output from video encoder 200 and input to video decoder 300. In some examples, portions of memory 106, 120 may be allocated as one or more video buffers for storing, for example, raw, decoded, and / or encoded video data.

[0023] The computer-readable medium 110 may represent any type of medium or device capable of transferring encoded video data from the source device 102 to the destination device 116. For example, the computer-readable medium 110 may represent a communication medium that enables the source device 102 to directly transmit encoded video data to the destination device 116 in real time, for example, over a radio frequency network or a computer-based network. The output interface 108 may modulate the transmission signal containing the encoded video data, and the input interface 122 may demodulate the received transmission signal according to a communication standard such as a wireless communication protocol. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from the source device 102 to the destination device 116.

[0024] In some examples, the source device 102 may output encoded data to the storage device 112 via the output interface 108. Similarly, the destination device 116 may access encoded data from the storage device 112 via the input interface 122. The storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0025] In some examples, the source device 102 may output the encoded video data to a file server 114 or another intermediate storage device capable of storing the encoded video data generated by the source device 102. The destination device 116 may access the stored video data from the file server 114 via streaming or download.

[0026] The file server 114 can be any type of server device capable of storing encoded video data and transmitting that encoded video data to the destination device 116. The file server 114 may represent a web server (for example, for a website), a server configured to provide file transfer protocol services (such as the File Transfer Protocol (FTP) or File Delivery over Unidirectional Transport (FLUTE) protocol), a Content Delivery Network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Service (MBMS) or Enhanced MBMS (eMBMS) server, and / or a Network Attached Storage (NAS) device. The file server 114 may, in addition or alternatively, implement one or more HTTP streaming protocols such as Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), Real Time Streaming Protocol (RTSP), or HTTP Dynamic Streaming.

[0027] The destination device 116 may access encoded video data from the file server 114 through any standard data connection, including an internet connection. This may include wireless channels (e.g., Wi-Fi connection), wired connections (e.g., digital subscriber line (DSL), cable modem, etc.), or a combination of both suitable for accessing encoded video data stored on the file server 114. The input interface 122 may be configured to operate according to one or more of the various protocols described above for retrieving or receiving media data from the file server 114, or other such protocols for retrieving media data.

[0028] The output interface 108 and input interface 122 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and input interface 122 include wireless components, the output interface 108 and input interface 122 may be configured to transfer data such as encoded video data according to cellular communication standards such as 4G, 4G-LTE (Long-Term Evolution), LTE Advanced, or 5G. In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and input interface 122 may be configured to transfer data such as encoded video data according to other wireless standards such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee®), or the Bluetooth® standard. In some examples, the source device 102 and / or destination device 116 may include their respective system-on-chip (SoC) devices. For example, the source device 102 may include an SoC device for performing functions related to the video encoder 200 and / or the output interface 108, and the destination device 116 may include an SoC device for performing functions related to the video decoder 300 and / or the input interface 122.

[0029] The techniques of this disclosure can be applied to video coding that supports any of a variety of multimedia applications, such as television broadcasting by radio waves, cable television transmission, satellite television transmission, internet streaming video transmission such as dynamic adaptive streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0030] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200, which is also used by the video decoder 300, such as syntax elements having values ​​that describe the characteristics and / or processing of video blocks or other coded units (e.g., slices, pictures, picture groups, sequences, etc.). The display device 118 displays the decoded picture of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0031] Although not shown in Figure 1, in some examples the video encoder 200 and video decoder 300 may each be integrated with an audio encoder and / or audio decoder, and may include a suitable MUX-DEMUX unit or other hardware and / or software to handle multiplexed streams containing both audio and video in a common data stream. Where applicable, the MUX-DEMUX unit may conform to the ITU H.223 Multiplexer Protocol or other protocols such as the User Datagram Protocol (UDP).

[0032] The video encoder 200 and video decoder 300 can each be implemented as one or more suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technique is partially implemented in software, the device may store instructions for the software in a suitable non-temporary computer-readable medium and execute those instructions in hardware using one or more processors to perform the technique of this disclosure. Each of the video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, any of which may be integrated as part of a composite encoder / decoder (codec) within each device. A device including the video encoder 200 and / or video decoder 300 may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a mobile phone.

[0033] The video encoder 200 and video decoder 300 may operate in accordance with video coding standards such as ITU-T H.265, also known as High Efficiency Video Coding (HEVC), or extensions thereof such as multiview and / or scalable video coding extensions. Alternatively, the video encoder 200 and video decoder 300 may operate in accordance with other proprietary or industry standards, such as ITU-T H.266, also known as Versatile Video Coding (VVC). The draft of the VVC standard is described in Bross et al., "Versatile Video Coding Editorial Refinements on Draft 10," ITU-T SG 16 WP 3 and the Joint Video Experts Team (JVET) of ISO / IEC JTC 1 / SC 29 / WG 11, 18th meeting held remotely, October 7-16, 2020, JVET-T2001-v1 (hereinafter "VVC Draft 10"). However, the techniques described herein are not limited to any specific coding standard.

[0034] Generally, the video encoder 200 and video decoder 300 may perform block-based coding of pictures. The term “block” generally refers to a structure containing data to be processed (e.g., to be coded, coded, or otherwise used in the coding and / or decoding process). For example, a block may contain a two-dimensional matrix of samples of luminance and / or chromaticity data. Generally, the video encoder 200 and video decoder 300 may code video data represented in YUV (e.g., Y, Cb, Cr) format. That is, rather than coding red, green, and blue (RGB) data for samples of a picture, the video encoder 200 and video decoder 300 may code luminance and chromaticity components, and the chromaticity component may include both red and blue chromaticity components. In some examples, the video encoder 200 converts the received RGB-formatted data to a YUV representation before coding, and the video decoder 300 converts the YUV representation to RGB format. Alternatively, pre-processing and post-processing units (not shown) may perform these conversions.

[0035] This disclosure may generally refer to coding a picture (e.g., encoding and decoding) as including the process of encoding or decoding the data of the picture. Similarly, this disclosure may refer to coding a block of a picture as including the process of encoding or decoding the data for the block, e.g., predictive and / or residual coding. An encoded video bitstream generally contains a set of values ​​for syntax elements that represent the coding decision (e.g., coding mode) and the division of the picture into blocks. Thus, references to coding a picture or a block should generally be understood as coding values ​​for the syntax elements that make up the picture or block.

[0036] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transformation units (TUs). According to HEVC, a video coder (such as video encoder 200) divides coding tree units (CTUs) into CUs according to a quadtree structure. That is, the video coder divides the CTUs and CUs into four equal, non-overlapping squares, and each node in the quadtree has either zero or four child nodes. Nodes without child nodes are sometimes called "leaf nodes," and the CU of such a leaf node may contain one or more PUs and / or one or more TUs. The video coder may further divide the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents a division of TUs. In HEVC, PUs represent intra-predicted data, while TUs represent residual data. Intra-predicted CUs contain intra-predicted information such as intra-mode indications.

[0037] As another example, a video encoder 200 and a video decoder 300 may be configured to operate according to VVC. According to VVC, a video coder (such as the video encoder 200) divides a picture into multiple coding tree units (CTUs). The video encoder 200 may divide the CTUs according to a tree structure such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple division types, such as the separation of CUs, PUs, and TUs in HEVC. The QTBT structure has two levels: a first level divided according to quadtree divisions and a second level divided according to binary tree divisions. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0038] In MTT partitioning structures, blocks can be partitioned using quadtree (QT) partitions, binary tree (BT) partitions, and one or more types of triple tree (TT) (also called ternary tree (TT)) partitions. A triple tree or ternary tree partition is a partition in which a block is divided into three subblocks. In some examples, a triple tree or ternary tree partition divides a block into three subblocks without dividing the original block through a center. The partition types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.

[0039] In some examples, the video encoder 200 and video decoder 300 may use a single QTBT or MTT structure to represent the luminance component and the chromaticity component, respectively, while in other examples, the video encoder 200 and video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for both chromaticity components (or two QTBT / MTT structures for each chromaticity component).

[0040] The video encoder 200 and video decoder 300 may be configured to use a quadtree partition, QTBT partition, MTT partition, or other partition structure according to HEVC. For illustrative purposes, the description of the techniques of this disclosure will be presented in relation to QTBT partitions. However, it should be understood that the techniques of this disclosure may also be applicable to video coders configured to use quadtree partitions or other types of partitions.

[0041] In some examples, a CTU includes a coding tree block (CTB) of a luma sample, two corresponding CTBs of a chroma sample for a picture with three sample arrays, or a CTB of a sample for a picture coded using three separate color planes and syntax structures used to code a monochrome picture or sample. A CTB can be an N×N block of samples for some value N, such that the division of components into the CTB is a partition. A component is a single sample from one array or one of three arrays (luma and two chromas) that make up a picture in a 4:2:0, 4:2:2, or 4:4:4 color format, or a single sample from an array or array that makes up a picture in a monochrome format. In some examples, a coding block is an M×N block of samples for some values ​​M and N, such that the division of the CTB into the coding block is a partition.

[0042] Blocks (e.g., CTUs or CUs) can be grouped in various ways within a picture. For example, a brick may refer to a rectangular area of ​​a row of CTUs within a particular tile in a picture. A tile can be a rectangular area of ​​CTUs within a particular tile column or row in a picture. A tile column refers to a rectangular area of ​​CTUs with a height equal to the height of the picture and a width specified by a syntax element (e.g., in a picture parameter set). A tile row refers to a rectangular area of ​​CTUs with a height specified by a syntax element (e.g., in a picture parameter set) and a width equal to the width of the picture.

[0043] In some examples, a tile may be divided into multiple bricks, each brick containing one or more CTU rows within the tile. A tile that is not divided into multiple bricks may also be called a brick. However, a brick that is a true subset of a tile may not be called a tile.

[0044] Bricks within a picture can also be arranged in a slice. A slice can be an integer number of picture bricks that can exclusively reside within a single Network Abstraction Layer (NAL) unit. In some examples, a slice may contain either several complete tiles or only a series of consecutive complete bricks of a single tile.

[0045] This disclosure may use "N×N" and "N to N" interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) with respect to vertical and horizontal dimensions, for example, 16×16 samples or 16 to 16 samples. Generally, a 16×16 CU has 16 samples vertically (y=16) and 16 samples horizontally (x=16). Similarly, an N×N CU generally has N samples vertically and N samples horizontally, where N represents a non-negative integer. The samples in a CU can be arranged as rows and columns. Furthermore, a CU does not necessarily have to have the same number of samples horizontally as vertically. For example, a CU may have N×M samples, where M is not necessarily equal to N.

[0046] The video encoder 200 encodes video data for the CU, representing prediction information and / or residual information, as well as other information. The prediction information indicates how the CU should be predicted in order to form a prediction block for the CU. The residual information generally represents the sample-by-sample difference between the CU sample before encoding and the prediction block.

[0047] To predict a CU (Critical Unit), the video encoder 200 can generally form prediction blocks for the CU through inter-prediction or intra-prediction. Inter-prediction generally refers to predicting the CU from data of a previously coded picture, while intra-prediction generally refers to predicting the CU from previously coded data of the same picture. To perform inter-prediction, the video encoder 200 can generate prediction blocks using one or more motion vectors. The video encoder 200 can generally perform motion search to identify a reference block that closely matches the CU, for example, with respect to the difference between the CU and the reference block. The video encoder 200 can calculate a difference metric using absolute difference sum (SAD), squared difference sum (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can predict the current CU using unidirectional or bidirectional prediction.

[0048] Some examples of VVC also offer an affine motion compensation mode, which can be considered an interpredictive mode. In affine motion compensation mode, the video encoder 200 may determine two or more motion vectors representing non-translational motion, such as zooming in or zooming out, rotation, projection motion, or other irregular motion types.

[0049] To perform intra-prediction, the video encoder 200 may select an intra-prediction mode to generate prediction blocks. Several examples of VVCs provide 67 intra-prediction modes, including various directional modes, as well as planar and DC modes. Generally, the video encoder 200 selects an intra-prediction mode that describes neighboring samples to the current block (e.g., a block of CUs) from which samples of the current block should be predicted. Such samples can generally be above, above and to the left of, or to the left of, the current block, in the same picture as the current block, assuming that the video encoder 200 codes the CTUs and CUs in raster scanning order (left to right, top to bottom).

[0050] The video encoder 200 encodes data representing the prediction mode for the current block. For example, in interprediction mode, the video encoder 200 may encode data representing which of the various available interprediction modes is used, as well as motion information for the corresponding mode. For unidirectional or bidirectional interprediction, for example, the video encoder 200 may encode motion vectors using advanced motion vector prediction (AMVP) mode or merge mode. The video encoder 200 may encode motion vectors for affine motion compensation mode using similar modes.

[0051] According to predictions such as intra-prediction or inter-prediction of a block, the video encoder 200 may calculate residual data for the block. Residual data, such as residual blocks, represents the sample-by-sample difference between the block and the predicted block for the block, formed using the corresponding prediction mode. The video encoder 200 may apply one or more transformations to the residual blocks to generate transformed data in the transformation region instead of the sample region. For example, the video encoder 200 may apply a discrete cosine transform (DCT), integer transform, wavelet transform, or a conceptually similar transform to the residual video data. In addition, the video encoder 200 may apply secondary transformations such as mode-dependent non-separable secondary transform (MDNSST), signal-dependent transform, or Carunen-Löwe ​​transform (KLT) following the initial transformation. The video encoder 200 generates transformation coefficients following the application of one or more transformations.

[0052] As described above, following any transformation to generate the transformation coefficients, the video encoder 200 may perform quantization of the transformation coefficients. Quantization generally refers to the process of quantizing the transformation coefficients to reduce the amount of data used to represent them as much as possible, thereby achieving further compression. By performing the quantization process, the video encoder 200 may reduce the bit depth associated with some or all of the transformation coefficients. For example, the video encoder 200 may truncate an n-bit value to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 may perform a bitwise right shift of the value to be quantized.

[0053] Following quantization, the video encoder 200 may scan the transformation coefficients and generate a one-dimensional vector from a two-dimensional matrix containing the quantized transformation coefficients. The scan may be designed to place higher-energy (and therefore lower-frequency) transformation coefficients at the beginning of the vector and lower-energy (and therefore higher-frequency) transformation coefficients at the end. In some examples, the video encoder 200 may generate a serialized vector using a predetermined scan order to scan the quantized transformation coefficients, and then entropy-code the quantized transformation coefficients of the vector. In other examples, the video encoder 200 may perform an adaptive scan. After scanning the quantized transformation coefficients to form a one-dimensional vector, the video encoder 200 may entropy-code the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). The video encoder 200 may also entropy-code values ​​for syntax elements that describe metadata related to the encoded video data for use by the video decoder 300 when decoding the video data.

[0054] To perform CABAC, the video encoder 200 may assign a context within a context model to the symbols to be transmitted. The context may relate, for example, to whether the adjacent values ​​of the symbols are 0. Probability decisions may be based on the context assigned to the symbols.

[0055] The video encoder 200 may further generate syntax data for the video decoder 300, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, for example, in the form of a picture header, block header, slice header, or other syntax data such as a sequence parameter set (SPS), picture parameter set (PPS), or video parameter set (VPS). The video decoder 300 may similarly decode such syntax data to determine how to decode the corresponding video data.

[0056] In this way, the video encoder 200 can generate a bitstream containing encoded video data, for example, syntax elements describing the division of a picture into blocks (e.g., CUs) and prediction and / or residual information for those blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.

[0057] Generally, the video decoder 300 performs the reverse process of the process performed by the video encoder 200 in order to decode the encoded video data of the bitstream. For example, the video decoder 300 may decode values ​​for syntax elements of the bitstream using CABAC in a substantially similar, but reverse, manner to the CABAC encoding process of the video encoder 200. The syntax elements may define the CUs of the CTUs by defining partitioning information for partitioning the picture into CTUs, and the partitions of each CTU according to a corresponding partitioning structure such as a QTBT structure. The syntax elements may further define prediction and residual information for blocks of video data (e.g., CUs).

[0058] Residual information may be represented, for example, by quantized transformation coefficients. The video decoder 300 may reconstruct the residual block for the block by inverse quantizing and inverse transforming the quantized transformation coefficients of the block. The video decoder 300 uses the signaled prediction mode (intra-prediction or inter-prediction) and associated prediction information (for example, motion information for inter-prediction) to form a predicted block for the block. The video decoder 300 may then combine the predicted block and the residual block (sample by sample) to reconstruct the original block. The video decoder 300 may perform additional processing, such as performing a deblocking process, to reduce visual artifacts along the block boundaries.

[0059] According to the technique of the present disclosure, the method comprises the steps of: applying a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector; and decoding the block based on at least one improved motion vector, wherein the multipath DMVR comprises: a first pass which is block-based and applied to a block of video data; a second pass which is subblock-based and applied to at least one second pass subblock of the block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass which is subblock-based and applied to at least one third pass subblock of the block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0060] According to the techniques of the present disclosure, a device includes a memory configured to store video data and one or more processors implemented by circuitry and communicatively coupled to the memory, the one or more processors configured to apply a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector, and to decode the block based on the at least one improved motion vector, the multipath DMVR comprising: a first pass which is block-based and applied to a block of video data; a second pass which is subblock-based and applied to at least one second pass subblock of a block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass which is subblock-based and applied to at least one third pass subblock of a block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0061] According to the techniques of the present disclosure, a non-temporary computer-readable medium stores instructions that, when executed, cause one or more processors to apply a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector, and to decode the block based on the at least one improved motion vector, wherein the multipath DMVR comprises a first pass that is block-based and applied to a block of video data; a second pass that is subblock-based and applied to at least one second pass subblock of a block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass that is subblock-based and applied to at least one third pass subblock of a block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0062] According to the technique of the present disclosure, the device includes means for applying a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector, and means for decoding the block based on at least one improved motion vector, wherein the multipath DMVR comprises a first pass that is block-based and applied to a block of video data, a second pass that is subblock-based and applied to at least one second pass subblock of the block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data, and a third pass that is subblock-based and applied to at least one third pass subblock of the block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0063] According to the technique of the present disclosure, the method includes the steps of applying a multipath decoder-side motion vector improvement (DMVR) to the motion vector for a block of video data to determine an improved motion vector, and coding the block based on the improved motion vector.

[0064] According to the techniques of the present disclosure, the device includes a memory configured to store video data and one or more processors that are circuit-implemented and communicatively coupled to the memory, the one or more processors being configured to perform any of the techniques of the present disclosure.

[0065] According to the techniques of this disclosure, the device includes at least one means for performing any of the techniques of this disclosure.

[0066] According to the techniques of this disclosure, a computer-readable storage medium encodes instructions that, when executed, cause a programmable processor to perform one of the techniques of this disclosure.

[0067] This disclosure may refer in general to “signaling” any information, such as syntax elements. The term “signaling” may generally refer to the communication of values ​​for syntax elements and / or other data used to decode the encoded video data. That is, the video encoder 200 may signal values ​​for syntax elements in the bitstream. In general, signaling refers to generating values ​​in the bitstream. As stated above, the source device 102 may transfer the bitstream to the destination device 116 substantially in real time or non-real time, which may occur, for example, when the destination device 116 stores the syntax elements in the storage device 112 for later retrieval.

[0068] Figures 2A and 2B are conceptual diagrams showing an exemplary quadtree-binary tree (QTBT) structure 130 and its corresponding coding tree unit (CTU) 132. Solid lines represent quadtree partitions, and dotted lines represent binary tree partitions. At each partition (i.e., non-leaf) node of the binary tree, one flag is signaled to indicate which partition type (i.e., horizontal or vertical) is used; in this example, 0 indicates a horizontal partition and 1 indicates a vertical partition. In the case of a quadtree partition, there is no need to indicate the partition type, as the quadtree node divides the block horizontally and vertically into four subblocks of equal size. Thus, the video encoder 200 may encode syntax elements (such as partition information) for the domain tree level (i.e., solid lines) of the QTBT structure 130, and the video decoder 300 may decode syntax elements (such as partition information) for the predictive tree level (i.e., dashed lines) of the QTBT structure 130. The video encoder 200 may encode video data, such as prediction data and transformation data, for CUs represented by terminal leaf nodes of the QTBT structure 130, and the video decoder 300 may decode it.

[0069] In general, CTU132 in Figure 2B can be associated with parameters that define the size of the blocks corresponding to the nodes of the QTBT structure 130 at the first and second levels. These parameters may include the CTU size (representing the size of CTU132 in the sample), the minimum quadtree size (MinQTSize, representing the minimum allowed quadtree leaf node size), the maximum binary tree size (MaxBTSize, representing the maximum allowed binary tree root node size), the maximum binary tree depth (MaxBTDepth, representing the maximum allowed binary tree depth), and the minimum binary tree size (MinBTSize, representing the minimum allowed binary tree leaf node size).

[0070] The root node of a QTBT structure corresponding to a CTU may have four child nodes at the first level of the QTBT structure, each of which may be partitioned according to a quadtree partition. That is, a node at the first level may have a leaf node (no child nodes) or four child nodes. An example of QTBT structure 130 represents such a node, including a parent node and child nodes with solid lines for branching. If a node at the first level is not larger than the maximum allowable binary tree root node size (MaxBTSize), the node may be further partitioned by its respective binary tree. Binary tree partitioning of a single node may be repeated until the node resulting from the partition reaches the minimum allowable binary tree leaf node size (MinBTSize) or the maximum allowable binary tree depth (MaxBTDepth). An example of QTBT structure 130 represents a node with dashed lines for branching. Binary tree leaf nodes are called coding units (CUs), which are used for prediction (e.g., in-picture or between-picture predictions) and transformations without further partitioning. As discussed above, CU can also be called “video block” or “block”.

[0071] In one example of a QTBT partitioned structure, the CTU size is set to 128×128 (a chroma sample and two corresponding 64×64 chroma samples), MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. To generate a quadtree leaf node, the quadtree partition is first applied to the CTU. The quadtree leaf node can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the quadtree leaf node is 128×128, the size exceeds MaxBTSize (i.e., 64×64 in this example), so the quadtree leaf node is not further partitioned by a binary tree. Otherwise, the quadtree leaf node is further partitioned by a binary tree. Therefore, a quadtree leaf node is also the root node of a binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (4 in this example), no further partitioning is allowed. A binary tree node with a width equal to MinBTSize (4 in this example) suggests that no further vertical partitioning (i.e., partitioning by width) is allowed for that binary tree node. Similarly, a binary tree node with a height equal to MinBTSize suggests that no further horizontal partitioning (i.e., partitioning by height) is allowed for that binary tree node. As mentioned above, a leaf node of a binary tree is called a CU and is further processed according to prediction and transformation without further partitioning.

[0072] Figure 3 is a block diagram showing an exemplary video encoder 200 capable of performing the techniques of this disclosure. Figure 3 is provided for illustrative purposes and should not be considered a limitation of the techniques as broadly illustrated and described herein. For illustrative purposes, this disclosure describes a video encoder 200 using the VVC (ITU-T H.266) and HEVC (ITU-T H.265) techniques. However, the techniques of this disclosure may be performed by video encoding devices configured according to other video coding standards.

[0073] In the example shown in Figure 3, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any or all of the video data memory 230, mode selection unit 202, residual generation unit 204, conversion processing unit 206, quantization unit 208, inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a filter unit 216, a DPB 218, and an entropy coding unit 220 may be implemented in one or more processors or processing circuits. For example, the units of the video encoder 200 may be implemented as one or more circuits or logic elements as part of a hardware circuit, or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuits to perform these and other functions.

[0074] The video data memory 230 can store video data to be encoded by the components of the video encoder 200. The video encoder 200 may receive video data to be stored in the video data memory 230 from, for example, a video source 104 (Figure 1). The DPB 218 may act as a reference picture memory, storing reference video data for use by the video encoder 200 in predicting subsequent video data. The video data memory 230 and DPB 218 may be formed by any of various memory devices, such as dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and DPB 218 may be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 may be on-chip with the other components of the video encoder 200, as shown, or off-chip relative to those components.

[0075] In this disclosure, references to the video data memory 230 should not be interpreted as being limited to memory inside the video encoder 200 unless otherwise stated, nor should they be interpreted as being limited to memory outside the video encoder 200 unless otherwise stated. Rather, references to the video data memory 230 should be understood as reference memory that stores video data received by the video encoder 200 for encoding (for example, video data for the current block to be encoded). Memory 106 in Figure 1 may also temporarily store outputs from various units of the video encoder 200.

[0076] The various units in Figure 3 are shown to help understand the operations performed by the video encoder 200. The units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide a specific function, with predefined operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks, offering flexibility in the operations they can perform. For example, a programmable circuit may execute software or firmware that operates it, defined by software or firmware instructions. Fixed-function circuits may execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0077] The video encoder 200 may include a programmable core formed from an arithmetic logic unit (ALU), an basic function unit (EFU), digital circuits, analog circuits, and / or programmable circuits. In an example where the operation of the video encoder 200 is performed using software executed by the programmable circuits, memory 106 (Figure 1) may store instructions (e.g., object code) of the software that the video encoder 200 receives and executes, or another memory (not shown) within the video encoder 200 may store such instructions.

[0078] The video data memory 230 is configured to store the received video data. The video encoder 200 can retrieve a picture of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 may be raw video data to be encoded.

[0079] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-prediction unit 226. The mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes. For example, the mode selection unit 202 may include a palette unit, an intra-block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, and so on.

[0080] The mode selection unit 202 generally coordinates multiple coding paths to test combinations of coding parameters and the rate distortion values ​​obtained for such combinations. Coding parameters may include the division of the CTU to the CU, the prediction mode for the CU, the transformation type for the residual data of the CU, and the quantization parameters for the residual data of the CU. The mode selection unit 202 can ultimately select a combination of coding parameters that yields a better rate distortion value than other tested combinations.

[0081] The video encoder 200 divides the picture retrieved from the video data memory 230 into a series of CTUs, and may encapsulate one or more CTUs within a slice. The mode selection unit 202 may divide the picture's CTUs according to a tree structure, such as a QTBT structure or the HEVC quadtree structure described above. As described above, the video encoder 200 may form one or more CUs from dividing the CTUs according to a tree structure. Such CUs may also be commonly referred to as “video blocks” or “blocks”.

[0082] In general, the mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate a predicted block for the current block (e.g., the current CU, or in HEVC, the overlapping portion of PU and TU). For intra-prediction of the current block, the motion estimation unit 222 may perform a motion search to identify one or more well-matching reference blocks among one or more reference pictures (e.g., one or more previously coded pictures stored in DPB 218). Specifically, the motion estimation unit 222 may calculate a value representing how similar a potential reference block is to the current block, for example, according to the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), etc. The motion estimation unit 222 may generally perform these calculations using the sample-by-sample difference between the current block and the reference block under consideration. The motion estimation unit 222 may identify the reference block with the lowest value resulting from these calculations, indicating the reference block that best matches the current block.

[0083] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to the position of a current block in the current picture. The motion estimation unit 222 may then provide the motion vectors to the motion compensation unit 224. For example, in the case of unidirectional interpretation, the motion estimation unit 222 may provide a single motion vector, while in the case of bidirectional interpretation, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then use the motion vectors to generate prediction blocks. For example, the motion compensation unit 224 may use the motion vectors to extract data for a reference block. As another example, if the motion vectors have non-integer sample precision, the motion compensation unit 224 may interpolate values ​​for the prediction blocks according to one or more interpolation filters. Furthermore, in the case of bidirectional interpretation, the motion compensation unit 224 may extract data for two reference blocks identified by each motion vector and combine the extracted data, for example, through a sample-by-sample average or a weighted average.

[0084] As another example, in the case of intra-prediction or intra-prediction coding, the intra-prediction unit 226 may generate a prediction block from samples adjacent to the current block. For example, in directional mode, the intra-prediction unit 226 may generally mathematically combine the values ​​of adjacent samples and populate these calculated values ​​in a direction defined across the current block to produce a prediction block. As another example, in DC mode, the intra-prediction unit 226 may calculate the average of the samples adjacent to the current block and generate a prediction block that includes this obtained average for each sample in the prediction block.

[0085] The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives a raw, unencoded version of the current block from the video data memory 230 and the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block relative to the current block. In some examples, the residual generation unit 204 may also determine the differences between sample values ​​in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, the residual generation unit 204 may be formed using one or more subtractor circuits that perform binary subtraction.

[0086] In an example where the mode selection unit 202 divides CUs into PUs, each PU may be associated with a luma prediction unit and a corresponding chroma prediction unit. The video encoder 200 and video decoder 300 may support PUs of various sizes. As shown above, the size of a CU may refer to the size of the luma coding block of the CU, and the size of a PU may refer to the size of the luma prediction unit of the PU. Assuming that the size of a particular CU is 2N×2N, the video encoder 200 may support PU sizes of 2N×2N or N×N for intra-prediction, and 2N×2N, 2N×N, N×2N, N×N, or similar symmetric PU sizes for inter-prediction. The video encoder 200 and video decoder 300 may also support asymmetric divisions for PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-prediction.

[0087] In cases where the mode selection unit 202 does not further subdivide the CUs into PUs, each CU may be associated with a lumacoding block and a corresponding chromacoding block. As shown above, the size of a CU may refer to the size of the lumacoding block within the CU. The video encoder 200 and video decoder 300 may support CU sizes of 2N×2N, 2N×N, or N×2N.

[0088] In some examples, such as intra-block copy mode coding, affine mode coding, and other video coding techniques like linear model (LM) mode coding, the mode selection unit 202 generates a predicted block for the current block being coded via the respective units associated with the coding technique. In some examples, such as palette mode coding, the mode selection unit 202 does not need to generate a predicted block; instead, it may generate syntax elements indicating a scheme for reconstructing the block based on the selected palette. In such modes, the mode selection unit 202 may provide these syntax elements to the entropy coding unit 220 to be coded.

[0089] As described above, the residual generation unit 204 receives video data for the current block and the corresponding predicted block. The residual generation unit 204 then generates residual blocks for the current block. To generate residual blocks, the residual generation unit 204 calculates the sample-by-sample difference between the predicted block and the current block.

[0090] The transformation processing unit 206 applies one or more transformations to the residual block to generate a block of transformation coefficients (referred herein to as the "transformation coefficient block"). The transformation processing unit 206 may apply various transformations to the residual block to form the transformation coefficient block. For example, the transformation processing unit 206 may apply a discrete cosine transform (DCT), a direction transform, a Carunenlobe transform (KLT), or a conceptually similar transformation to the residual block. In some examples, the transformation processing unit 206 may perform multiple transformations on the residual block, such as linear and quadratic transformations, including a rotation transform. In some examples, the transformation processing unit 206 does not apply any transformations to the residual block.

[0091] The quantization unit 208 can quantize the transformation coefficients in the transformation coefficient block to produce a quantized transformation coefficient block. The quantization unit 208 can quantize the transformation coefficients in the transformation coefficient block according to the quantization parameter (QP) value associated with the current block. The video encoder 200 can adjust the degree of quantization applied to the transformation coefficient block associated with the current block by adjusting the QP value associated with the CU (for example, via the mode selection unit 202). Quantization may result in a loss of information, and therefore the quantized transformation coefficients may be less precise than the original transformation coefficients generated by the transformation processing unit 206.

[0092] The inverse quantization unit 210 and the inverse transformation processing unit 212 can reconstruct residual blocks from transformation coefficient blocks by applying inverse quantization and inverse transformation, respectively, to the quantized transformation coefficient blocks. The reconstruction unit 214 can generate reconstructed blocks corresponding to the current blocks (which may be with some distortion) based on the reconstructed residual blocks and prediction blocks generated by the mode selection unit 202. For example, the reconstruction unit 214 can generate reconstructed blocks by adding samples from the reconstructed residual blocks to the corresponding samples from the prediction blocks generated by the mode selection unit 202.

[0093] The filter unit 216 may perform one or more filtering operations on the reconstructed block. For example, the filter unit 216 may perform a deblocking operation to reduce blockingness artifacts along the edges of the CU. In some examples, the operation of the filter unit 216 may be skipped.

[0094] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in an example where the filter unit 216 does not operate, the reconstruction unit 214 may store the reconstructed blocks in the DPB 218. In an example where the filter unit 216 operates, the filter unit 216 may store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may take a reference picture from the DPB 218, which is formed from the reconstructed (and possibly filtered) blocks, to interpret blocks of the picture to be encoded later. In addition, the intraprediction unit 226 may use the reconstructed blocks in the DPB 218 of the current picture to intrapret other blocks in the current picture.

[0095] In general, the entropy coding unit 220 can entropy code syntax elements received from other functional components of the video encoder 200. For example, the entropy coding unit 220 can entropy code a quantized conversion coefficient block from the quantization unit 208. As another example, the entropy coding unit 220 can entropy code prediction syntax elements from the mode selection unit 202 (e.g., motion information for inter-prediction or intra-mode information for intra-prediction). The entropy coding unit 220 can perform one or more entropy coding operations on syntax elements, which are another example of video data, to generate entropy-coded data. For example, the entropy coding unit 220 may perform context-adaptive variable-length coding (CAVLC) operation, CABAC operation, variable-length to variable-length (V2V) coding operation, syntax-based context-adaptive binary arithmetic coding (SBAC) operation, probability interval partitioned entropy (PIPE) coding operation, exponential Golomb coding operation, or another type of entropy coding operation on the data. In some examples, the entropy coding unit 220 may operate in a bypass mode in which syntax elements are not entropically coded.

[0096] The video encoder 200 may output a bitstream containing entropy-encoded syntax elements required to reconstruct a slice or block of a picture. Specifically, the entropy encoding unit 220 may output a bitstream.

[0097] The behavior described above is described in relation to blocks. Such a description should be understood as the behavior for lumacoding blocks and / or chromacoding blocks. As described above, in some examples, lumacoding blocks and chromacoding blocks are the luma and chroma components of CU. In some examples, lumacoding blocks and chromacoding blocks are the luma and chroma components of PU.

[0098] In some cases, actions performed for a lumacoding block do not need to be repeated for a chromacoding block. For example, actions to identify the motion vector (MV) and reference picture for a lumacoding block do not need to be repeated to identify the MV and reference picture for a chromablock. Rather, the MV for the lumacoding block may be scaled to determine the MV for the chromablock, and the reference picture may be the same. In another example, the intra-prediction process may be the same for both lumacoding and chromacoding blocks.

[0099] Figure 4 is a block diagram showing an exemplary video decoder 300 capable of performing the techniques of this disclosure. Figure 4 is provided for illustrative purposes and is not intended to limit the techniques that are more broadly illustrated and described in this disclosure. For illustrative purposes, this disclosure describes a video decoder 300 using VVC (ITU-T H.266) and HEVC (ITU-T H.265) techniques. However, the techniques of this disclosure may be performed by video coding devices configured according to other video coding standards.

[0100] In the example shown in Figure 4, the video decoder 300 includes a coding picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transformation processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoding picture buffer (DPB) 314. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transformation processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 may be implemented in one or more processors or processing circuits. For example, the units of the video decoder 300 may be implemented as one or more circuits or logic elements as part of a hardware circuit, or as part of a processor, ASIC, or FPGA. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuits to perform these and other functions.

[0101] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-prediction unit 318. The prediction processing unit 304 may include additional units for performing predictions according to other prediction modes. For example, the prediction processing unit 304 may include a pallet unit, an intra-block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, and so on. In other examples, the video decoder 300 may include more, fewer, or different functional components. The motion compensation unit 316 may include a multipath DMVR unit (MPDMVR) 317, which will be discussed in the following discussion of the motion compensation unit 316.

[0102] The CPB memory 320 can store video data, such as an encoded video bitstream, to be decoded by the components of the video decoder 300. Video data stored in the CPU memory 320 may be obtained, for example, from a computer-readable medium 110 (Figure 1). The CPU memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. The CPB memory 320 may also store video data other than syntax elements of the coded picture, such as temporary data representing outputs from various units of the video decoder 300. The DPB 314 generally stores the decoded picture, and the video decoder 300 may output and / or use this decoded picture as reference video data when decoding subsequent data or pictures from the encoded video bitstream. The CPB memory 320 and DPB 314 may be formed by any of various memory devices, such as DRAM, MRAM, RRAM, or other types of memory devices, including SDRAM. The CPU memory 320 and DPB 314 may be provided by the same memory device or by separate memory devices. In various examples, the CPB memory 320 may be on-chip with the other components of the video decoder 300, or it may be off-chip relative to those components.

[0103] As an addition or alternative, in some examples, the video decoder 300 may retrieve coded video data from memory 120 (Figure 1). That is, memory 120 may store data such as those discussed above together with the CPB memory 320. Similarly, memory 120 may store instructions to be executed by the video decoder 300 when some or all of the functions of the video decoder 300 are implemented in software to be performed by the processing circuit of the video decoder 300.

[0104] The various units shown in Figure 4 are presented to help understand the operations performed by the video decoder 300. The units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to Figure 3, fixed-function circuits refer to circuits that provide a specific function, with predefined operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks, offering flexibility in the operations they can perform. For example, a programmable circuit may execute software or firmware that operates it, in a manner defined by software or firmware instructions. Fixed-function circuits may execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0105] The video decoder 300 may include a programmable core formed from an ALU, EFU, digital circuitry, analog circuitry, and / or programmable circuitry. In an example where the operation of the video decoder 300 is performed by software running on the programmable circuitry, on-chip memory or off-chip memory may store software instructions (e.g., object code) that the video decoder 300 receives and executes.

[0106] The entropy decoding unit 302 can receive video data encoded from the CPB and reconstruct the syntax elements by entropy decoding the video data. The prediction processing unit 304, the inverse quantization unit 306, the inverse transformation processing unit 308, the reconstruction unit 310, and the filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.

[0107] Generally, the video decoder 300 reconstructs the picture block by block. The video decoder 300 may perform the reconstruction operation for each block individually (where the block currently being reconstructed, i.e., decoded, may be called the "current block").

[0108] The entropy decoding unit 302 can entropy decode the syntax elements that define the quantized transformation coefficients of the quantized transformation coefficient block, as well as transformation information such as quantization parameters (QP) and / or transformation mode indications. The inverse quantization unit 306 may use the QP associated with the quantized transformation coefficient block to determine the degree of quantization and, similarly, the degree of inverse quantization that the inverse quantization unit 306 should apply. The inverse quantization unit 306 may, for example, perform a bitwise left shift operation to inverse quantize the quantized transformation coefficients. In this way, the inverse quantization unit 306 may form a transformation coefficient block containing the transformation coefficients.

[0109] After the inverse quantization unit 306 has formed a transformation coefficient block, the inverse transformation processing unit 308 may apply one or more inverse transformations to the transformation coefficient block to generate a residual block associated with the current block. For example, the inverse transformation processing unit 308 may apply an inverse DCT, an inverse integer transformation, an inverse Carunenlebe transformation (KLT), an inverse rotation transformation, an inverse direction transformation, or another inverse transformation to the transformation coefficient block.

[0110] Furthermore, the prediction processing unit 304 generates prediction blocks according to the prediction information syntax elements entropy-decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is interpredicted, the motion compensation unit 316 may generate a prediction block. In this case, the prediction information syntax elements may indicate a reference picture in the DPB 314 from which the reference block should be extracted, as well as a motion vector that identifies the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can generally perform interprediction processing in substantially the same manner as described with respect to the motion compensation unit 224 (Figure 3).

[0111] In some examples, the motion compensation unit 316 may include a multi-pass DMVR unit 317. The multi-pass DMVR unit 317 may apply multi-pass DMVR to motion vectors for blocks of video data to determine improved motion vectors. The multi-pass DMVR may include a first pass, which is block-based and applied to blocks of video data. The multi-pass DMVR may include a second pass, which is sub-block-based and applied to at least one second pass sub-block of a block of video data. The multi-pass DMVR may include a third pass, which is sub-block-based and applied to at least one third pass sub-block of a block of video data. The width of the second pass sub-block may be less than or equal to the width of the block of video data, and the height of the second pass sub-block may be less than or equal to the height of the block of video data. The width of the third pass sub-block may be less than or equal to the width of the second pass sub-block, and the height of the third pass sub-block may be less than or equal to the height of the second pass sub-block. Further examples and explanations of multi-pass DMVR techniques are described later in this disclosure.

[0112] As another example, if the prediction information syntax element indicates that the current block is to be intra-predicted, the intra-prediction unit 318 may generate a predicted block according to the intra-prediction mode indicated by the prediction information syntax element. Again, the intra-prediction unit 318 may generally perform intra-prediction processing in substantially the same manner as described with respect to the intra-prediction unit 226 (Figure 3). The intra-prediction unit 318 may retrieve data from the DPB 314 for samples adjacent to the current block.

[0113] The reconstruction unit 310 may reconstruct the current block using the predicted block and the residual block. For example, the reconstruction unit 310 may reconstruct the current block by adding samples from the residual block to the corresponding samples from the predicted block.

[0114] The filter unit 312 may perform one or more filtering operations on the reconstructed block. For example, the filter unit 312 may perform a deblocking operation to reduce blockingness artifacts along the edges of the reconstructed block. In all examples, the filter unit 312 may not necessarily perform any operations.

[0115] The video decoder 300 may store the reconstructed blocks in the DPB 314. For example, in an example where the filter unit 312 does not operate, the reconstruction unit 310 may store the reconstructed blocks in the DPB 314. In an example where the filter unit 312 operates, the filter unit 312 may store the filtered reconstructed blocks in the DPB 314. As discussed above, the DPB 314 may provide the prediction processing unit 304 with reference information such as the current picture for intra-prediction and samples of previously decoded pictures for subsequent motion compensation. Furthermore, the video decoder 300 may output the decoded picture (e.g., decoded video) from the DPB 314 for later presentation on a display device such as the display device 118 in Figure 1.

[0116] Thus, the video decoder 300 represents an example of a video decoding device, comprising a memory configured to store video data and one or more processors implemented in circuitry and communicably coupled to the memory, the one or more processors configured to apply a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector, and to decode the block based on the at least one improved motion vector, the multipath DMVR comprising a first pass that is block-based and applied to a block of video data; a second pass that is subblock-based and applied to at least one second pass subblock of a block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass that is subblock-based and applied to at least one third pass subblock of a block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0117] The video decoder 300 also represents an example of a video decoding device, which includes a memory configured to store video data and one or more processing units implemented in the circuit, the processing units being configured to apply a multipath decoder-side motion vector improvement (DMVR) to the motion vector for a block of video data to determine an improved motion vector, and to decode the block based on the improved motion vector.

[0118] This disclosure relates to decoder-side motion vector derivation techniques (e.g., template matching, bilateral matching, decoder-side MV improvements, bidirectional optical flow, etc.). The techniques of this disclosure may be applied to any existing video codec such as HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), and Essential Video Coding (EVC), or may be efficient coding tools in any future video coding standard. This section first considers the HEVC and JEM techniques and ongoing work on Versatile Video Coding (VVC) related to this disclosure.

[0119] Video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, and ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), including their Scalable Video Coding (SVC) and Multi-view Video Coding (MVC) extensions. The algorithmic description of Versatile Video Coding and Test Model 10 (VTM 10.0) may be called JVET-T2002, available from https: / / jvet-experts.org / .

[0120] The CU structure and motion vector prediction in HEVC are discussed here. In HEVC, the largest coding unit in a slice is called a coding tree block (CTB) or coding tree unit (CTU). A CTB may contain a quadtree whose nodes are coding units.

[0121] (Technically, an 8x8 CTB size may be supported.) The size of the CTB can range from 16x16 to 64x64 in the HEVC main profile. Coding units (CUs) can be the same size as the CTB or as small as 8x8. Each CU is coded using one mode, i.e., intermode or intramode. When a CU is intercoded, it may be further divided into two or four predictive units (PUs), or it may remain as a single PU if no further division is applied. When two PUs exist within a single CU, the two PUs may each be a rectangle half the size of the CU (half the size of the CU), or they may be two rectangles of different sizes, one being 1 / 4 the size of the CU and the other 3 / 4 the size of the CU.

[0122] When a CU is intercoded, each PU has one set of motion information, which is derived using a unique inter-prediction mode.

[0123] Here, motion vector prediction is discussed. The HEVC standard has two interpretation modes for the PU, named merge mode (skipping is considered a special case of merging) and advanced motion vector prediction (AMVP) mode.

[0124] In either AMVP mode or merge mode, a motion vector (MV) candidate list is maintained for multiple motion vector predictors. Currently, the MV, and the reference index in merge mode, are generated by taking one candidate from the MV candidate list. For example, the video decoder 300 may maintain an MV candidate list.

[0125] The MV candidate list contains up to five candidates for merge mode and only two candidates for AMVP mode. A merge candidate may contain MVs corresponding to both a set of motion information, such as a reference picture list (lists 0 and 1) and its corresponding reference index. When a merge candidate is identified by its merge index, the reference picture used for prediction of the current block, as well as the associated motion vector, are determined. On the other hand, in AMVP mode for each potential prediction direction from either list 0 or list 1, the AMVP candidate contains only MVs, so the video encoder 200 may explicitly signal the reference index along with the MV predictor (MVP) index to the MV candidate list. In AMVP mode, the predicted MV can be further refined.

[0126] The video decoder 300 can similarly derive candidates for both modes from the same spatial and temporal adjacent blocks.

[0127] Figures 5A and 5B are conceptual diagrams showing exemplary spatially adjacent MV candidates for merge mode and AMVP mode, respectively. Spatial MV candidates are derived from adjacent blocks shown in Figures 5A and 5B for a given PU (PU0), but the method of generating candidates from blocks differs between merge mode and AMVP mode.

[0128] In merge mode, up to four spatial MV candidates for PU0 500 can be derived in the order shown in Figure 5A, where the numbers increase in the order of left (0,A1), top (1,B1), top right (2,B0), bottom left (3,A0), and top left (4,B2). For example, video decoder 300 may derive up to four spatial MV candidates for PU0 500 using the order described above.

[0129] As shown in Figure 5B, in AMVP mode, the adjacent blocks of PU0 502 are divided into two groups: the left group consisting of blocks 0 and 1, and the upper group consisting of blocks 2, 3, and 4. For example, the video decoder 300 may divide the adjacent blocks into the left group and the upper group. For each group, potential candidates among the adjacent blocks that reference the same reference picture as the reference picture indicated by the signaled reference index have the highest priority when selected to form the final candidate for the group. It is possible that not all adjacent blocks contain motion vectors pointing to the same reference picture. Therefore, if no such candidate can be found, the difference in time distance can be compensated for by scaling the first available candidate to form the final candidate.

[0130] Time-motion vector prediction in HEVC is discussed here. The video decoder 300 may add time-motion vector predictor (TMVP) candidates to the MV candidate list after any spatial motion vector candidates, if valid and available. The motion vector derivation process for TMVP candidates is the same for both merge mode and AMVP mode. However, the target reference index for TMVP candidates in merge mode may always be set to 0.

[0131] Figures 6A and 6B are conceptual diagrams illustrating exemplary TMVP candidates and MV scaling, respectively. The primary block location for TMVP candidate derivation is the lower right block outside the same-position PU, shown in Figure 6A as block "T" 600, to compensate for biases to the upper and left blocks used to generate spatial adjacency candidates. However, if the block is currently located outside the row of the CTB (shown as block 602), or if motion information is unavailable, the block is replaced by the central block 604 of PU0 606.

[0132] The motion vector for the TMVP candidate is derived from the co-position PU of the co-position picture, shown at the slice level. The motion vector of the co-position PU is called the co-position MV.

[0133] Similar to the time-direct mode in AVC, in order to derive TMVP candidates, the same-position MV610 needs to be scaled to compensate for the time-distance difference, as shown in Figure 6B. For example, the video decoder 300 may scale the same-position MV610 to compensate for the time-distance difference.

[0134] Other aspects of motion prediction in HEVC are discussed here. Several aspects of merge mode and AMVP mode are worth mentioning as follows:

[0135] Scaling of Motion Vectors: The value of an MV is proportional to the distance between pictures at presentation time. An MV associates two pictures: a reference picture and a picture containing the motion vector (e.g., a containing picture or a picture containing a block predicted using the motion vector). When an MV is used to predict another MV, the time distance between the containing picture and the reference picture is calculated based on the picture order count (POC) value.

[0136] For a motion vector to be predicted, both the associated containing picture and the reference picture may be different. Therefore, a new distance (based on POC) is calculated. The MV is scaled based on these two POC distances. For example, the video decoder 300 may calculate the new distance based on POC, or it may scale the MV based on two POC distances. For a spatially adjacent candidate, the containing pictures of the two MVs may be the same, but the reference pictures may be different. In HEVC, MV scaling is applied to both TMVP and AMVP for spatially adjacent candidates and temporally adjacent candidates.

[0137] Generation of artificial motion vector candidates: If the MV candidate list is incomplete, artificial MV candidates may be generated and inserted at the end of the list until the list contains all candidates (for example, the list is filled).

[0138] In merge mode, there are two types of artificial MV candidates: synthetic candidates derived solely for B slices, and zero candidates used solely for AMVP when the first type does not provide sufficient artificial candidates.

[0139] For each pair of candidates already present in the candidate list and possessing the necessary motion information, a bidirectional composite MV candidate is derived by combining the MV of the first candidate that references a picture in List 0 with the MV of the second candidate that references a picture in List 1.

[0140] Pruning process for candidate insertion: Candidates from different blocks may coincidentally be the same, which reduces the efficiency of the merge / AMVP candidate list. A pruning process may be applied to solve this problem. During the pruning process, the video decoder 300 compares a candidate with other candidates in the current candidate list to some extent to avoid inserting identical candidates. To reduce complexity, the pruning process may be applied to a limited number of candidates rather than comparing each possible candidate with all other existing candidates.

[0141] Template matching prediction is discussed here. Template matching (TM) prediction is a special merge mode based on the Frame-Rate Up Conversion (FRUC) technique. In this TM prediction mode, block motion information is not signaled but is derived on the decoder side by the video decoder 300. TM prediction is applied to both AMVP mode and normal merge mode. In AMVP mode, the selection of MVP candidates is determined using basic template matching to select the candidate with the smallest difference between the current block template and the reference block template. In normal merge mode, the video encoder 200 signals a TM mode flag to indicate the use of TM, and then TM is applied to merge candidates indicated by the merge index for MV improvement.

[0142] Figure 7 is a conceptual diagram illustrating exemplary template matching in the search area around the initial MV. As shown in Figure 7, template matching may be used to derive motion information for the current CU. Deriving motion information may involve finding the best match between template 700 in the current picture 702 (the adjacent block above and / or to the left of the current CU) and block 704 in the reference picture 706 (for example, the same size as the template). Using AMVP candidates selected based on the initial matching error, the candidate MVP is refined by template matching. Using merge candidates indicated by signaled merge indices, the merged MVs of candidates corresponding to reference picture list0 (L0) and reference picture list1 (L1) are refined independently by template matching, and then the less accurate MVs are further refined again using more accurate MVs as previous references. For example, the video decoder 300 may receive and parse signaled merge indices and apply template matching to the merged MVs to refine them.

[0143] Cost function: When the motion vector points to a non-integer sample position, the video decoder 300 may use motion-compensated interpolation. To reduce complexity, bilinear interpolation is used for both template matching and generating the template on the reference picture, instead of the usual 8-tap discrete cosine transform-interpolation filter (DCT-IF) interpolation. The matching cost C for template matching can be calculated as follows: C=SAD+w·(|MV x -MV x s |+|MV y -MV y s |) Here, w is a weighting coefficient that is empirically set to 4, and MV and MV s These represent the current MV under test and the initial MV (e.g., the MVP candidate in AMVP mode or the merged motion vector in merge mode), respectively. The absolute difference sum (SAD) can be used as the matching cost for template matching.

[0144] When motion correction (TM) is used, motion is improved by using only luma samples. The derived motion is used for both luma and chroma for motion compensation (MC) interpretation. After the motion video (MV) is determined, the final MC is performed using an 8-tap interpolation filter for luma and a 4-tap interpolation filter for chroma. For example, the video decoder 300 may improve motion using only luma samples.

[0145] Search Method: MV refinement can be a pattern-based MV search using a template matching cost criterion. Two search patterns are supported for MV refinement: diamond search and cross search. For example, the video decoder 300 may use either diamond search or cross search for MV refinement. The MV is searched directly using the diamond pattern with accuracy of a quarter-lumasample motion vector difference (MVD), then using the cross pattern with accuracy of a quarter-lumasample MVD, and subsequently refined to an eighth-lumasample MVD using the cross pattern. The search range for MV refinement can be set to equal (-8, +8) lumasamples around the initial MV.

[0146] Bilateral matching prediction is discussed here. Bilateral matching (also known as bilateral merge) (BM) prediction is another merge mode based on the FRUC technique. Once a decision is made to apply the BM mode for a given block, two initial MVs (MV0 and MV1) are derived by using signaled merge candidate indices to select merge candidates from a constructed merge list. The video decoder 300 can perform a bilateral matching search around MV0 and MV1 and derive the final MV0' and MV1' based on the minimum bilateral matching cost.

[0147] Figures 8A and 8B are conceptual diagrams illustrating examples where MVD0 and MVD1 are proportional based on time distance, and examples where MVD0 and MVD1 are mirror images regardless of time distance, respectively. The motion vector differences MVD0 (denoted by MV0'-MV0) and MVD1 (denoted by MV1'-MV1) pointing to two reference blocks can now be proportional to the time distance (TD) between picture 804 and the two reference pictures 806 and 808, for example, TD0 800 and TD1 802. Figure 8A shows an example of MVD0 and MVD1, where TD1 802 is four times TD0 800.

[0148] However, there is an optional design in which MVD0 and MVD1 are mirror images of each other regardless of the time distances TD0 and TD1. Figure 8B shows an example of mirror images of MVD0 and MVD1, where TD1 812 is four times TD0 810.

[0149] Figure 9 is a conceptual diagram showing an example of a 3x3 square search pattern in the search range [-8,8]. Bilateral matching may involve performing a local search around the initial MV0 and MV1 to derive the final MV0' and MV1'. To apply the local search, the video decoder 300 may apply a 3x3 square search pattern, looping through the search range [-8,8]. In each iteration of the search, the bilateral matching costs of the eight surrounding MVs in the search pattern are calculated and compared to the bilateral matching cost of the central MV. The MV with the smallest bilateral matching cost becomes the new central MV in the next iteration of the search. The local search terminates when the current central MV has the smallest cost within the 3x3 square search pattern, or when the local search reaches a predetermined maximum number of iterations. For example, the video decoder 300 may perform bilateral matching as described herein. In the example in Figure 9, the initial MV900 is used, and a 3x3 search pattern 902 is explored around the initial MV900. In the first iteration, the MV with the lowest cost among the initial 8 MVs is MV904. In the second iteration, the video decoder 300 then repeats the search pattern 902 around MV904. In this example, the MV that is finally selected after N iterations is MV906.

[0150] Figure 10 is a conceptual diagram illustrating an exemplary decoder-side motion vector improvement. To improve the accuracy of the merge-mode MV, decoder-side motion vector improvement (DMVR) may be applied as shown in VVD Draft 10. For example, video decoder 300 may apply DMVR. In the biprediction operation, improved MVs are searched around the initial MV in reference picture list0 (L0) and reference picture list1 (L1). The DMVR method calculates the strain between two candidate blocks in L0 and L1. For example, video decoder 300 may calculate the strain between two candidate blocks. As shown in Figure 10, the SAD between blocks 1000 and 1002 is calculated based on each MV candidate around the initial MV. For example, video decoder 300 may determine the SAD between blocks 1000 and 1002. The MV candidate with the lowest SAD becomes the improved MV and is used to generate the bipredicted signal.

[0151] The improved MV derived by the DMVR technique is used to generate interprediction samples and is also used in time motion vector prediction for coding future pictures. The video decoder 300 may use the original MV in the deblocking process and in spatial motion vector prediction for coding future CUs.

[0152] The DMVR in VVC Draft 10 is a subblock-based merge mode that uses a predetermined maximum PU of 16x16 lumasamples. When the width and / or height of a CU is greater than 16 lumasamples, the CU may be further divided into subblocks with widths and / or heights equal to 16 lumasamples. For example, the video decoder 300 may further divide a larger CU into subblocks with widths and / or heights equal to 16 lumasamples.

[0153] An exemplary search scheme is discussed here. In DVMR, the search points and MV offsets around the initial MV follow the mirroring rule of the MV difference discussed above. In other words, any point identified by the video decoder 300 implementing the DMVR, denoted by a candidate MV pair (MV0, MV1), follows the following two equations: MV0' = MV0 + MV_offset MV1' = MV1 - MV_offset Here, MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer samples from the initial MV. This search includes an integer sample offset search stage and a non-integer sample refinement stage.

[0154] For integer sample offset search, a full 25-point search may be applied. For example, the video decoder 300 may perform a full 25-point search. The SAD of the initial MV pair is calculated first. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of the DMVR ends. Otherwise, the SADs of the remaining 24 points are calculated and verified in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To reduce the disadvantage of uncertainty in DMVR refinement, the original MV may be preferred during the DMVR process. The SAD between reference blocks referenced by the initial MV candidates may be reduced by 1 / 4 of the SAD value.

[0155] An integer sample search is followed by a non-integer sample refinement. For example, the video decoder 300 may perform an integer sample search and then a non-integer sample refinement. To reduce computational complexity, the non-integer sample refinement may be derived by using a parametric error surface equation rather than an additional search using SAD comparison. The non-integer sample refinement is conditionally invoked based on the output of the integer sample search stage. If the integer sample search stage ends with the center having the smallest SAD in either the first or second iteration of the search, the non-integer sample refinement is further applied.

[0156] In parametric error surface-based sub-pixel offset estimation, the center position cost and the costs at four adjacent positions from the center are used to fit a two-dimensional (2-D) parabolic error surface equation of the following form. E(x,y)=A(x - x min ) 2 +B(y - y min ) 2 +C where (x min ,y min ) corresponds to the non-integer position where the cost is minimum, A and B are constants, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, the position of the minimum value (x min ,y min ) is calculated as follows. x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0)))

[0157] Since all cost values are integers and the minimum value is E(0,0), the values of x min and y min are automatically constrained to be between -8 and 8. This corresponds to a half-pel offset when using the accuracy of 1 / 16 pel MV in VVC Draft 10. To obtain the accurate refined delta MV of sub-pixels, the calculated non-integer (x min , y min ]>) is added to the integer-distance refined MV.

[0158] Bilinear interpolation and sample padding are discussed here. These techniques can be applied by the video decoder 300. In VVC Draft 10, the resolution of the MV is 1 / 16 luma samples. Samples at non-integer positions can be interpolated using an 8-tap interpolation filter. In DMVR, the search points surround the initial non-integer phase MV with an offset of integer samples. Therefore, samples at those non-integer positions need to be interpolated for DMVR search. To reduce computational complexity, a bilinear interpolation filter is used to generate non-integer samples for search in DMVR. Another effect is that by using a bilinear filter with a search range of 2 samples, DVMR accesses fewer reference samples compared to a normal motion compensation process. After the improved MV is obtained using DMVR search, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than a normal motion compensation process, samples that are not needed for the interpolation process based on the original MV but are needed for the interpolation process based on the improved MV are padded from the available samples.

[0159] Exemplary activation conditions for DMVR are discussed here. For example, DMVR may be enabled if all of the following conditions are met: 1) CU level merge mode is used with bi-prediction MV, 2) with respect to the current picture, one reference picture is in the past and another reference picture is in the future, 3) the distance from both reference pictures to the current picture (e.g., POC difference) is the same, 4) CU has more than 64 luma samples, 5) both the height and width of the CU are greater than or equal to 8 luma samples, 6) Bi-prediction with CU weights (BCW) weight indices show equal weights, 7) Weighted prediction (WP) is not enabled for the current block, and 8) Combined inter-intra prediction (CIIP) mode is not used for the current block.

[0160] Bidirectional optical flow is discussed here. The video decoder 300 may use bidirectional optical flow (BDOF) to improve the dual prediction signals of luma samples in the CU at the 4x4 subblock level. As its name suggests, the BDOF mode is based on the concept of optical flow, which assumes that the motion of the object is smooth. For each 4x4 subblock, motion improvement (v) is achieved by minimizing the difference between the L0 prediction sample and the L1 prediction sample. x ,v y The following is calculated. Motion improvements are then used to adjust the predicted sample values ​​within the 4x4 subblock. The following steps are applied in the BDOF process.

[0161] First, the horizontal and vertical gradients of the two prediction signals,

number

number

number

number

[0162] Next, the autocorrelation and crosscorrelation of gradients S1, S2, S3, S5, and S6 are calculated as follows: S1 = Σ (i,j)∈Ω │ψ x (i,j)│, S3=Σ (i,j)∈Ω θ(i,j)·(-sign(ψx (i,j))) S2 = Σ (i,j)∈Ω ψ x (i,j)·sign(ψ y (i,j)) S5=Σ (i,j)∈Ω │ψ y (i,j)│ S6=Σ (i,j)∈Ω θ(i,j)·(-sign(ψ y (i,j)))

[0163] Here

number

number

[0164] Movement improvements (v x ,v y ) is then derived using the terms of cross-correlation and autocorrelation, as follows:

number

number

number

number

[0165] Based on motion improvements and gradients, the following adjustments are calculated for each sample in the 4x4 subblock.

number

[0166] Here, shift5 is Max(3,15-BitDepth), and the variable o offset This is set to equal to (1 << (shift5-1)).

[0167] These values ​​are selected so that the multiplier in the BDOF process does not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits.

[0168] Figure 11 is a conceptual diagram showing an exemplary extended CU region used in BDOF. To derive the gradient values, several prediction samples I from the list k (k=0,1) are currently outside the CU boundary. (k) (i,j) may need to be generated. As shown in Figure 11, BDOF uses one extended row / column around the boundary of CU1100. To control the computational complexity of generating out-of-bounds prediction samples, prediction samples within the extended area (e.g., outermost positions) are generated by directly taking reference samples at nearby integer positions without interpolation (using the floor() operation on the coordinates), and a standard 8-tap motion-compensated interpolation filter is used to generate prediction samples within CU1100 (e.g., diagonal or patterned positions within CU1100). These extended sample values ​​may only be used in gradient calculations. If any sample values ​​and gradient values ​​outside the boundary of CU1100 are needed for the remaining steps in the BDOF process, they may be padded (e.g., iterated) from the nearest neighbor.

[0169] BDOF is used to improve the biprediction signal of a CU at the 4x4 subblock level (e.g., subblock 1102). In one example, BDOF may be applied to a CU if it satisfies all of the following conditions: 1) the CU is coded using a "true" biprediction mode, e.g., one of two reference pictures is before the current picture in display order and the other is after the current picture in display order; 2) the CU is not coded using affine mode or ATMVP merge mode; 3) the CU has more than 64 luma samples; 4) both the height and width of the CU are greater than or equal to 8 luma samples; 5) the BCW weight index shows equal weights; 6) WP is not currently valid for the CU; and 7) CIIP mode is not currently used for the CU.

[0170] In VVC Draft 10, the DMVR is subblock-based and includes up to 16 × 16 luma samples. The improved MV for each subblock has a delta MV (Δhor, Δver) from the original MV. Δhor and Δver are the motion vector offsets in the horizontal and vertical directions, respectively. The range of values ​​for Δhor and Δver is determined by the search range of the DMVR. In VVC Draft 10, the search range of the DMVR is [-2, 2]. Therefore, the improved motion vectors have an offset of up to ±2 pel from the original MV in both the horizontal and vertical directions.

[0171] The range of ±2 Pel values ​​for DeltaMV may be too small for some blocks. For blocks with the best DeltaMV outside the range of ±2 Pel values ​​for DeltaMV, the video decoder 300 cannot derive an optimal improved MV using a DMVR with such a range of values.

[0172] The range of the delta MV value can be broadened by widening the DMVR search range. For example, the DMVR search range can be widened to [-8,8]. Thus, the improved motion vector will have an offset of up to ±8 Pels from the original MV in both the horizontal and vertical directions.

[0173] However, expanding the search range increases the complexity of the DMVR process. For example, when expanding to a fixed search range [-8,8], the video decoder 300 needs to perform more than 11 times more DMVR searches compared to a search range of [-2,2] for the coded block of the DMVR. In addition, even if the derived subsets of MVs are similar or identical, subsets of subblocks within the coded block of the DMVR may have similar improved MVs, and the subblock-based DMVR process includes MV improvements for each subblock. On the other hand, a subarea of ​​a subblock may have a different optimal improved MV than other subareas of the subblock. Since the DMVR of VVC Draft 10 is 16×16 luma sample subblock-based, the video decoder 300 cannot derive different improved MVs in, for example, 8×8 or 4×4 subareas within a 16×16 subblock.

[0174] Techniques that can improve the DMVR process are disclosed herein.

[0175] Example 1. In this example, the improved motion vectors of the subblocks within the W×H coding block are derived by a Multi-Pass DMVR process. A given number N may represent the total number of passes in the Multi-Pass DMVR technique. The video decoder 300 may utilize these Multi-Pass DMVR techniques.

[0176] Figure 12 is a conceptual diagram illustrating an exemplary 3-pass DMVR technique. In this example, 32 × 16 coding blocks 1200 represent the initial MV orgIt begins with. The first pass can be block-based. Therefore, the first pass can currently use an entire block 1200 such as PU or CU. The video decoder 300 that implements the first pass 1202 is an improved MV. pass1 The second pass can be subblock-based. In this example, the video decoder 300 can divide block 1200 into two 16x16 subblocks, subblocks 1204A to 1204B. The video decoder 300 implementing the second pass can generate subblock 1204A (MV (pass2,0) ) and subblock 1204B (MV (pass2,1) An improved MV can be generated for each of the subblocks 1208A to 1208G. In this example, the video decoder 300 may divide block 1200 into eight 8x8 subblocks, subblocks 1208A to 1208H. As shown, the video decoder 300 implementing the third pass may generate an improved MV for each of the subblocks 1208A to 1208G.

[0177] For example, the video decoder 300 may apply a multipath DMVR to the motion vector for a block of video data (e.g., block 1200) to determine an improved motion vector, and decode the block based on the improved motion vector. The multipath DMVR may include a first pass that is block-based and applied to a block of video data; a second pass that is subblock-based and applied to at least one second pass subblock of the block of video data, wherein the width of the second pass subblock is less than or equal to the width of the first pass block and the height of the second pass subblock is less than or equal to the height of the first pass block; and a third pass that is subblock-based and applied to at least one third pass subblock, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0178] The multi-pass DMVR technique uses the original motion vector MV of the W×H coding block. orgIt starts with . The coding block can be PU or CU. The first pass can be block-based. The first pass is an improved motion vector MV for the entire W×H coding block. pass1 This can be derived. MV pass1 This can be saved and used as an initial motion vector for subsequent passes.

[0179] The second pass may be subblock-based, for example, based on one or more subblocks of a W×H coding block. A subblock in the second pass (second pass SB) may have a predetermined maximum dimension sbW_1×sbH_1. A W×H coding block may be divided into K1 subblocks (second pass SB) such that K1≧1. Each second pass SB may have dimensions M1×N1 such that M1≦W and N1≦H. Each second pass SB has an initial motion vector MV pass1 The second pass may have (for example, the MV derived from the first pass). The second pass may have an improved motion vector MV for each second pass SB. (pass2,i) We can also derive this, where i represents the index of the second pass SB, and 0 ≤ i ≤ K1-1. MV (pass2,i) This can be saved and used as an initial motion vector for subsequent passes.

[0180] The third path may be subblock-based, for example, based on one or more subblocks of each subblock of the second path. A subblock in the third path (third path SB) has a predetermined maximum dimension sbW_2 × sbH_2, where sbW_2 ≤ sbW_1 and sbH_2 ≤ sbH_1. Each i-th second path SB in the second path may be divided into K2 subblocks (third path SBs), where K2 ≥ 1. The total number of third path SBs in a W × H coding block may be K2 * K1. Each third path SB may have dimension M2 × N2, where M2 ≤ sbW_1 and N2 ≤ sbH_1. Each third path SB in the i-th second path SB has an initial motion vector MV (pass2,i)(For example, the MV derived during the second pass) The third pass may have an improved motion vector MV for each third pass SB. (pass3,j) We derive the following, where j represents the index of the third pass SB, and 0 ≤ j ≤ K2 * K1 - 1. MV (pass3,j) This can be saved and used as an initial motion vector for subsequent passes.

[0181] In some examples, the multi-pass DMVR technique continues up to the pth pass. The video decoder 300, which performs the MV improvement, performs the MV for each subblock in the pth pass (pth pass SB). (passP,i) It is also possible to derive the following, where i represents the index of the first P path SB within the W×H coding block. MV (passP,i) This can be saved and used to derive the predicted block for the current coding block. MV (passP,i) This represents an improved MV for the i-th subblock.

[0182] Example 2. As in Example 1, when both the p-th pass and the preceding pass (the (p-1)th pass) of the DMVR technique are subblock-based, the dimensions of the p-th pass subblock may be less than or equal to the dimensions of the subblock in the preceding pass.

[0183] As in Example 1, the range of values ​​for the delta motion vector MV(Δhor, Δver) in the p-th path may be predetermined. For example, minDeltaHorPassP ≤ Δhor ≤ maxDeltaHorPassP and minDeltaVerPassP ≤ Δver ≤ maxDeltaVerPassP. When the p-th path is not the first path (for example, p > 1), the range of values ​​for Δhor and Δver in the p-th path may be less than or equal to the range of values ​​in the preceding path. For example, minDeltaHorPassP ≥ minDeltaHorPass(P-1), maxDeltaHorPassP ≤ maxDeltaHorPass(P-1), minDeltaVerPassP ≥ minDeltaVerPass(P-1), and maxDeltaVerPassP ≤ maxDeltaVerPass(P-1). Since the p-th pass can start from the improved motion vector of the preceding pass, the overall range of values ​​for the delta (final improved) motion vector is extended compared to a single-pass DMVR.

[0184] As in Example 1, when the video decoder 300 decides to divide the current coding block into K subblocks, the subblocks can be in the raster scan order of the current coding block, from the top left to the bottom right.

[0185] Example 3 - Skipping the p-th pass of the DMVR technique. As in Example 1, a given number N may represent the total number of passes in the multi-pass DMVR technique. A video decoder 300 implementing the multi-pass DMVR technique may skip one or more passes to derive the final improved MV. In other words, the video decoder 300 may derive the final improved motion vector by applying a subset of the multi-pass DMVR technique. Skipping the p-th pass of the DMVR technique may reduce the complexity of the video decoder 300.

[0186] The decision of whether to skip the p-th pass in the DMVR technique can be based on the results of the preceding passes in the DMVR technique. For example, if the preceding passes derive a relatively optimal improved motion vector, the p-th pass may be skipped.

[0187] For example, the video decoder 300 may apply a shortened multipath DMVR to the motion vector for a block. The video decoder 300 may decide to skip a given path of the multipath DMVR for a block and, based on the decision to skip the given path, may skip a given path of the multipath DMVR for a block. For example, the decision to skip a given path may be based on the result of the preceding path, such as when the improved MV of the preceding path is relatively optimal (for example, further improvements may not result in any change to the MV in terms of MV (sub-per) resolution, or the cost of further improvements may outweigh the benefits of further improvements).

[0188] Example 4 - Subblock-based first-pass DMVR technique. In some hardware designs, the maximum size for the motion compensation process may be constrained, and larger coding blocks may be divided into multiple subblocks for hardware processing. In some examples, the multi-pass DMVR technique may start with a subblock size of min{P,W} × min{Q,H} for the first pass, where P and Q are predetermined integer values ​​determined by hardware constraints.

[0189] As in Examples 1 and 3, the first pass of the DMVR technique can be block-based. When the multi-pass DMVR technique starts with a sub-block-based pass, the first pass of the DMVR technique may also be known as the sub-block-based first-pass DMVR technique or the skip first-pass DMVR technique. The video decoder 300 may apply the sub-block-based first-pass DMVR technique.

[0190] Example 5 - Skip the p-th pass of the DMVR technique for a sub-area of a coding block. Assuming a W×H coding block, as in Examples 1 and 3, a given number N can represent the total number of passes of the multi-pass DMVR technique. The video decoder 300 can derive improved motion vectors for a sub-area of the coding block by applying the N passes of the DMVR technique. The video decoder 300 may also derive improved motion vectors for different sub-areas of the coding block by applying M passes of the DMVR technique, where M < N. In other words, the video decoder 300 may skip one or more passes to derive the final improved motion vector for a given sub-area of the coding block. The sub-area may include one or more sub-blocks of the coding block.

[0191] For example, the video decoder 300 applies a shortened multi-pass DMVR to the motion vectors of a block. The video decoder 300 determines to skip a given sub-block-based pass of the multi-pass DMVR for a particular sub-area of the block (e.g., different sub-areas of the coding block referred to in the previous paragraph) that includes one or more sub-blocks, and may skip the given sub-block-based pass of the multi-pass DMVR for the particular sub-area based on the determination to skip the given sub-block-based pass. For example, the determination to skip a given sub-block-based pass may be based on the results of the previous pass when the improved MV of the previous pass is relatively optimal (e.g., further improvement may not result in a change to the MV in terms of MV (sub-pel) accuracy, or the cost of further improvement may exceed the benefit of further improvement).

[0192] Example 6 - Deriving an improved motion vector in the p-th pass of a DMVR. This example illustrates several decoder-side motion vector improvement techniques. In a multi-pass DMVR technique, the video decoder 300 may apply a bilateral matching-based motion vector improvement, as discussed below, through at least one pass, and / or apply a BDOF-based motion vector improvement through at least one pass. In other words, at least one pass of a multi-pass DMVR may include applying a BDOF, and / or at least one pass of a multi-pass DMVR may include applying bilateral matching. In one example, the first pass includes applying bilateral matching, the second pass includes applying bilateral matching, and the third pass includes applying a BDOF.

[0193] Figure 13 is a conceptual diagram illustrating an exemplary BDOF motion vector improvement. The derivation of an improved motion vector by bidirectional optical flow is described here. In this example, the video decoder 300 can derive an improved motion vector in the p-th pass DMVR technique by using bidirectional optical flow (BDOF). Improvements to BDOF MV may include the following: Mv0'=Mv0+bioMv Mv1'=Mv1-bioMv Here, Mv0 and Mv1 represent the initial Mv of the p-th path in the current block / subblock within reference picture 0 1300 and reference picture 1 1302, respectively; Mv0' and Mv1' represent the BDOF-enhanced MV of the current block within reference picture 0 1300 and reference picture 1 1302, respectively; and bioMv is the BDOF delta MV.

[0194] In the BDOF MV improvement process, bioMv(Δhor,Δver) can be derived from the following steps. 1) As discussed above, from the prediction signals predSig0 and predSig1, the horizontal and vertical gradients,

number

number

number

number

[0195] A video decoder 300 that derives improved motion vectors by bilateral matching is described here. Bilateral matching involves searching around two initial motion vectors MV0 and MV1 in the p-th path in a given local search area in reference picture 0 and reference picture 1, respectively. The final MV0' and MV1' are derived based on the minimum bilateral matching cost.

[0196] The local search area for bilateral matching for a coding block has a horizontal search range, e.g., [sMinHor, sMaxHor] and a vertical search range, e.g., [sMinVer, sMaxVer]. The local search area for bilateral matching for a coding block may be (sMaxHor - sMinHor + 1) × (sMaxVer - sMinVer + 1).

[0197] As in Example 2, if there is a predetermined range of values ​​for the delta motion vector MV(Δhor,Δver) in the p-th pass, the value of the search range can be determined by the range of values ​​for the delta motion vector in the p-th pass DMVR technique, as follows: sMinHor≧minDeltaHorPassP sMaxHor≦maxDeltaHorPassP sMinVer≧minDeltaVerPassP sMaxVer≦maxDeltaVerPassP

[0198] Further decoder-side motion vector refinement methods are described herein. The refined motion vectors may be derived by alternative decoder-side motion vector derivation techniques, such as template matching or decoder-side motion vector derivation (DMVD). A video decoder 300 implementing a p-pass multi-pass DMVR technique may use one of these motion vector refinement methods described herein. However, the details of the DMVR technique may differ from those described herein, and may still be within the scope of this disclosure.

[0199] Example 7 - Deriving the prediction signal in the p-th pass DMVR technique by applying an interpolation filter or by using the prediction signal from a preceding pass. As in Example 6, the motion vector improvement technique in the p-th pass starts with the initial motion vector in the p-th pass and the prediction signal in the reference picture. The prediction signal in the reference picture can be derived by applying an interpolation filter along with the initial motion vector information in the reference picture.

[0200] In this example, the video decoder 300 is 1) The predicted signal is derived using the p-th pass DMVR technique by applying an interpolation filter. The interpolation filter may be determined in the p-th pass by an MV improvement technique (e.g., bilateral matching, BDOF, etc.), and / or 2) By using the preceding prediction signal, the prediction signal is derived using the p-pass DMVR technique.

[0201] A video decoder 300 is described here that derives a predicted signal using the p-th pass DMVR technique by applying an interpolation filter. In the DMVR technique for bilateral matching or improved motion vector derivation, some simplified interpolation filter may be used to generate motion compensation results for the search. For example, a bilinear interpolation filter may be used to generate non-integer samples for the search process in bilateral matching or DMVR.

[0202] In some examples, when applying BDOF-based techniques, as in Example 6, the input may be samples generated by motion compensation using the original (unsimplified) interpolation filter in order to derive an improved motion vector in the p-th pass.

[0203] In other examples, when applying BDOF-based techniques, as in Example 6, the input may be samples generated by motion compensation using a simplified interpolation filter, such as a bilinear interpolation filter, in order to derive an improved motion vector in the p-th pass.

[0204] A video decoder 300 is described here that derives a prediction signal in the p-th pass DMVR technique by using the prediction signal of the preceding pass. For example, the video decoder 300 may decide whether to use the prediction signal of the preceding pass by checking the accuracy of the delta motion vector in the preceding pass. For example, 1) When the delta motion vector in the preceding path has integer Pell precision, the prediction signal in the p-th path of the DMVR technique may be derived by using the prediction signal of the preceding path, and 2) When the improved motion vector in the preceding path is identical to the initial motion vector in the preceding path, the prediction signal in the p-th path of the DMVR technique may be derived by using the prediction signal of the preceding path.

[0205] Example 8 - Example of 3-pass decoder-side motion improvement. In this example, the video decoder 300 uses a 3-pass decoder-side motion improvement technique. In this example, the process involves three passes as follows: 1) The first pass is block-based. The improved motion vector is derived by applying bilateral matching-based motion vector improvement. The delta motion value range is, for example, [-8,8] horizontally and, for example, [-8,8] vertically. 2) The second pass is sub-block-based. The improved motion vector is derived by applying bilateral matching-based motion vector improvement. The maximum sub-block dimension is, for example, 16 × 16 luma samples. For example, the sub-block of the second pass has a predetermined maximum width of 16 luma samples and a predetermined maximum height of 16 luma samples. The delta motion value range is, for example, [-8,8] horizontally and, for example, [-8,8] vertically. 3) The third pass is sub-block-based. The improved motion vector is derived by applying BDOF-based motion vector improvements. The maximum subblock dimensions are, for example, 8 x 8 luma samples. For example, the subblock of the third pass has a predetermined maximum width of 8 luma samples and a predetermined maximum height of 8 luma samples. The delta motion value range is, for example, [-2,2] in the horizontal direction and, for example, [-2,2] in the vertical direction.

[0206] For example, the delta motion range for at least one of the first or second path may be [-8,8] horizontally and [-8,8] vertically, and the delta motion range for the third path may be [-2,2] horizontally and [-2,2] vertically.

[0207] The aforementioned techniques can be applied by the video decoder 300 of the video coding system. The following is a detailed example of multipath DMVR. The video decoder 300 may implement the techniques described herein by all or a subset of the following steps in order to decode interpredicted blocks in a picture from a bitstream. 1) The position component (x,y) is derived as the top-left Luma position of the current block by decoding the syntax element in the bitstream. 2) The size of the current block is derived as width W and height H by decoding the syntax elements in the bitstream. 3) By decoding the elements in the bitstream, we determine that the current block is the block that can be interpreted. 4) By decoding the elements in the bitstream, the motion vector components (mvL0 and mvL1) and reference indices (refPicL0 and refPicL1) of the current block are derived. 5) A flag is inferred from decoding the elements in the bitstream, and the flag indicates whether a decoder-side motion vector derivation (e.g., DMVR, bilateral merge, template matching, etc.) is currently applied to the block. The flag inference method may be, but is not limited to, the same as the DMVR activation conditions discussed earlier in this disclosure. In another example, this flag may be explicitly signaled in the bitstream to avoid complex conditional checks by the video decoder 300. 6) (Pass 1) If the decision is not to apply DMVR (Bilateral Merge or Template Matching) to the current block, according to the value of the flag mentioned above, then the motion vectors mvL0 and mvL1 are set as the motion vectors for MV0pass1 and MV1pass1, respectively; otherwise (if the decision is to apply DMVR to the current block), then the following applies: (a) Set the mvL0 and mvL1 of the current block as the initial motion vectors for the current block. (b) Determine the variables sHor and sVer as follows. sHor = maximum(maxDeltaHorPass1, W × sFactor) sVer = maximum(maxDeltaVerPass1, H × sFactor) Here, maxDeltaHorPass1 is a predetermined variable (e.g., 8). maxDeltaVerPass1 is a predetermined variable (e.g., 8). sFactor is a predetermined variable (e.g., 0.5). sHor is the search range [-sHor, sHor] of DMVR in the horizontal direction. sVer is the search range [-sVer, sVer] of DMVR in the vertical direction. (c) By using the derived mvL0 and refPicL0, derive the prediction signal predSig0 from reference picture 0. The width of predSig0 is equal to W + 2 × sHor. The height of predSig0 is equal to H + 2 × sVer. (d) By using the derived mvL1 and refPicL1, derive the prediction signal predSig1 from reference picture 1. The width of predSig1 is equal to W + 2 × sHor. The height of predSig0 is equal to H + 2 × sVer. (e) Set the variable minCostPass1 to the maximum cost value. (f) Set the variable best delta MV (Δhor_best, Δver_best) to delta MV (0, 0). (g) Loop through each or a subset of the delta MVs (Δhor, Δver) within the search range of the current block. -sVer ≤ Δver ≤ sVer, -sHor ≤ Δhor ≤ sHor. (i) Derive the bilateral matching cost bilCost at the current delta MV (Δhor, Δver). (ii) If bilCost is less than minCostPass1, (a) Set minCostPass1 to be equal to bilCost. (b) Set the best delta MV(Δhor_best,Δver_best) to be equal to MV(Δhor,Δver). (h) Improved motion vector (mvL0+MV(Δhor_best,Δver_best)) to MV0 pass1 This is derived as the motion vector. (i) The improved motion vector (mvL1-MV(Δhor_best,Δver_best)) is given to MV1 pass1 This is derived as the motion vector. 7) (Path 2) The number of subblocks in the horizontal direction numSbX and the number of subblocks in the vertical direction numSbY, the width of the subblocks sbWidthPass2, and the height sbHeightPass2 are derived as follows: numSbX=(W>thW)?(W / thW):1 numSbY=(H>thH)?(H / thH):1 sbWidthPass2=(W>thW)?thW:W sbHeightPass2=(H>thH)?thH:H Here, thW and thH are predetermined integer values ​​that represent the width and height of the largest subblock for the second path, respectively (for example, thW = thH = 16). (a) If the decision is to not apply DMVR (Bilateral Merge or Template Matching) to the current block, according to the value of the flag mentioned above, then the motion vector MV0 pass1 and MV1 pass1 Each of these is the motion vector MV0 for each subblock. (pass2,i) and MV1 (pass2,i) If set as such, and otherwise (if the decision is to apply the DMVR to the current block), then the following applies: (b) (Check whether to skip pass 2) Derive a variable costThPass2 equal to (thFactorPass2 × W × H), where thFactorPass2 is a predetermined value, for example thFactorPass2 = 1. If minCostPass1 is less than costThPass2, then MV0 pass1 and MV1 pass1 Each of these is the motion vector MV0 for each subblock. (pass2, i) and MV1 (pass2, i) If set as such, otherwise (when minCostPass1 is greater than or equal to costThPass2), the following applies: (i) Set the position component (sbX,sbY)=(x,y) as the top-left corner position of the first subblock of the current block. (ii) For each subblock, starting from the top left and going down to the bottom right, (a) Set the variable i = (sbY / sbHeightPass2)*(W / sbWidthPass2)+(sbX / sbWidthPass2) as the current subblock index. (b)MV0 pass1 and MV1 pass1 This is now set as the initial motion vector for the subblock. (c) Determine the variables sHor and sVer as follows: sHor=maximum(maxDeltaHorPass2, sbWidthPass2×sFactor) sVer=maximum(maxDeltaVerPass2, sbHeightPass2×sFactor) Here, maxDeltaHorPass2 is a predetermined variable (for example, 8), and maxDeltaVerPass2 is a predetermined variable (for example, 8). sFactor is a given variable (for example, 0.5). sHor specifies the horizontal search range [-sHor,sHor] for path 2. sVer specifies the vertical search range [-sVer,sVer] for path 2. (d) Derived MV0 pass1 By using refPicL0, the predicted signal predSig0 is derived from reference picture 0. The width of predSig0 is equal to sbWidthPass2 + 2 × sHor. The height of predSig0 is equal to sbHeightPass2 + 2 × sVer. (e) Derived MV1 pass1 By using refPicL1, the predicted signal predSig1 is derived from reference picture 1. The width of predSig1 is equal to sbWidthPass2 + 2 × sHor. The height of predSig0 is equal to sbHeightPass2 + 2 × sVer. (f) Set the variable minCostPass2 to the maximum cost value. (g) Set the variable best delta MV(Δhor_best,Δver_best) to delta MV(0,0). (h) Loop through each or a subset of DeltaMV(Δhor,Δver) within the current subblock's search range, where -sVer≦Δver≦sVer and -sHor≦Δhor≦sHor. (i) Derive the bilateral matching cost bilCost in the current delta MV(Δhor,Δver). (ii) If bilCost is less than minCostPass2, (a) Set minCostPass2 to be equal to bilCost. (b) Set the best delta MV(Δhor_best,Δver_best) to be equal to MV(Δhor,Δver). (i) Improved motion vector (MV0 pass1 +MV(Δhor_best,Δver_best)) to MV0 (pass2,i) This is derived as the motion vector. (j) Improved motion vector (MV1 pass1 -MV(Δhor_best,Δver_best)) to MV1 (pass2,i) This is derived as the motion vector. (k) Update the top-left luma position of the subblock as follows: sbX=(sbX+sbWidthPass2) <W?sbX+sbWidthPass2:0 sbY=(sbX+sbWidthPass2) <W?sbY:sbY+sbHeightPass2 8) Infer a flag from decoding elements in the bitstream, which indicates whether a bidirectional optical flow is currently applied to the block. The method for inferring the flag is not limited, but could be the same as in the example above. In another example, this flag may be explicitly signaled in the bitstream to avoid complex conditional checks in the decoder. 9) (Path 3) When the decision is to apply BDOF to the current block according to the value of the flag mentioned above, the following is true: (a) The number of subblocks in the horizontal direction numSbX and the number of subblocks in the vertical direction numSbY, the width of the subblock sbW, and the height sbH are derived as follows: numSbX=(W>thW)?(W / thW):1 numSbY=(H>thH)?(H / thH):1 sbWidthPass3=(W>thW)?thW:W sbHeightPass3=(H>thH)?thH:H Here, thW and thH are predetermined integer values ​​that represent the width and height of the largest subblock for the third path, respectively (for example, thW = thH = 8). (b) Derive a variable costThPass3 that is equal to (thFactorPass3 × sbWidth × sbHeight), where thFactorPass3 is a predetermined value, for example, thFactorPass3 = 32. (c) Set the position component (sbX,sbY)=(x,y) as the top-left corner position of the first subblock of the current block. (d) For each subblock, starting from the top left and moving to the bottom right, (i) Set the variable i = (sbY / sbHeightPass3) * (W / sbWidthPass3) + (sbX / sbWidthPass3) as the current sub-block index of Pass 3. (ii) Set the variable j = (sbY / sbHeightPass2) * (W / sbWidthPass2) + (sbX / sbWidthPass2) as the current sub-block index of Pass 2. (iii) MV0 (pass2,j) and MV1 (pass2,j) are set as the initial motion vectors for the current sub-block. (iv) Derive the prediction signal predSig0 from the reference picture 0 by using the derived MV0 (pass2,j) and refPicL0. (v) Derive the prediction signal predSig1 from the reference picture 1 by using the derived MV1 (pass2,j) and refPicL1. (vi) Derive the distortion cost distance between predSig0 and predSig1 of the current sub-block. (vii) If the distortion cost distance (checking whether to skip Sub-Area Pass 3) is less than costThPass3, MV0 (pass2,j) and MV1 (pass2,j) are respectively set as the improved motion vectors MV0 (pass3,i) and MV1 (pass3,i) for the current sub-block; otherwise (if the distortion cost distance is greater than or equal to costThPass3), the following holds. (a) As discussed above, derive the horizontal and vertical gradients

Number

Number

[0208] Example 9 - All passes in the multi-pass DMVR technique are skipped. When all passes in the multi-pass DMVR technique are skipped, the final improved motion vector MV for each subblock in the last pass (path P) (passP,i) This is the initial motion vector MV Org It is equal to.

[0209] For example, as in Example 8, the video decoder 300 may decide whether to apply a BDOF-based motion vector improvement (pass 3) to the current block, depending on the conditions under which a DMVR (e.g., bilateral merge or template matching) is applied to the current block. For example, in step 5 of Example 8 above, the video decoder 300 decides not to apply a DMVR to the current block, and all three passes of the multi-pass DMVR technique are skipped. The improved motion vector MV0 for each subblock in step 10 of Example 8 (pass3,i) and MV1 (pass3,i) These are equal to mvL0 and mvL1, respectively. For example, the video decoder 300 may decide not to apply the DMVR to a block. Based on the decision not to apply the DMVR to a block, the video decoder 300 may skip all paths of the multipath DMVR and decode the block based on the initial motion vector.

[0210] Figure 14 is a flowchart illustrating an exemplary multi-pass DMVR technique of the present disclosure. The video decoder 300 may apply multi-pass DMVR to the MV for blocks of video data to determine an improved MV (1400). For example, the video decoder 300 may apply multi-pass DMVR including a first pass which is block-based; a second pass which is sub-block-based, wherein the width of the second-pass subblock is less than or equal to the width of the first-pass block and the height of the second-pass subblock is less than or equal to the height of the first-pass block; and a third pass which is sub-block-based, wherein the width of the third-pass subblock is less than or equal to the width of the second-pass subblock and the height of the third-pass subblock is less than or equal to the height of the second-pass subblock.

[0211] The video decoder 300 may code blocks based on the improved MV (1402). For example, the video decoder 300 may predict blocks using the improved MV.

[0212] In some examples, at least one third-pass subblock of the video data block is at least one second-pass subblock of the video data block. againstThis is a subblock. In some examples, the video decoder 300 may apply a first pass to derive at least one first improved motion vector for a block of video data and use at least one first improved motion vector in a second pass. For example, the video decoder 300 may use a first improved motion vector as an initial motion vector for the second pass. In some examples, the video decoder 300 may apply a second pass to derive at least one second improved motion vector for at least one second pass subblock and use at least one second improved motion vector in a third pass. For example, the video decoder 300 may derive one or more second improved motion vectors for one or more subblocks of the second pass and use one or more second improved motion vectors as an initial motion vector for the third pass. In some examples, the video decoder 300 may apply a third pass to derive at least one third improved motion vector for at least one third pass subblock, thereby determining at least one improved motion vector as at least one third improved motion vector.

[0213] In some examples, at least one pass of the multipath DMVR includes applying BDOF or applying bilateral matching. In some examples, the first pass includes applying bilateral matching, the second pass includes applying bilateral matching, and the third pass includes applying BDOF.

[0214] In some examples, at least one second-pass subblock has a predetermined maximum width of 16 luma samples and a predetermined maximum height of 16 luma samples. In some examples, at least one third-pass subblock has a predetermined maximum width of 8 luma samples and a predetermined maximum height of 8 luma samples.

[0215] In some examples, the delta movement range for at least one of the first or second paths is [-8,8] horizontally and [-8,8] vertically, and the delta movement range for the third path is [-2,2] horizontally and [-2,2] vertically.

[0216] In some examples, the block of video data is the first block. In some examples, the video decoder 300 may apply a shortened multipath DMVR to the motion vector for the second block of video data. For example, the video decoder 300 may decide to skip a given path of the multipath DMVR for the second block, and based on that decision to skip a given path of the multipath DMVR for the second block, it may skip a given path of the multipath DMVR for the second block. In some examples, the video decoder 300 may decide to skip a given path based on the result of a preceding path.

[0217] In some examples, the video block is the first block. In some examples, the video decoder 300 may apply a shortened multipath DMVR to the motion vector for the second block of video data. For example, the video decoder 300 may decide to skip a given subblock-based path of the multipath DMVR for a particular sub-area of ​​the second block of video data, where the particular sub-area comprises one or more sub-blocks of the second block. For example, the video decoder 300 may skip a given subblock-based path of the multipath DMVR for a particular sub-area of ​​the second block based on its decision to skip a given subblock-based path of the multipath DMVR for a particular sub-area of ​​the second block. In some examples, the video decoder 300 may decide to skip a given path based on the result of a preceding path.

[0218] In some examples, the block is the first block of video data. In some examples, the video decoder 300 may decide not to apply the DMVR to the second block of video data. Based on the decision not to apply the DMVR to the second block, the video decoder 300 may skip all paths of the multipath DMVR for the second block and decode the second block based on the initial motion vector for the second block.

[0219] Figure 15 is a flowchart illustrating an exemplary method for encoding a current block using the technique of the present disclosure. A current block may comprise a current CU. While the video encoder 200 (Figures 1 and 3) is described, it should be understood that other devices may be configured to perform a similar method to that shown in Figure 15.

[0220] In this example, the video encoder 200 first predicts the current block (350). For example, the video encoder 200 may form a predicted block for the current block. The video encoder 200 may then compute the residual block for the current block (352). To compute the residual block, the video encoder 200 may compute the difference between the original unencoded block and the predicted block for the current block. The video encoder 200 may then transform the residual block and quantize the transformation coefficients of the residual block (354). The video encoder 200 may then scan the quantized transformation coefficients of the residual block (356). During or following the scan, the video encoder 200 may entropy encode the transformation coefficients (358). For example, the video encoder 200 may encode the transformation coefficients using CAVLC or CABAC. The video encoder 200 may then output the entropy encoded data of the block (360).

[0221] Figure 16 is a flowchart illustrating an exemplary method for decoding the current block of video data using the technique of the present disclosure. The current block may comprise a current CU. While a video decoder 300 (Figures 1 and 4) is described, it should be understood that other devices may be configured to perform a method similar to that of Figure 16.

[0222] The video decoder 300 may receive entropy-encoded data for the current block, such as entropy-encoded prediction information and entropy-encoded data of the transformation coefficients of the residual block corresponding to the current block (370). The video decoder 300 may entropy-decode the entropy-encoded data to determine the prediction information for the current block and reconstruct the transformation coefficients of the residual block (372). The video decoder 300 may predict the current block to compute a prediction block for the current block, for example, using an intra-prediction mode or an inter-prediction mode as indicated by the prediction information for the current block (374). As part of predicting the current block, the video decoder 300 may use any of the multi-pass DMVR techniques of this disclosure, including, but not limited to, the technique shown in Figure 14. The video decoder 300 may then back-scan the reconstructed transformation coefficients to create a block of quantized transformation coefficients (376). The video decoder 300 may then inverse-quantize the transformation coefficients and apply the inverse transformation to the transformation coefficients to generate a residual block (378). The video decoder 300 can finally decode the current block by combining the predicted block and the residual block (380).

[0223] This disclosure includes the following non-limiting clauses.

[0224] Clause 1A. A method for coding video data, comprising the steps of: applying a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine an improved motion vector; and coding a block based on the improved motion vector.

[0225] Clause 2A. The method according to Clause 1A, wherein the total number of paths in a multipath DMVR is a predetermined integer.

[0226] Clause 3A. The method according to Clause 1A or Clause 2A, wherein the multipath DMVR comprises a first path that is block-based, a second path that is subblock-based, and a third path that is subblock-based.

[0227] Clause 4A. The method according to Clause 3A, wherein the step of applying the first pass derives the first improved motion vector.

[0228] Clause 5A. The method according to Clause 4A, wherein an improved motion vector is used in a second path, the subblocks of the second path have a predetermined maximum width and a predetermined maximum height, and the step of applying the second path derives a second improved motion vector for at least one of each subblocks of the second path.

[0229] Clause 6A. The method according to Clause 5A, wherein each second improved motion vector is used in a third path, the subblocks of the third path have a predetermined maximum width and a predetermined maximum height, and the step of applying the third path derives a third improved motion vector for at least one of each subblocks of the third path.

[0230] Clause 7A. A method of any combination of Clauses 1A through 6A in which the multipath DMVR is iterative.

[0231] Clause 8A. A method according to any combination of Clauses 1A to 7A, wherein the subblocks of a path are less than or equal to the size of a block or subblock of a preceding path.

[0232] Clause 9A. A method of any combination of Clauses 1A to 8A, further comprising the steps of determining whether to skip a given path in a multipath DMVR, and skipping a given path in a multipath DMVR based on the decision to skip a given path.

[0233] Clause 10A. The method according to Clause 9A, wherein the step of deciding whether to skip a given path comprises the step of determining that a given improved motion vector from a preceding path is optimal.

[0234] Clause 11A. A method according to any combination of Clauses 1A to 10A, further comprising the step of determining that a given subblock size is the smaller of (P,W) and the smaller of (Q,H), wherein P and Q are predetermined integers.

[0235] Clause 12A. The method described in Clause 11A, wherein P and Q are based on hardware constraints.

[0236] Clause 13A. The method of any combination of Clauses 1A to 12A, further comprising the steps of determining whether to skip a subblock path of a multipath DMVR for a particular subblock, and skipping a subblock path of a multipath DMVR for a particular subblock based on the decision to skip a subblock path.

[0237] Clause 14A. The method of Clause 13A, wherein the step of determining whether to skip a subblock path comprises the step of determining that the improved motion vector of the subblock from the preceding path is optimal.

[0238] Clause 15A. The method according to any combination of Clauses 1A to 14A, wherein at least one path of a multipath DMVR comprises a step of applying bidirectional optical flow.

[0239] Clause 16A. The method of any combination of Clauses 1A to 15A, wherein at least one path of a multipath DMVR comprises a step of applying bilateral matching.

[0240] Clause 17A. A method of any combination of Clauses 1A to 16A, wherein at least one path of a multipath DMVR comprises a step of applying template matching.

[0241] Clause 18A. The method according to any combination of Clauses 1A to 17A, wherein at least one pass of a multi-pass DMVR comprises a step of applying an interpolation filter.

[0242] Clause 19A. The method according to any combination of Clauses 1A to 18A, wherein at least one pass of a multipass DMVR comprises the step of applying a simplified interpolation filter.

[0243] Clause 20A. The method of any combination of Clauses 1A to 19A, further comprising the steps of determining whether to skip all paths of a multipath DMVT and coding a block based on an initial motion vector based on the decision to skip all paths of a multipath DMVT.

[0244] Clause 21A. The method according to any one of Clauses 1A to 20A, wherein the coding step comprises a decryption step.

[0245] Clause 22A. The method according to any one of Clauses 1A to 21A, wherein the coding step comprises the encoding step.

[0246] Clause 23A. A device for coding video data, comprising one or more means for performing the method described in any of Clauses 1A to 22A.

[0247] Clause 24A. The device according to Clause 23A, wherein one or more means comprises one or more processors implemented in the circuit.

[0248] Clause 25A. The device described in either Clause 23A or 24A, further comprising memory for storing video data.

[0249] Clause 26A. A device as described in any combination of Clauses 23A to 25A, further comprising a display configured to display decoded video data.

[0250] Clause 27A. A device described in any combination of Clauses 23A through 26A, comprising one or more of the following: a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0251] Clause 28A. A device as described in any combination of Clauses 23A to 27A, wherein the device comprises a video decoder.

[0252] Clause 29A. A device as described in any combination of Clauses 23A to 28A, wherein the device comprises a video encoder.

[0253] Clause 30A. A computer-readable storage medium storing instructions, wherein, when the instructions are executed, causes one or more processors to perform the method described in any of Clauses 1A to 22A.

[0254] Clause 1B. A method for decoding video data, A method comprising the steps of: applying a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector; and decoding the block based on at least one improved motion vector, wherein the multipath DMVR comprises: a first pass which is block-based and applied to a block of video data; a second pass which is subblock-based and applied to at least one second pass subblock of the block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass which is subblock-based and applied to at least one third pass subblock of the block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0255] Clause 2B. At least one third-path subblock of a block of video data is at least one second-path subblock of a block of video data. against The method described in subblock, Clause 1B.

[0256] Clause 3B. The method according to Clause 1B or Clause 2B, wherein the step of applying the first pass derives at least one first improved motion vector for a block of video data, and at least one first improved motion vector is used in the second pass.

[0257] Clause 4B. The method according to Clause 3B, wherein the step of applying the second pass derives at least one second improved motion vector for at least one each second pass subblock, and at least one second improved motion vector is used in the third pass.

[0258] Clause 5B. The method according to Clause 4B, wherein the step of applying the third path derives at least one third improved motion vector for at least one each third path subblock, and at least one improved motion vector is determined to be at least one third improved motion vector.

[0259] Clause 6B. A method according to any combination of Clauses 1B to 5B, wherein at least one path of a multipath DMVR comprises a step of applying bidirectional optical flow (BDOF) or a step of applying bilateral matching.

[0260] Clause 7B. The method according to Clause 6B, comprising a first pass applying bilateral matching, a second pass applying bilateral matching, and a third pass applying BDOF.

[0261] Clause 8B. The method according to any combination of Clauses 1B to 7B, wherein at least one second pass subblock has a predetermined maximum width of 16 luma samples and a predetermined maximum height of 16 luma samples.

[0262] Clause 9B. The method according to any combination of Clauses 1B to 8B, wherein at least one third path subblock has a predetermined maximum width of 8 luma samples and a predetermined maximum height of 8 luma samples.

[0263] Clause 10B. The method according to any combination of Clauses 1B to 9B, wherein the delta motion range for at least one of the first or second paths is [-8,8] horizontally and [-8,8] vertically, and the delta motion range for the third path is [-2,2] horizontally and [-2,2] vertically.

[0264] Clause 11B. The method of any combination of Clauses 1B to 10B, wherein a block of video data is a first block, and the method further comprises the step of applying a shortened multipath DMVR to a motion vector for a second block of video data, wherein the applying step comprises the step of determining to skip a given path of the multipath DMVR for the second block, and the step of skipping a given path of the multipath DMVR for the second block based on the determination to skip a given path of the multipath DMVR for the second block.

[0265] Clause 12B. The method according to Clause 11B, wherein the step of deciding to skip a given path is based on the result of a preceding path.

[0266] Clause 13B. The method of any combination of Clauses 1B to 12B, wherein a block of video data is a first block, and the method further comprises the step of applying a multipath DMVR shortened to motion vectors for a second block of video data, wherein the applying step is a step of determining to skip a given subblock-based path of the multipath DMVR for a particular sub-area of ​​the second block of video data, the particular sub-area comprising one or more sub-blocks of the second block, and the step of skipping a given subblock-based path of the multipath DMVR for a particular sub-area of ​​the second block based on the determination to skip a given subblock-based path of the multipath DMVR for a particular sub-area of ​​the second block.

[0267] Clause 14B. The method described in Clause 13B, wherein the step of deciding to skip a given subblock-based path is based on the result of a preceding path.

[0268] Clause 15B. The method of any combination of Clauses 1B to 10B, wherein Block is a first block of video data, and the method further comprises the steps of determining not to apply a DMVR to a second block of video data, skipping all paths of a multipath DMVR for the second block based on the decision not to apply a DMVR to the second block, and decoding the second block based on an initial motion vector for the second block.

[0269] Clause 16B. A device for decoding video data, comprising a memory configured to store video data, and one or more processors implemented by circuitry and communicatively coupled to the memory, wherein one or more processors are configured to apply a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector, and to decode the block based on the at least one improved motion vector, wherein the multipath DMVR comprises a first pass that is block-based and applied to a block of video data, a second pass that is subblock-based and applied to at least one second pass subblock of a block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data, and a third pass that is subblock-based and applied to at least one third pass subblock of a block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0270] Clause 17B. At least one third-path subblock of a block of video data is at least one second-path subblock of a block of video data. against A subblock, the device described in Clause 16B.

[0271] Clause 18B. The device according to Clause 16B or Clause 17B, wherein one or more processors are configured to apply a first pass to derive at least one first improved motion vector for a block of video data and to use at least one first improved motion vector in a second pass.

[0272] Clause 19B. The device according to Clause 18B, wherein one or more processors are configured to apply a second pass to derive at least one second improved motion vector for at least one each second pass subblock and to use at least one second improved motion vector in a third pass.

[0273] Clause 20B. The device according to Clause 19B, wherein one or more processors are configured to apply a third path to derive at least one third improved motion vector for at least one each third path subblock, and to determine at least one improved motion vector as at least one third improved motion vector.

[0274] Clause 21B. A device according to any combination of Clauses 16B to 20B, wherein at least one path of the multipath DMVR comprises a step of applying bidirectional optical flow (BDOF) or a step of applying bilateral matching.

[0275] Clause 22B. The device according to Clause 21B, wherein a first pass applies bilateral matching, a second pass applies bilateral matching, and a third pass applies BDOF.

[0276] Clause 23B. The device according to any combination of Clauses 16B to 22B, wherein at least one second pass subblock has a predetermined maximum width of 16 luma samples and a predetermined maximum height of 16 luma samples.

[0277] Clause 24B. A device according to any combination of Clauses 16B to 23B, wherein at least one third-pass subblock has a predetermined maximum width of 8 luma samples and a predetermined maximum height of 8 luma samples.

[0278] Clause 25B. A device according to any combination of Clauses 16B to 24B, wherein the delta motion range for at least one of the first or second passes is [-8,8] horizontally and [-8,8] vertically, and the delta motion range for the third pass is [-2,2] horizontally and [-2,2] vertically.

[0279] Clause 26B. A device according to any combination of Clauses 16B to 25B, wherein a block of video data is a first block, and one or more processors are configured to apply a shortened multipath DMVR to motion vectors for a second block of video data, and one or more processors decide to skip a given path of the multipath DMVR for the second block in order to apply the shortened multipath DMVR to motion vectors for the second block, and are configured to skip a given path of the multipath DMVR for the second block based on the decision to skip a given path of the multipath DMVR for the second block.

[0280] Clause 27B. The device described in Clause 26B, wherein one or more processors are configured to determine to skip a given path based on the results of a preceding path.

[0281] Clause 28B. A device according to any combination of Clauses 16B to 27B, wherein a block of video data is a first block, and one or more processors are configured to apply a shortened multipath DMVR to motion vectors for a second block of video data, and to determine that, in order to apply the shortened multipath DMVR to motion vectors for the second block, one or more processors are configured to determine that, in order to apply the shortened multipath DMVR to motion vectors for a second block, one or more processors are configured to determine that, in order to apply the shortened multipath DMVR to a particular sub-area of ​​the second block, the particular sub-area comprises one or more sub-blocks of the second block, and to perform the task of performing

[0282] Clause 29B. The device described in Clause 28B, wherein one or more processors are configured to determine to skip a given subblock-based path based on the results of a preceding path.

[0283] Clause 30B. A device according to any combination of Clauses 16B to 25B, wherein Block is a first block of video data, and one or more processors are configured to further decide not to apply DMVR to a second block of video data, to skip all paths of multipath DMVR for the second block based on the decision not to apply DMVR to the second block, and to decode the second block based on the initial motion vector for the second block.

[0284] Clause 31B. A non-temporary computer-readable storage medium storing instructions, wherein, when an instruction is executed, causes one or more processors to apply a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector, and to decode the block based on at least one improved motion vector, wherein the multipath DMVR comprises: a first pass which is block-based and applied to a block of video data; a second pass which is subblock-based and applied to at least one second pass subblock of a block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass which is subblock-based and applied to at least one third pass subblock of a block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0285] Clause 33B. A device for coding video data, comprising means for applying a multipath decoder-side motion vector improvement (DMVR) to a motion vector for a block of video data to determine at least one improved motion vector, and means for decoding a block based on at least one improved motion vector, wherein the multipath DMVR comprises: a first pass which is block-based and applied to a block of video data; a second pass which is subblock-based and applied to at least one second pass subblock of a block of video data, wherein the width of the second pass subblock is less than or equal to the width of the block of video data and the height of the second pass subblock is less than or equal to the height of the block of video data; and a third pass which is subblock-based and applied to at least one third pass subblock of a block of video data, wherein the width of the third pass subblock is less than or equal to the width of the second pass subblock and the height of the third pass subblock is less than or equal to the height of the second pass subblock.

[0286] It should be noted that, depending on the example, some of the actions or events among the techniques described herein may be performed in a different order, and may be added, merged, or excluded entirely (for example, not all of the actions or events described may be necessary for the practice of the technique). Furthermore, in some examples, the actions or events may be performed not sequentially, but in parallel, for example, through multithreading, interrupt handling, or across multiple processors.

[0287] In one or more examples, the described functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on a computer-readable medium, transmitted through a computer-readable medium, or executed by a hardware-based processing unit. The computer-readable medium may include computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of computer programs from one place to another according to a communication protocol, for example. Thus, the computer-readable medium may generally correspond to (1) a non-temporary tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described herein. A computer program product may include a computer-readable medium.

[0288] As an example, and not an limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer. Any connection is also appropriately called computer-readable media. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other temporary media, but instead refer to non-temporary tangible storage media. As used herein, the terms "disk" and "disc" include compact discs (CDs), laser discs, optical discs, digital multipurpose discs (DVDs), floppy disks, and Blu-ray discs, where a disk typically reproduces data magnetically, and a disc reproduces data optically using a laser. Combinations of the above should also be included within the scope of computer-readable media.

[0289] Instructions may be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits. Therefore, the terms “processor” and “processing circuit” as used herein may refer to any of the above-described structures or any other structure suitable for implementing the techniques described herein. In addition, in some embodiments, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a composite codec. Furthermore, the techniques may be fully implemented by one or more circuits or logic elements.

[0290] The techniques of this disclosure may be implemented in a wide variety of devices or apparatus, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units have been described in this disclosure to highlight the functional aspects of devices configured to perform the disclosed techniques, but these do not necessarily require implementation by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit, or provided by a set of interoperable hardware units, including one or more processors as described above, along with appropriate software and / or firmware.

[0291] Various examples were described. These and other examples fall within the scope of the following claims. [Explanation of symbols]

[0292] 102 Source Device 104 Video Sources 106 memory 108 Output Interfaces 110 Computer-readable media 112 Storage Devices 114 File Server 116 Destination Devices 118 Display Devices 120 memory 122 Input Interfaces 130 QTBT structure 132 Coding Tree Units 200 video encoders 202 Mode Selection Unit 204 Residual Generation Unit 206 Conversion Processing Unit 208 Quantization Units 210 Inverse Quantization Unit 212 Inverse Transform Processing Unit 214 Reconstruction Unit 216 Filter Unit 218 Decoded picture buffer 220 Entropy Coding Units 222 Motion Estimation Unit 224 Motion Compensation Unit 226 Intra Prediction Units 230 video data memory 300 video decoders 302 Entropy Decoding Unit 304 Predictive Processing Unit 306 Inverse Quantization Unit 308 Inverse Transform Processing Unit 310 Reconstruction Unit 312 Filter Unit 314 DPB 316 Motion Compensation Unit 317 MPDMVR 318 Intra Prediction Units 320 CPB memory 500 PU0 502 PU0 600 blocks 602 blocks 610 Same position MV 700 templates 702 Current Picture 706 Reference Picture 800 TD0 802 TD1 804 Current Picture 806 Reference Picture 808 Reference Picture 810 TD0 812 TD1 900 Initial MV 902 3x3 search pattern 904 MV 906 The last selected MV 1000 blocks 1002 blocks 1100 CU 1102 Subblock 1200 coding blocks 1202 First Pass 1204A~1204B Subblock 1208A~1208H Subblock

Claims

1. A method for decoding video data, The steps include applying a multipath decoder-side motion vector improvement (DMVR) to the motion vector for a block of video data in order to determine at least one improved motion vector, The step of decoding the block based on the at least one improved motion vector is included. The aforementioned multipath DMVR A first path which is block-based and applied to the blocks of the video data, A second path that is subblock-based and applied to at least one second path subblock of the block of the video data, wherein the width of the second path subblock is less than or equal to the width of the block of the video data, and the height of the second path subblock is less than or equal to the height of the block of the video data, A third path that is subblock-based and applied to at least one third path subblock of the block of the video data, wherein the width of the third path subblock is less than or equal to the width of the second path subblock, the height of the third path subblock is less than or equal to the height of the second path subblock, and the at least one third path subblock of the block of the video data is a subblock with respect to the at least one second path subblock of the block of the video data. A method that includes [a certain feature].

2. The step of applying the first pass derives at least one first improved motion vector for the block of video data, and the at least one first improved motion vector is used in the second pass. The step of applying the second pass derives at least one second improved motion vector for each of at least one second pass subblocks, and the at least one second improved motion vector is used in the third pass. The method according to claim 1, wherein the step of applying the third pass derives at least one third improved motion vector for each of at least one third pass subblocks, and the at least one improved motion vector is determined to be the at least one third improved motion vector.

3. At least one path of the multipath DMVR includes applying bidirectional optical flow (BDOF) or bilateral matching, The method according to claim 1, wherein the first pass comprises applying bilateral matching, the second pass comprises applying bilateral matching, and the third pass comprises applying BDOF.

4. The at least one second pass subblock has a predetermined maximum width of 16 luma samples and a predetermined maximum height of 16 luma samples, and / or The method according to claim 1, wherein the at least one third pass subblock has a predetermined maximum width of 8 luma samples and a predetermined maximum height of 8 luma samples.

5. The method according to claim 1, wherein the delta motion range for at least one of the first or second path is [-8,8] horizontally and [-8,8] vertically, and the delta motion range for the third path is [-2,2] horizontally and [-2,2] vertically.

6. The block of the video data is a first block, and the method further comprises the step of applying a shortened multipath DMVR to a motion vector for a second block of the video data, the applying step is A step of deciding to skip a given path of the multipath DMVR for the second block based on the result of the preceding path, The process includes the step of skipping the given path of the multipath DMVR for the second block based on the decision to skip the given path of the multipath DMVR for the second block, or The block of the video data is a first block, and the method further comprises the step of applying a shortened multipath DMVR to a motion vector for a second block of the video data, the applying step is A step of determining, based on the results of a preceding pass, to fly a given subblock-based path of the multipath DMVR for a specific sub-area of ​​the second block of the video data, wherein the specific sub-area comprises one or more sub-blocks of the second block; The method according to claim 1, further comprising the step of flying the given subblock-based path of the multipath DMVR for the particular sub-area of ​​the second block, based on the decision to fly the given subblock-based path of the multipath DMVR for the particular sub-area of ​​the second block.

7. The block is the first block of the video data, and the method further The steps include deciding not to apply DMVR to the second block of the aforementioned video data, Based on the decision not to apply the DMVR to the second block, the steps include skipping all paths of the multipath DMVR for the second block, The method according to claim 1, further comprising the step of decoding the second block based on an initial motion vector for the second block.

8. A device for decoding video data, A memory configured to store the aforementioned video data, The circuit comprises one or more processors that are implemented in a circuit and are communicably coupled to the memory, and the one or more processors To determine at least one improved motion vector, a multipath decoder-side motion vector improvement (DMVR) is applied to the motion vector for the block of video data, The block is configured to decode based on at least one improved motion vector, The aforementioned multipath DMVR A first path which is block-based and applied to the blocks of the video data, A second path that is subblock-based and applied to at least one second path subblock of the block of the video data, wherein the width of the second path subblock is less than or equal to the width of the block of the video data, and the height of the second path subblock is less than or equal to the height of the block of the video data, A third path that is subblock-based and applied to at least one third path subblock of the block of the video data, wherein the width of the third path subblock is less than or equal to the width of the second path subblock, the height of the third path subblock is less than or equal to the height of the second path subblock, and the at least one third path subblock of the block of the video data is a subblock with respect to the at least one second path subblock of the block of the video data. A device equipped with the following features.

9. The one or more processors are configured to apply the first pass in order to derive at least one first improved motion vector for the block of video data and to use the at least one first improved motion vector in the second pass. The one or more processors are configured to apply the second pass in order to derive at least one second improved motion vector for each of at least one second pass subblocks, and to use the at least one second improved motion vector in the third pass. The device according to claim 8, wherein one or more processors are configured to apply the third pass to derive at least one third improved motion vector for each of at least one third pass subblocks, and to determine the at least one improved motion vector as the at least one third improved motion vector.

10. The device according to claim 8, wherein at least one path of the multipath DMVR includes applying bidirectional optical flow (BDOF), or at least one path of the multipath DMVR includes applying bilateral matching, wherein the first path includes applying bilateral matching, the second path includes applying bilateral matching, and the third path includes applying BDOF.

11. The at least one second pass subblock has a predetermined maximum width of 16 luma samples and a predetermined maximum height of 16 luma samples, and / or The device according to claim 8, wherein the at least one third pass subblock has a predetermined maximum width of 8 luma samples and a predetermined maximum height of 8 luma samples.

12. The device according to claim 8, wherein the delta motion range for at least one of the first or second passes is [-8,8] horizontally and [-8,8] vertically, and the delta motion range for the third pass is [-2,2] horizontally and [-2,2] vertically.

13. The block of the video data is a first block, and one or more processors are configured to apply a shortened multipath DMVR to the motion vector for a second block of the video data, and in order to apply the shortened multipath DMVR to the motion vector for the second block, the one or more processors Based on the results of the preceding pass, it is decided to skip a given path of the multipath DMVR for the second block. Based on the decision to skip the given path of the multipath DMVR for the second block, the multipath DMVR is configured to skip the given path for the second block. or The block of the video data is a first block, and one or more processors are configured to apply a shortened multipath DMVR to the motion vector for a second block of the video data, and in order to apply the shortened multipath DMVR to the motion vector for the second block, the one or more processors Determining, based on the results of a preceding pass, to perform a given subblock-based pass of the multipath DMVR for a specific sub-area of ​​the second block of the video data, wherein the specific sub-area comprises one or more sub-blocks of the second block, The device according to claim 8, configured to perform the given subblock-based path of the multipath DMVR for the particular sub-area of ​​the second block, based on the decision to perform the given subblock-based path of the multipath DMVR for the particular sub-area of ​​the second block.

14. The block is the first block of the video data, and the one or more processors further It was decided not to apply DMVR to the second block of the aforementioned video data. Based on the decision not to apply the DMVR to the second block, all paths of the multipath DMVR for the second block are skipped. The device according to claim 8, configured to decode the second block based on an initial motion vector for the second block.

15. A computer program that, when executed, causes one or more processors to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Decoder Side Motion Vector Refinement in Video Coding

    US20190020895A1

  • Hardware Friendly Constrained Motion Vector Refinement

    US20190238883A1

  • Method and Apparatus of Bilateral Template MV Refinement for Video Coding

    US20200128258A1

  • Decoder-side motion vector refinement

    US20200169748A1

  • Difference calculation based on patial position

    WO2020103852A1