Decoder-side control point motion vector refinement for affine inter prediction in video coding

By using multiple CPMV and DMVR technologies in video decoding, the problem of controlling point motion vector refinement in affine motion prediction is solved, and a more efficient and accurate video decoding process is achieved.

CN119948875APending Publication Date: 2025-05-06QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380068358.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-28
Filing Date
2023-09-29
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When using affine motion prediction, it is difficult to refine the control point motion vector during video decoding, resulting in limited accuracy and efficiency of prediction blocks.

Method used

Video data blocks are predicted by using multiple control point motion vectors (CPMVs), and decoder side motion vector refinement (DMVR) techniques, such as bilateral matching or template matching, refine each CPMV independently to form a more accurate prediction block.

Benefits of technology

This method can improve the accuracy and efficiency of video decoding, reduce the bit rate of the bit stream, and reduce the delay of decoding video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948875A_ABST
    Figure CN119948875A_ABST
Patent Text Reader

Abstract

An example apparatus for decoding video data includes a memory configured to store video data; and a processing system comprising one or more processors implemented in a circuit, the processing system configured to: refine a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; and decoding the current block using the prediction block. In some examples, the CPMVs may each be decoded using a respective merge index and a respective motion vector difference (MVD).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims the benefit of U.S. Application No. 18 / 476,931, filed on September 28, 2023, which claims the benefit of U.S. Provisional Application No. 63 / 487,101, filed on February 27, 2023; U.S. Patent Application No. 18 / 476,931, filed on September 28, 2023, claims the benefit of U.S. Provisional Patent Application No. 63 / 379,045, filed on October 11, 2022; U.S. Patent Application No. 18 / 476,931, filed on September 28, 2023, claims the benefit of U.S. Provisional Patent Application No. 63 / 377,800, filed on September 30, 2022, and the entire contents of each application are incorporated herein by reference. Technical Field

[0003] The present invention relates to video decoding, including video encoding and video decoding. Background Art

[0004] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, digital live broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio telephones, so-called "smart phones", video teleconferencing devices, video streaming devices, and the like. Digital video devices implement video coding techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 / Versatile Video Coding (VVC), and extensions of such standards, as well as proprietary video codecs / formats, such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media. Video devices may more efficiently send, receive, encode, decode, and / or store digital video information by implementing such video coding techniques.

[0005] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction with respect to reference samples in neighboring blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction with respect to reference samples in neighboring blocks in the same picture or temporal prediction with respect to reference samples in other reference pictures. Pictures may be referred to as frames, and reference pictures may be referred to as reference frames. Summary of the invention

[0006] In general, this disclosure describes techniques for refining control point motion vectors during a video decoding (e.g., reproduction) process when affine motion prediction is used to predict a block of video data. Multiple CPMVs may be used to predict a block of video data predicted using affine motion prediction. For example, for a four-parameter affine model, two CPMVs may be used, e.g., corresponding to the top left and top right corners of the block. As another example, for a six-parameter affine model, three CPMVs may be used, e.g., corresponding to the top left, top right, and bottom left corners of the block. In accordance with the techniques of this disclosure, each of the CPMVs may be refined using one or more of a variety of decoder-side motion vector refinement (DMVR) techniques, such as bilateral matching or template matching.

[0007] As an example, each CPMV may be identified using a corresponding merge index. In addition, a motion vector difference (MVD) may be coded for each CPMV. The video decoder may perform a DVMR process on an initial CPMV identified by a merge index to form an intermediate refined CPMV, and then add the MVD to the intermediate refined CPMV to form a refined CPMV. The video decoder may perform a DMVR process independently for each CPMV. In this way, the video decoder may perform the DMVR process in parallel, which may reduce the delay associated with decoding video data.

[0008] In one example, a method for decoding video data includes: refining a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; and decoding the current block using the prediction block. Refining the first CPMV may include determining a first predicted CPMV based on a first merge index value; refining the first predicted CPMV using the first DMVR process to form a first intermediate refined CPMV; decoding first motion vector difference MVD data; and adding the first MVD data to the first intermediate refined CPMV to form the first refined CPMV. Refining the second CPMV may include determining a second predicted CPMV based on a second merge index value; refining the second predicted CPMV using the second DMVR process to form a second intermediate refined CPMV; decoding second motion vector difference MVD data; and adding the second MVD data to the second intermediate refined CPMV to form the second refined CPMV.

[0009] In another example, an apparatus for decoding video data includes: a memory configured to store video data; and a processing system including one or more processors implemented in a circuit, the processing system configured to: refine a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; refine a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; form a prediction block for the current block using the first refined CPMV and the second refined CPMV; and decode the current block using the prediction block.

[0010] In another example, a computer-readable storage medium has instructions stored thereon that, when executed, cause a processor to: refine a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; refine a second CPMV of the current block of video data using a second DMVR process, independent of the first DMVR process, to form a second refined CPMV of the current block; form a prediction block for the current block using the first refined CPMV and the second refined CPMV; and decode the current block using the prediction block.

[0011] In another example, an apparatus for decoding video data includes: a device for refining a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; a device for refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; a device for forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; and a device for decoding the current block using the prediction block.

[0012] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a block diagram illustrating an example video encoding and decoding system that may perform the techniques of this disclosure.

[0014] Figure 2A and 2B A conceptual diagram illustrating an example of control point motion vectors for affine motion vector prediction.

[0015] Figure 3 A conceptual diagram illustrating an example of decoder-side motion vector refinement (DMVR) using bilateral matching.

[0016] Figure 4 is a conceptual diagram illustrating an example of DMVR using template matching.

[0017] Figure 5 A conceptual diagram illustrating an example of individual control point motion vector (CPMV) refinement in accordance with the techniques of this disclosure.

[0018] Figure 6 is a block diagram illustrating an example video encoder that may perform the techniques of this disclosure.

[0019] Figure 7 is a block diagram illustrating an example video decoder that may perform the techniques of this disclosure.

[0020] Figure 8 is a flowchart illustrating an example method for encoding a current block according to the techniques of this disclosure.

[0021] Fig. 9 is a flowchart illustrating an example method for decoding a current block according to the techniques of this disclosure.

[0022] Fig.10 2 is a flow chart illustrating an example method for encoding a current block using refined control point motion vectors (CPMVs) when performing affine motion compensation in accordance with the techniques of this disclosure.

[0023] Fig.11 2 is a flow chart illustrating an example method for decoding a current block using refined CPMV when performing affine motion compensation according to the techniques of this disclosure. DETAILED DESCRIPTION

[0024] In general, this disclosure describes techniques for refining control point motion vectors during a video decoding (e.g., rendering) process when affine motion prediction is used to predict a block of video data. Multiple CPMVs may be used to predict a block of video data predicted using affine motion prediction. For example, for a four-parameter affine model, two CPMVs may be used, e.g., corresponding to the top-left and top-right corners of the block. As another example, for a six-parameter affine model, three CPMVs may be used, e.g., corresponding to the top-left, top-right, and bottom-left corners of the block.

[0025] According to the techniques of this disclosure, each of the CPMVs may be refined using one or more of a variety of decoder-side motion vector refinement (DMVR) techniques, such as bilateral matching (BM) or template matching. In particular, rather than determining a common offset for all CPMVs according to the techniques of this disclosure, a video coder (encoder or decoder) may determine an offset for each CPMV separately. In this way, the refined CPMVs may be used to generate more accurate prediction blocks, which in turn may reduce the bit rate of the bitstream, thereby improving the performance of the video coder.

[0026] Figure 1 1 is a block diagram illustrating an example video encoding and decoding system 100 that may perform techniques of this disclosure. The techniques of this disclosure generally relate to coding (encoding and / or decoding) video data. In general, video data includes any data used to process video. Thus, video data may include original, uncoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling data.

[0027] like Figure 1, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. In particular, source device 102 provides the video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 may include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, mobile devices, tablet computers, set-top boxes, telephone handsets such as smart phones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receiver devices, or the like. In some cases, source device 102 and destination device 116 may be equipped for wireless communication, and thus may be referred to as wireless communication devices.

[0028] exist Figure 1 In the example of , source device 102 includes video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes input interface 122, video decoder 300, memory 120, and display device 118. According to the present invention, video encoder 200 of source device 102 and video decoder 300 of destination device 116 may be configured to apply techniques for refining control point motion vectors. Therefore, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, source devices and destination devices may include other components or arrangements. For example, source device 102 may receive video data from an external video source (e.g., an external camera). Similarly, destination device 116 may be connected to an external display device rather than including an integrated display device.

[0029] like Figure 1 The system 100 shown is only one example. In general, any digital video encoding and / or decoding device may perform techniques for refining control point motion vectors. The source device 102 and the destination device 116 are merely examples of such decoding devices in which the source device 102 generates decoded video data for transmission to the destination device 116. The present invention refers to a "coding" device as a device that performs decoding (encoding and / or decoding) of data. Therefore, the video encoder 200 and the video decoder 300 represent examples of decoding devices, in particular, a video encoder and a video decoder, respectively. In some examples, the source device 102 and the destination device 116 can operate in a substantially symmetrical manner, so that each of the source device 102 and the destination device 116 includes video encoding and decoding components. Therefore, the system 100 can support one-way or two-way video transmission between the source device 102 and the destination device 116, such as for video streaming, video playback, video broadcasting, or video telephony.

[0030] In general, video source 104 represents a source of video data (i.e., raw, undecoded video data) and provides a series of consecutive pictures (also referred to as "frames") of video data to video encoder 200, which encodes the data of the pictures. Video source 104 of source device 102 may include a video capture device, such as a camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As another alternative, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 encodes captured, pre-captured, or computer-generated video data. Video encoder 200 may rearrange the pictures from the order in which they are received (sometimes referred to as "display order") into a decoding order for decoding. Video encoder 200 may generate a bitstream including the encoded video data. Source device 102 may then output the encoded video data onto computer-readable medium 110 via output interface 108 for receipt and / or retrieval by, for example, input interface 122 of destination device 116 .

[0031] The memory 106 of the source device 102 and the memory 120 of the destination device 116 represent general purpose memories. In some examples, the memories 106, 120 can store raw video data, such as raw video from the video source 104 and raw, decoded video data from the video decoder 300. Additionally or alternatively, the memories 106, 120 can store software instructions that can be executed by, for example, the video encoder 200 and the video decoder 300, respectively. Although the memory 106 and the memory 120 are shown separately from the video encoder 200 and the video decoder 300 in this example, it should be understood that the video encoder 200 and the video decoder 300 can also include internal memory for functionally similar or equivalent purposes. In addition, the memories 106, 120 can store encoded video data, such as output from the video encoder 200 and input to the video decoder 300. In some instances, portions of the memories 106, 120 can be allocated as one or more video buffers, such as to store raw, decoded and / or encoded video data.

[0032] The computer-readable medium 110 may represent any type of medium or device capable of transmitting the encoded video data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium that enables the source device 102 to send the encoded video data directly to the destination device 116 in real time, such as via a radio frequency network or a computer-based network. According to a communication standard (e.g., a wireless communication protocol), the output interface 108 can modulate a transmission signal including the encoded video data, and the input interface 122 can demodulate the received transmission signal. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device 102 to the destination device 116.

[0033] In some examples, source device 102 may output the encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0034] In some examples, source device 102 may output the encoded video data to a file server 114 or another intermediate storage device that may store the encoded video data generated by source device 102. Destination device 116 may access the stored video data from file server 114 via streaming or downloading.

[0035] The file server 114 may be any type of server device capable of storing encoded video data and sending the encoded video data to the destination device 116. The file server 114 may represent a web server (e.g., for a website), a server configured to provide file transfer protocol services such as the File Transfer Protocol (FTP) or the File Delivery over Unidirectional Transport (FLUTE) protocol, a content delivery network (CDN) device, a hypertext transfer protocol (HTTP) server, a Multimedia Broadcast Multicast Service (MBMS) or an enhanced MBMS (eMBMS) server, and / or a network attached storage (NAS) device. The file server 114 may additionally or alternatively implement one or more HTTP streaming protocols such as HTTP Dynamic Adaptive Streaming (DASH), HTTP Live Streaming (HLS), Real-Time Streaming Protocol (RTSP), HTTP Dynamic Streaming, or the like.

[0036] Destination device 116 may access the encoded video data from file server 114 through any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on file server 114. Input interface 122 may be configured to operate according to any one or more of the various protocols discussed above for retrieving or receiving media data from file server 114 or other such protocols for retrieving media data.

[0037] Output interface 108 and input interface 122 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components that operate according to any of the various IEEE 802.11 standards, or other physical components. In instances where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transmit data (e.g., encoded video data) according to a cellular communication standard (e.g., 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, or the like). In some instances where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to transmit data (e.g., encoded video data) according to other wireless standards (e.g., IEEE 802.11 specifications, IEEE 802.15 specifications (e.g., ZigBee 5G, 5G ... TM ), Bluetooth TM Standard or the like). In some examples, source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include a SoC device for performing functions attributed to video encoder 200 and / or output interface 108, and destination device 116 may include a SoC device for performing functions attributed to video decoder 300 and / or input interface 122.

[0038] The techniques of this disclosure may be applied to support video decoding for any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions (e.g., Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0039] The input interface 122 of the destination device 116 receives an encoded video bitstream from the computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, or the like). The encoded video bitstream may include signaling information defined by the video encoder 200 and also used by the video decoder 300, such as syntax elements with values ​​describing characteristics and / or processing of video blocks or other coding units (e.g., slices, pictures, groups of pictures, sequences, etc.). The display device 118 displays decoded pictures of the decoded video data to a user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0040] although Figure 1 Not shown, but in some examples, the video encoder 200 and the video decoder 300 can each be integrated with an audio encoder and / or an audio decoder, and can include appropriate MUX-DEMUX units or other hardware and / or software to process a multiplexed stream including both audio and video in a common data stream.

[0041] The video encoder 200 and the video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs) application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the technology of the present invention. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in a corresponding device. A device including the video encoder 200 and / or the video decoder 300 may include an integrated circuit, a microprocessor, and / or a wireless communication device, such as a cellular phone.

[0042] The video encoder 200 and the video decoder 300 may operate according to a video coding standard such as ITU-T H.265, also known as high-efficiency video coding (HEVC) or an extension thereof, such as multi-view and / or scalable video coding extension. Alternatively, the video encoder 200 and the video decoder 300 may operate according to other proprietary or industry standards such as ITU-T H.266, also known as universal video coding (VVC). In other examples, the video encoder 200 and the video decoder 300 may operate according to a proprietary video codec / format such as AOMedia Video 1 (AV1), an extension of AV1, and / or a successor version of AV1 (e.g., AV2). In other examples, the video encoder 200 and the video decoder 300 may operate according to other proprietary formats or industry standards. However, the techniques of the present invention are not limited to any particular coding standard or format. In general, the video encoder 200 and the video decoder 300 may be configured to perform the techniques of the present invention in conjunction with any video coding technique for refining control point motion vectors.

[0043] Typically, the video encoder 200 and the video decoder 300 can perform block-based decoding of a picture. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used in an encoding and / or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Typically, the video encoder 200 and the video decoder 300 can decode video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, the video encoder 200 and the video decoder 300 can decode luminance components and chrominance components instead of decoding red, green, and blue (RGB) data of samples of a picture, where the chrominance components may include both red hue chrominance components and blue hue chrominance components. In some examples, the video encoder 200 converts the received RGB formatted data to a YUV representation before encoding, and the video decoder 300 converts the YUV representation to an RGB format. Alternatively, a pre-processing and post-processing unit (not shown) may perform these conversions.

[0044] This disclosure may generally refer to the coding (e.g., encoding and decoding) of a picture to include the process of encoding or decoding data of a picture. Similarly, this disclosure may refer to the coding of a block of a picture to include the process of encoding or decoding data of the block, such as prediction and / or residual coding. A coded video bitstream typically includes a series of values ​​for syntax elements that represent coding decisions (e.g., coding modes) and partitioning of a picture into blocks. Thus, references to coding a picture or block should generally be understood to mean coding the values ​​of the syntax elements that form the picture or block.

[0045] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder (such as video encoder 200) divides a coding tree unit (CTU) into CUs according to a quadtree structure. That is, the video coder partitions the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. A node without child nodes may be referred to as a "leaf node", and the CU of such a leaf node may include one or more PUs and / or one or more TUs. The video coder may further partition the PU and TU. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of a TU. In HEVC, a PU represents inter-prediction data, and a TU represents residual data. An intra-predicted CU includes intra-prediction information, such as an intra-mode indication.

[0046] As another example, the video encoder 200 and the video decoder 300 may be configured to operate according to VVC. According to VVC, a video decoder (such as the video encoder 200) divides a picture into a plurality of coding tree units (CTUs). The video encoder 200 may partition the CTU according to a tree structure (e.g., a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure removes the concept of multiple partition types, such as the separation between CU, PU, ​​and TU of HEVC. The QTBT structure includes two levels: a first level partitioned according to quadtree partitioning and a second level partitioned according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0047] In the MTT partition structure, blocks may be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) (also referred to as ternary tree (TT)) partitioning. A ternary tree or ternary tree partition is a partition in which a block is divided into three sub-blocks. In some examples, the ternary tree or ternary tree partition divides a block into three sub-blocks without dividing the original block through the center. Partition types in MTT (e.g., QT, BT, and TT) may be symmetric or asymmetric.

[0048] When operating according to the AV1 codec, the video encoder 200 and the video decoder 300 may be configured to decode the video data in blocks. In AV1, the largest decoding block that can be processed is called a super block. In AV1, a super block may be 128×128 luma samples or 64×64 luma samples. However, in subsequent video coding formats (e.g., AV2), super blocks may be defined by different (e.g., larger) luma sample sizes. In some instances, a super block is the top level of a block quadtree. The video encoder 200 may also divide the super block into smaller decoding blocks. The video encoder 200 may divide the super block and other decoding blocks into smaller blocks using square or non-square partitioning. Non-square blocks may include N / 2×N, N×N / 2, N / 4×N, and N×N / 4 blocks. The video encoder 200 and the video decoder 300 may perform separate prediction and transform processing on each decoding block.

[0049] AV1 also defines tiles of video data. A tile is a rectangular array of super blocks that can be decoded independently of other tiles. That is, the video encoder 200 and the video decoder 300 can encode and decode the coded blocks within the tile respectively without using video data from other tiles. However, the video encoder 200 and the video decoder 300 can perform filtering across tile boundaries. The size of the tile can be uniform or uneven. Tile-based decoding can implement parallel processing and / or multithreading for encoder and decoder implementations.

[0050] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma component and the chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the corresponding chroma components).

[0051] The video encoder 200 and the video decoder 300 may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, super block partitioning, or other partitioning structures.

[0052] In some examples, a CTU includes a coding tree block (CTB) of luma samples, two corresponding CTBs of chroma samples for a picture with three sample arrays, or a CTB of samples for a monochrome picture or a picture coded using three separate color planes and syntax structures for coding samples. A CTB may be an N×N block of samples for some value of N, such that the division of a component into a CTB is a partition. A component may be an array or a single sample of one of the three arrays (luma and two chroma) for a picture in 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample of an array for a picture in monochrome format. In some examples, a coding block is an M×N block of samples for some values ​​of M and N, such that the division of a CTB into a coding block is a partition.

[0053] Blocks (e.g., CTUs or CUs) may be grouped in a picture in various ways. As an example, a brick may refer to a rectangular region of a CTU row within a particular tile in a picture. A tile may be a rectangular region of a CTU within a particular tile column and a particular tile row in a picture. A tile column refers to a rectangular region of a CTU having a height equal to the height of the picture and a width specified by a syntax element (e.g., in a picture parameter set). A tile row refers to a rectangular region of a CTU having a height specified by a syntax element (e.g., in a picture parameter set) and a width equal to the width of the picture.

[0054] In some instances, a tile may be partitioned into multiple bricks, each of which may include one or more CTU rows within the tile. Tiles that are not partitioned into multiple bricks may also be referred to as bricks. However, bricks that are true subsets of tiles may not be referred to as tiles. Bricks in a picture may also be arranged in slices. A slice may be an integer number of bricks of a picture, which may be exclusively included in a single network abstraction layer (NAL) unit. In some instances, a slice includes a continuous sequence of several complete tiles or only complete bricks of one tile.

[0055] This disclosure may use "NxN" and "N by N" interchangeably to refer to the sample size of a block (e.g., a CU or other video block) in terms of vertical and horizontal dimensions, e.g., 16x16 samples or 16 by 16 samples. In general, a 16x16 CU will have 16 samples in the vertical direction (y=16) and 16 samples in the horizontal direction (x=16). Likewise, an NxN CU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. The samples in a CU may be arranged in rows and columns. Furthermore, a CU does not necessarily need to have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU may include NxM samples, where M is not necessarily equal to N.

[0056] The video encoder 200 encodes the video data of the CU representing the prediction and / or residual information and other information. The prediction information indicates how to predict the CU in order to form the prediction block of the CU. The residual information generally indicates the sample-by-sample difference between the samples of the CU before encoding and the prediction block.

[0057] In order to predict a CU, the video encoder 200 can generally form a prediction block of the CU by inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting a CU from data of a previously decoded picture, while intra-frame prediction generally refers to predicting a CU from previously decoded data of the same picture. In order to perform inter-frame prediction, the video encoder 200 can use one or more motion vectors to generate a prediction block. The video encoder 200 can generally perform a motion search to identify a reference block that closely matches the CU, for example, in terms of the difference between the CU and the reference block. The video encoder 200 can use the sum of absolute differences (SAD), the sum of square differences (SSD), the mean absolute difference (MAD), the mean square difference (MSD), or other such difference operations to calculate a difference metric to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional prediction or bidirectional prediction to predict the current CU.

[0058] Some examples of VVC also provide an affine motion compensation mode, which can be considered an inter-prediction mode. In the affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion (such as zooming in or out, rotation, perspective motion, or other irregular motion types).

[0059] According to the technology of the present invention, the video encoder 200 may determine that a current block of video data will be predicted using affine motion compensation. The video encoder 200 may determine the number of control point motion vectors (CPMVs) to be used to predict the current block, for example, two, three, or four CPMVs. The video encoder 200 may also encode data representing the CPMVs (e.g., a merge index for each CPMV). In addition, the video encoder 200 may use the merge index to obtain the corresponding intermediate CPMV for each CPMV, and then use a decoder-side motion vector refinement (DMVR) technique to refine the intermediate CPMV, for example, as discussed in more detail below. After refining the intermediate CPMV, the video encoder 200 may calculate a motion vector difference (MVD) value representing the difference between the actual CPMV to be used and the intermediate refined CPMV. The MVD value may include a corresponding motion magnitude and a corresponding direction value. The video encoder 200 may use the actual CPMV to generate a prediction block for the current block. The video encoder 200 may signal the merge index and the MVD value in the bitstream.

[0060] To perform intra prediction, the video encoder 200 can select an intra prediction mode to generate a prediction block. Some instances of VVC provide sixty-seven intra prediction modes, including various directional modes, as well as plane modes and DC modes. Typically, the video encoder 200 selects an intra prediction mode that describes neighboring samples of a current block (e.g., a block of a CU) and predicts samples of the current block from the neighboring samples. Assuming that the video encoder 200 decodes CTUs and CUs in a raster scan order (from left to right, from top to bottom), such samples can typically be above, above and to the left, or to the left of the current block in the same picture as the current block.

[0061] The video encoder 200 encodes data representing the prediction mode of the current block. For example, for the inter-frame prediction mode, the video encoder 200 may encode data indicating which of the various available inter-frame prediction modes is used and motion information of the corresponding mode. For example, for unidirectional or bidirectional inter-frame prediction, the video encoder 200 may encode motion vectors using advanced motion vector prediction (AMVP) or merge mode. The video encoder 200 may use a similar mode to encode motion vectors for affine motion compensation mode.

[0062] AV1 includes two general techniques for encoding and decoding coding blocks of video data. The two general techniques are intra prediction (e.g., intra prediction or spatial prediction) and inter prediction (e.g., inter prediction or temporal prediction). In the context of AV1, when predicting a current block of a video data frame using an intra prediction mode, the video encoder 200 and the video decoder 300 do not use video data from other video data frames. For most intra prediction modes, the video encoder 200 encodes a block of the current frame based on the difference between the sample values ​​in the current block and the prediction values ​​generated from reference samples in the same frame. The video encoder 200 determines the prediction values ​​generated from the reference samples based on the intra prediction mode.

[0063] After prediction (such as intra-frame prediction or inter-frame prediction of a block), the video encoder 200 can calculate residual data for the block. The residual data (e.g., a residual block) represents the sample-by-sample difference between the block and the predicted block of the block formed using the corresponding prediction mode. The video encoder 200 can apply one or more transforms to the residual block to generate transform data in a transform domain rather than a sample domain. For example, the video encoder 200 can apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. In addition, the video encoder 200 can apply a secondary transform after the first transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), etc. The video encoder 200 generates transform coefficients after applying one or more transforms.

[0064] As described above, after any transform to produce transform coefficients, the video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to the process of quantizing the transform coefficients to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. By performing the quantization process, the video encoder 200 may reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 may round down an n-bit value to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 may perform a bitwise right shift of the value to be quantized.

[0065] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place higher energy (and therefore lower frequency) transform coefficients in front of the vector and lower energy (and therefore higher frequency) transform coefficients in the back of the vector. In some examples, the video encoder 200 can scan the quantized transform coefficients using a predefined scan order to generate a serialized vector, and then entropy encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 can perform adaptive scanning. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy encode the one-dimensional vector, for example, according to context adaptive binary arithmetic coding (CABAC). The video encoder 200 can also entropy encode the values ​​of the syntax elements describing the metadata associated with the encoded video data for use by the video decoder 300 when decoding the video data.

[0066] To perform CABAC, video encoder 200 may assign context within a context model to a symbol to be sent. The context may relate to, for example, whether neighboring values ​​of the symbol are zero values. The probability determination may be based on the context assigned to the symbol.

[0067] The video encoder 200 may also generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, to the video decoder 300, for example, in a picture header, a block header, a slice header, or other syntax data, such as a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS). The video decoder 300 may also decode such syntax data to determine how to decode the corresponding video data.

[0068] In this way, the video encoder 200 can generate a bitstream including the encoded video data, for example, a syntax element describing the division of a picture into blocks (e.g., CUs) and prediction and / or residual information of the blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.

[0069] In general, the video decoder 300 performs a process that is inverse to the process performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 may decode the values ​​of the syntax elements of the bitstream using CABAC in a manner substantially similar to, but inverse to, the CABAC encoding process of the video encoder 200. The syntax elements may define partitioning information for partitioning a picture into CTUs and partitioning each CTU according to a corresponding partitioning structure (e.g., a QTBT structure) to define CUs of the CTUs. The syntax elements may further define prediction and residual information for a block (e.g., a CU) of video data.

[0070] The residual information may be represented by, for example, quantized transform coefficients. The video decoder 300 may inverse quantize and inverse transform the quantized transform coefficients of the block to reproduce a residual block of the block. The video decoder 300 forms a prediction block of the block using the signaled prediction mode (intra-frame or inter-frame prediction) and related prediction information (e.g., motion information for inter-frame prediction).

[0071] According to the technology of the present invention, the video decoder 300 may determine that affine motion compensation will be used to predict the current block. The video decoder 300 may receive and decode syntax data indicating the number of control point motion vectors (CMPVs) of the current block. For each of the CPMVs, the video decoder 300 may receive a corresponding merge index and a corresponding motion vector difference (MVD) value. The MVD value may include an amplitude and a direction value. The video decoder 300 may use the merge index to determine the corresponding initial CPMV, and then use the corresponding different DMVR processes to refine the initial CPMV to form a corresponding intermediate refined CPMV. The video decoder 300 may then apply the corresponding MVD value to the corresponding intermediate refined CPMV to form a refined CPMV. The video decoder 300 may then use the refined CPMV to form a prediction block for the current block.

[0072] The video decoder 300 may then combine the predicted block and the residual block (on a sample-by-sample basis) to reproduce the original block.The video decoder 300 may perform additional processing, such as performing a deblocking process to reduce visual artifacts along the boundaries of blocks.

[0073] This disclosure may generally refer to "signaling" certain information, such as syntax elements. The term "signaling" may generally refer to the communication of values ​​of syntax elements and / or other data used to decode encoded video data. That is, video encoder 200 may signal the values ​​of syntax elements in a bitstream. In general, signaling refers to producing the values ​​in the bitstream. As mentioned above, source device 102 may transmit the bitstream to destination device 116 substantially in real time or in non-real time, such as may occur when syntax elements are stored to storage device 112 for later retrieval by destination device 116.

[0074] Figure 2A and 2B is a conceptual diagram illustrating an example of a control point motion vector for affine motion vector prediction. In particular, Figure 2A An example of a current block 130A predicted with two control point motion vectors 132A, 134A is depicted. Figure 2B An example of a current block 130B predicted with three CPMVs 132B, 134B, 136B is depicted.

[0075] The affine motion model can be described as:

[0076]

[0077] In this example, (v x ,v y ) is the motion vector at coordinate (x, y), and a, b, c, d, e, and f are six affine parameters. This disclosure refers to this affine motion model as a 6-parameter affine motion model. In a typical video coder (e.g., video encoder 200 or video decoder 300), a picture is divided into blocks for block-based coding.

[0078] According to the techniques of this disclosure, video encoder 200 and video decoder 300 may independently perform a decoder-side motion vector refinement (DMVR) process on CPMV 132A and 132B. That is, video encoder 200 and video decoder 300 may each perform a first DMVR process on CPMV 132A, and independently perform a different second DMVR process on CPMV 132B. Although DMVR refers to "decoder side," video encoder 200 may perform the same refinement process as video decoder 300, such that CPMV data encoded by video encoder 200 accurately matches CPMV data decoded and reconstructed by video decoder 300.

[0079] According to the techniques of the present disclosure, to encode and refine CPMV 132A, video encoder 200 may determine a first merge candidate that closely matches CPMV 132A. Video encoder 200 may then perform a first DMVR process on the first merge candidate to generate a first intermediate refined CPMV. Video encoder 200 may then calculate a first motion vector difference (MVD) value representing the difference between CPMV 132A and the intermediate refined CPMV. Video encoder 200 may use CPMV 132A to predict current block 130A, and also encode data representing the merge candidate and the first MVD.

[0080] Similarly, according to the techniques of this disclosure, video encoder 200 may determine a second merge candidate that closely matches CPMV 134A. Video encoder 200 may then perform a second DMVR process on the second merge candidate to generate a second intermediately refined CPMV. Video encoder 200 may then calculate a second MVD value that represents the difference between CPMV 134A and the second intermediately refined CPMV. Video encoder 200 may use CPMV 134A to predict current block 130A, and also encode data representing the second merge candidate and the second MVD.

[0081] The video decoder 300 may receive data representing first and second merge candidates and first and second MVDs, and an encoded version of the current block 130. The video decoder 300 may perform a first DMVR process on the first merge to generate a first intermediately refined CPMV from the first merge candidate, then add the first MVD value to the first intermediately refined CPMV to reconstruct the CPMV 132A, which represents the first refined CPMV in this example. The video decoder 300 may also perform a second DMVR process on the second merge candidate to generate a second intermediately refined CPMV from the second merge candidate, then add the second MVD value to the second intermediately refined CPMV to reconstruct the CPMV 134A. The video decoder 300 may then use the CPMVs 132A and 134A to generate a prediction block, and use the prediction block to decode and reconstruct the current block 130A.

[0082] The affine motion model of a block can also be described by three motion vectors (MV) at three different positions that are not in the same row, as well as For example Figure 2B The CPMVs 132B, 134B and 136B are shown in FIG. 3. The 3 positions are often referred to as control points and the 3 motion vectors are referred to as control point motion vectors (CPMVs).

[0083] In the case where the 3 control points are located at the 3 corners of the block, such as Figure 2B As shown, affine motion can be described as:

[0084]

[0085] Where blkW and blkH are the widths of the block ( Figure 2A and 2B w) and height ( Figure 2A and 2B h) in.

[0086] In this example, according to the technology of the present invention, the video encoder 200 may perform three different DVMR processes for each of the CPMVs 132B, 134B, and 136B. For each of these CPMVs, the video encoder 200 may determine a corresponding merge candidate and then perform a corresponding DMVR process on the merge candidate to generate an intermediate refined CPMV. The video encoder 200 may calculate corresponding MVDs representing the differences between the original CPMVs 132B, 134B, and 136B and the corresponding intermediate refined CPMVs.

[0087] Similarly, according to the technology of the present invention, the video decoder 300 may receive data representing merge candidates, MVD values, and an encoded version of the current block 130B (e.g., including quantized transform coefficients). The video decoder 300 may perform a corresponding DMVR process on the corresponding merge candidate to generate an intermediate refined CPMV and then add the corresponding MVD value to the corresponding intermediate refined CPMV to reconstruct the CPMVs 132B, 134B, and 136B. Finally, the video decoder 300 may decode the current block 130B, e.g., generate a prediction block using the CPMVs 132B, 134B, and 136B, decode the quantized transform coefficients to construct a residual block, and then add the residual block to the prediction block on a pixel-by-pixel basis to reconstruct the current block 130B.

[0088] In the affine mode, different motion vectors may be derived for each pixel in a block according to an associated affine motion model. Thus, motion compensation may be performed on a pixel-by-pixel (or sample-by-sample) basis. However, to reduce complexity, sub-block-based motion compensation may be performed, where the block is divided into multiple sub-blocks (which have a smaller block size) and each sub-block is associated with one motion vector for block-based motion compensation. The motion vector for each sub-block may be derived using the representative coordinates of the sub-block. Typically, the center position is used. In one example, the block is divided into non-overlapping sub-blocks. If the block width is blkW, the block height is blkH, the sub-block width is sbW, and the sub-block height is sbH, then there are blkH / sbH rows of sub-blocks and blkW / sbW sub-blocks in each row. For a six-parameter affine motion model, the motion vector of the sub-block (referred to as sub-block MV) at the i-th row (0 <= i < blkW / sbW) and the j-th column (0 <= j < blkH / sbH) may be derived as:

[0089]

[0090] The sub-block MV may be rounded to a predefined precision and stored in a motion buffer for motion compensation and motion vector prediction.

[0091] A simplified four-parameter affine model (for scaling and rotational motion) is described as:

[0092]

[0093] Similarly, the 4-parameter affine model for the block (as in Figure 2A In the example of ) can be represented by two corners of the block (usually the upper left and upper right) as well as Description. The playing field is then described as:

[0094]

[0095] The sub-block MV at the i-th row and j-th column can be derived as:

[0096]

[0097] After performing sub-block based affine motion compensation, the prediction signal can be refined by adding an offset, which can be derived based on the pixel-by-pixel motion and gradient of the prediction signal. The offset at position (m,n) can be calculated as:

[0098] ΔI(m,n)=g x (m,n)*Δv x (m,n)+g y (m,n)*Δv y (m,n)

[0099] In this example, g x (m,n) and g y (m,n) are the horizontal and vertical gradients of the predicted signal respectively. Δv x (m,n) and Δv y (m,n) is the difference in x and y components between the motion vector calculated at collinear pixel position (m,n) and the sub-block MV.

[0100] Let the coordinates of the upper left sample of the sub-block be (0,0) and the center of the sub-block be Given the affine motion parameters a, b, c and d, Δv x (m,n) and Δv y (m,n) can be derived as:

[0101]

[0102] In the control point based affine motion model, the affine motion parameters a, b, c and d can be calculated from the CPMV as:

[0103]

[0104] Figure 3A conceptual diagram illustrating an example of decoder-side motion vector refinement (DMVR) using bilateral matching. In ITU-T H.266 / Versatile Video Coding (VVC), decoder-side motion vector refinement (DMVR) based on bilateral matching is applied to increase the accuracy of the MV of the bi-predictive merge candidate. The BM method calculates the SAD between two candidate blocks in reference picture lists L0 and L1.

[0105] like Figure 3 As illustrated in , the sum of absolute differences (SAD) between reference blocks 142 and 144 based on each MV candidate around the initial MV (MV0, MV1) may be calculated for the current block 140. The MV candidate (MV0', MV1') with the lowest SAD becomes the refined MV and is used to generate a bidirectional prediction signal. The SAD of the initial MV is subtracted by 1 / 4 of the SAD value to serve as a regularization term. The temporal distances from the two reference pictures to the current picture (i.e., the picture order count (POC) difference) may be the same. Therefore, MVD0 may have an opposite sign to MVD1.

[0106] In VVC, the refinement search range is two integer brightness samples from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. A twenty-five (25) point full search can be applied to the integer sample offset search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than a threshold, then the integer sample phase of DMVR terminates. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search phase.

[0107] In H.266 / VVC, integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using parameterized error surface equations instead of an additional search using SAD comparisons. Fractional sample refinement is conditionally called based on the output of the integer sample search stage. When the integer sample search stage terminates with a center with minimum SAD in the first or second iteration search, fractional sample refinement is further applied.

[0108] In the parametric error surface based sub-pixel offset estimation, the center location cost and the costs at four neighboring locations from the center are used to fit a 2-D parabolic error surface equation of the following form:

[0109] E(x,y)=A(x min ) 2 +B(yy min ) 2 +C (1)

[0110] Among them, (x min ,y min) corresponds to the fractional position with the minimum cost, and corresponds to the minimum cost value. By solving the above equation using the cost values ​​of the five search points, (x min ,y min ) can be calculated as:

[0111] x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (1)

[0112] y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2)

[0113] x min and min The value of is automatically constrained between -8 and 8, since all cost values ​​are positive and the minimum is E(0,0). This corresponds to a half-pixel shift with 1 / 16 pixel MV accuracy in VVC. min ,y min ) is added to the integer distance refinement MV to obtain the sub-pixel accurate refinement ΔMV.

[0114] In VVC, the resolution of the MV is 1 / 16 luma sample. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points are around the initial fractional pixel MV with an integer sample offset, so the samples at those fractional positions need to be interpolated for the DMVR search process. In order to reduce the computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect is that by using a bilinear filter, DVMR does not access more reference samples than the normal motion compensation process in the case of a 2-sample search range.

[0115] After obtaining the refined MV through the DMVR search process, a conventional 8-tap interpolation filter is applied to generate the final prediction. In order to avoid accessing more reference samples than the normal motion compensation (MC) process, the samples that are not required by the interpolation process based on the original MV but are required by the interpolation process based on the refined MV are filled from those available samples.

[0116] When the width and / or height of a CU is greater than 16 luma samples, the CU may be further divided into sub-blocks with width and / or height equal to 16 luma samples for the DMVR process.

[0117] In VVC, DMVR can be applied to CUs coded with the following modes and / or features:

[0118] CU-level merge mode with bi-predicted MV

[0119] • Relative to the current picture, one reference picture is in the past and the other reference picture is in the future.

[0120] ● The distances from the two reference pictures to the current picture (i.e., the POC difference) are the same

[0121] ●Both reference images are short-term reference images

[0122] The CU has more than 64 luma samples

[0123] ●CU height and CU width are both greater than or equal to 8 luma samples

[0124] ●BCW weight index indicates equal weights

[0125] ●WP is not enabled for the current block

[0126] CIIP mode is not used for the current block

[0127] In Chen et al., “Non-EE2: DMVR for affine merge coded blocks” (ITU-TSG 16WP3 and ISO / IEC JTC 1 / SC In the document JVET-AA0144-v2 (hereinafter, "JVET-AA0144") of the Joint Video Experts Group (JVET) of 29th session, 27th meeting, July 13-22, decoder-side motion vector refinement (DMVR) is proposed for bidirectional predictive affine merge candidates. If the candidate satisfies the DMVR condition, the translational MV offset is added to all CPMVs of the candidate in the affine merge list. And the MV offset is derived by minimizing the cost of bilateral matching, which is similar to conventional DMVR. In JVET-AA0144, affine motion compensation is performed to generate predictors in both directions. The motion vector offset search process is the same as the first stage of multi-pass DMVR (prediction unit level) in ECM. A 3x3 square search pattern is used to loop through the search range [-8, +8] in the horizontal direction and loop through the search range [-8, +8] in the vertical direction to find the best integer MV offset. A half-pixel search is then performed near the best integer position, and finally an error surface estimation is performed to find an MV offset with 1 / 16 accuracy.

[0128] Video encoders such as video encoder 200 and video decoder 300 may also be configured to perform adaptive DMVR. In general, adaptive DMVR allows different search strategies or methods for bilateral matching to be assigned to different decoded blocks. For example, video encoder 200 may select a search strategy and signal the search strategy in the bitstream using the value of a syntax element encoded in the bitstream. In this way, video decoder 300 may use the value of the syntax element to determine the search strategy. The search strategy may include constraints and / or relationships between MVD0 and MVD1 that may be applied during the bilateral matching search process. Such constraints may include, for example, 1) a mirrored MVD, where MVD0 and MVD1 have the same magnitude but opposite signs, i.e., MVD0=-MVD1 (original DMVR). As another example, 2) MVD0 may be 0 (i.e., both the x and y components of MVD0 are zero), so that MVD0 is fixed when searching around MVD1 to derive the refined MVD1', and MVD0' is set equal to MVD0 (adaptive DMVR). As yet another example, 3) MVD1 may be 0, ie, when searching around MVD0 to derive MVD0', MVD1 is fixed and MVD1' is equal to MVD1 (adaptive DMVR).

[0129] The first syntax element may indicate mode information (e.g., conventional DMVR or adaptive DMVR). The three options described above may be categorized by the first syntax element. Option 1) corresponds to conventional DMVR applied to a coded block when a conventional merge candidate satisfies a DMVR condition, while options 2) and 3) may be applied when a coded block uses a specified new merge mode, where all candidates satisfy the specified DMVR condition. Options 2) and 3) may be distinguished from each other using a mode flag or a merge index.

[0130] Figure 4 is a conceptual diagram illustrating an example of DMVR using template matching. Template matching is another example DMVR process. In template matching, an error measure, such as the sum of absolute differences (SAD), may be minimized for a template region adjacent to a current block relative to a reference region. For example, in Figure 4 , reference templates 154A, 154B having minimum errors with respect to upper template region 152A and left template region 152B of current block 150 may be identified. According to the techniques of the present invention, a similar template matching process may be applied to refine the CPMV of a block predicted using an affine mode.

[0131] Figure 5 A conceptual diagram illustrating an example of individual control point motion vector (CPMV) refinement according to the techniques of the present invention. In particular, Figure 5A current block 330 is depicted having representative blocks 332A, 332B, and 332C at the upper left, upper right, and lower left corners, respectively, of the current block 330. In this example, the current block 330 is predicted using three CPMVs, each corresponding to a respective one of the upper left, upper right, and lower left corners of the current block 330.

[0132] According to the techniques of this invention, each of these CPMVs can use DMVR techniques (such as those described above with respect to Figure 3 and 4 For example, Figure 5 As shown, each of the representative blocks 332A-332C is contained within a corresponding search area 334A-334C. To refine the CPMV associated with, for example, the representative block 332A, the video encoder 200 or the video decoder 300 may perform a bilateral search or a template matching search within the search area 334A of one or more corresponding reference pictures indicated by the original motion information of the CPMV to detect a motion vector that best reduces the cost value according to a cost metric such as the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), etc. The video encoder 200 or the video decoder 300 may refine each CPMV independently in a similar manner.

[0133] Figure 1 The video encoder 200 and the video decoder 300 may be configured to perform DMVR according to the techniques of the present invention to independently refine the CPMV of blocks of video data predicted using affine modes. Although JVET-AA0144 describes refining only the e and f parameters of the affine motion model, i.e., adding the same offset to all CPMVs, the video encoder 200 and the video decoder 300 may perform a decoder-side CPMV refinement method for an affine motion model according to the techniques of the present invention, including refining all 6 parameters (or 4 parameters for a 4-parameter model) using, for example, bilateral matching or template matching. Without losing generosity, in some examples of the techniques of the present invention, a 6-parameter affine motion model represented by 3 CPMVs at the corners of the current block is assumed. The method for a 4-parameter affine motion model can be similarly derived.

[0134] The video encoder 200 and the video decoder 300 may calculate the bilateral matching cost of a given bidirectional affine motion vector as follows. First, the video encoder 200 or the video decoder 300 may add an offset to each CPMV in both directions to update the CPMV. Then, the video encoder 200 or the video decoder 300 may apply affine motion compensation according to the updated CPMV to generate predictors in both directions. Then, the video encoder 200 or the video decoder 300 may use a predefined cost criterion (such as the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean square difference (MSD), etc.) to calculate the distortion between the generated predictors.

[0135] Affine motion compensation can be the same as in VVC, i.e., sub-block based motion compensation followed by prediction refinement (PROF) using optical flow. Affine motion compensation can also be simplified relative to VVC. For example, PROF can be skipped, sub-block sizes can be larger, and / or the interpolation filters can be bilinear interpolation filters instead of the 6-tap or 8-tap interpolation filters of VVC. However, this is not the focus of the present invention, and any type of affine motion compensation can be used to generate the predictor for bilateral matching cost calculation.

[0136] In some examples, the DMVR process may be performed by the video encoder 200 and the video decoder 300 before processing the motion vector difference (MVD) value of the CPMV. That is, the video encoder 200 may determine a merge candidate of a CPMV that closely matches the CPMV, perform the DMVR process on the CPMV to generate an intermediate refined CPMV, and then calculate an MVD value representing the difference between the intermediate refined CPMV and the actual CPMV. Similarly, the video decoder 300 may determine a merge candidate (e.g., using a merge index value received from the video encoder 200), perform the DMVR process on the CPMV merge candidate to generate an intermediate refined CPMV, and then apply the MVD value to the intermediate refined CPMV to reconstruct the actual CPMV.

[0137] The video encoder 200 and the video decoder 300 can be configured to decode data such as a high-level syntax (HLS), a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, or a block header to control whether the described techniques are applied to affine merge candidates.

[0138] In some examples, video encoder 200 and video decoder 300 may refine each CPMV independently using conventional DMVR (eg, two-sided matching) of a representative block. Figure 5In the embodiment, each CPMV may be associated with a corresponding one of the representative blocks 332A-332C. The video encoder 200 or the video decoder 300 may search for each CPMV within the search areas 334A-334C, respectively, based on, for example, bilateral matching or template matching of the representative blocks 332A-332C.

[0139] The representative blocks 332A to 332C may be blocks containing CPMV positions, such as corner samples of the current block 330. For example, the representative block 332A may be a block including the CPMV position as its center, that is, the coordinates of the upper left sample of the block, which may be (xk-halfDx, yk-halfDy), where (xk, yk) are the coordinates of the CPMV, halfDx is half the width of the representative block 332A, and halfDy is half the height of the representative block 332A. Examples of halfDx and halfDy include 2, 4, 8, etc. (e.g., other power-of-two values). The block width and height may also be adaptive depending on the width and height of the current block 330. Various methods of conventional DMVR for representative blocks may be applied, such as various search modes, cost criteria, etc.

[0140] The input and output of the process are the CPMVs of the two prediction directions, denoted as mvAffineInit[i][k] and mvAffineBest[i][k], where i∈{0,1} represents the index of the prediction direction, and k∈{0,1,2} represents the index of the CPMV. For each k in {0,1,2}, the video encoder 200 and the video decoder 300 can use mvAffineInit[i][k] (i=0,1) as input to perform conventional DMVR on the representative block k and output cpmv[i][k]. In this example, cpmv[0][k]=mvAffineInit[0][k]+mvOffset[k], cpmv[1][k]=mvAffineInit[1][k]-mvOffset[k].

[0141] In some examples, mvAffineBest[i][k] is set equal to cpmv[i][k] for all i=0,1 and k=0,1,2.

[0142] In some instances, if the bilateral matching cost of the current block calculated using cpmv[i][k] is less than the cost calculated using mvAffineInit[i][k], then mvAffineBest[i][k] is set equal to cpmv[i][k], otherwise mvAffineBest[i][k] is set equal to mvAffineInit[i][k], for all i=0,1 and k=0,1,2.

[0143] In some examples, both mvAffineInit[i][k] and CPMV[i][k] are used as candidates for CPMV k, and then mvAffineBest[i][k] is set equal to the best combination of those candidates that produces the minimum bilateral matching cost for the current block. The following pseudo code is an example of this process:

[0144] For all i=0,1 and k=0,1,2, set cpmvCand[i][k][0]=mvAffineInit[i][k] and cpmvCand[i][k][1]=cpmv[i][k]:

[0145] Set minCost equal to a predefined maximum value (e.g., the largest 64-bit integer value)

[0146] Loop cpmvIdx0 from 0 to 1

[0147] Loop cpmvIdx1 from 0 to 1

[0148] Loop cpmvIdx2 from 0 to 1

[0149] cpmvTemp[i][0] is set equal to cpmvCand[i][0][cpmvIdx0] for i=0,1

[0150] cpmvTemp[i][1] is set equal to cpmvCand[i][1][cpmvIdx1] for i=0,1

[0151] cpmvTemp[i][2] is set equal to cpmvCand[i][2][cpmvIdx2] for i=0,1

[0152] Use cpmvTemp[i][k] (i=0, 1, k=0, 1, 2) to calculate the bilateral matching cost of the current block, denoted as costTemp.

[0153] If costTemp is less than minCost:

[0154] minCost is set equal to costTemp, and mvAffineBest[i][k]

[0155] Equal to cpmvTemp[i][k], i = 0, 1, k = 0, 1, 2

[0156] The bilateral matching refinement process generally includes 3 stages. The first stage (stage 1) is to perform a bilateral matching search with integer pixel (pel) precision, where the motion vector offset added to the CPMV is a multiple of integer pixels. The second stage (stage 2) is to perform a bilateral matching search with fractional pixel precision, where the motion vector offset added to the CPMV is a multiple of fractional pixels. The last stage (stage 3) uses a parameter error surface model to estimate the optimal motion vector offset based on data from the fractional pixel stage (stage 2). The output CPMV from stage 1 or stage 2 can be used as a replacement for the final output of the bilateral matching search result. This can be applied to all three examples mentioned above, where CPMV[i][k] can be assumed to be the CPMV output from stage 3 and can be replaced with the CPMV output from stage 1 or stage 2.

[0157] In the third example mentioned above, not only the final stage CPMV output CPMV[i][k], but also the CPMV input mvAffineInit[i][k] to the two-sided matching search process can be replaced by the output of stage 1 or stage 2. Either mvAffineInit[i][k] or CPMV[i][k] can be replaced with the CPMV output from stage 1 or stage 2, or both can be replaced simultaneously.

[0158] When performing the bilateral matching search in stage 1 and stage 2, the video encoder 200 or the video decoder 300 may calculate the bilateral matching cost. Based on the bilateral matching cost, the video encoder 200 or the video decoder 300 may sort the corresponding CPMV offsets (or equivalently output CPMVs) and use the best N results in the third example above to search for the best CPMV combination. For each best N results of CPMV, the loop becomes:

[0159] Loop cpmvIdx0 from 0 to N

[0160] Loop cpmvIdx1 from 0 to N

[0161] Loop cpmvIdx2 from 0 to N

[0162] …

[0163] For different CPMV positions (top-left CPMV, top-right CPMV, bottom-left CPMV, etc.), different numbers of candidates may be used, where the loop may be further adapted to:

[0164] Loop cpmvIdx0 from 0 to 1

[0165] Loop cpmvIdx1 from 0 to M

[0166] Loop cpmvIdx2 from 0 to N

[0167] …

[0168] Instead of performing two-sided matching to refine each of the CPMVs for both directions in the above process, the video encoder 200 and the video decoder 300 may perform an adaptive two-sided matching search as an alternative or additional option for refining any or all CPMVs for either direction.

[0169] When used as a replacement, the cpmv[i][k] initially assumed to be a two-sided match search result can be replaced with an adaptive two-sided match search result. Not only can all CPMV two-sided match search results be replaced with the adaptive two-sided match search results, but also a subset of CPMVs can be replaced with the adaptive two-sided match search results. It is also possible to replace one CPMV with the adaptive two-sided match search result of reference list 0, and simultaneously replace another CPMV with the adaptive two-sided match search result of reference list 1.

[0170] Instead of replacement, the adaptive two-sided matching search results can be used as a new variety of the best combination of the search for refined CPMVs. In one example, both the adaptive two-sided matching search results of reference list 0 and reference list 1 can be added to the for loop, as follows:

[0171] For all i=0, 1, k=0, 1, 2 and j=0, 1, 2, set cpmvCand[i][k][0]=mvAffineInit[i][k], cpmvCand[i][k][1]=cpmv[i][k], cpmvCand[i][k][2]=CPMVAdaptL0[i][k] and cpmvCand[i][k][3]=cpmvadaptL1[i][k], where cpmvadaptL0 and cpmvadaptL1 are the adaptive BM search results.

[0172] Set minCost equal to a predefined maximum value (e.g., the largest 64-bit integer value)

[0173] Loop cpmvIdx0 from 0 to 3

[0174] Loop cpmvIdx1 from 0 to 3

[0175] Loop cpmvIdx2 from 0 to 3

[0176] cpmvTemp[i][0] is set equal to cpmvCand[i][0][cpmvIdx0],

[0177] For i=0,1

[0178] cpmvTemp[i][1] is set equal to cpmvCand[i][1][cpmvIdx1],

[0179] For i=0,1

[0180] cpmvTemp[i][2] is set equal to cpmvCand[i][2][cpmvIdx2],

[0181] For i=0,1

[0182] …

[0183] Since the adaptive two-sided matching search requires additional operations and calculations, the overall CPMV search complexity will increase. To avoid the additional two-sided matching search, in one example, the two-sided matching search results can be used to generate pseudo-adaptive search results, for example, for the first CPMV cpmv[1][0], to generate a pseudo-adaptive list 0 result, cpmv[0][0] remains unchanged, and cpmv[1][0] is replaced with the initial input MV mvAffineInit[1][0].

[0184] CPMV refinement may not necessarily be performed on all CPMVs. Early termination can be applied based on previous search result information. The basic concept of affine DMVR is introduced above. When CPMV search is used as a subsequent step after affine DMVR, information from the previous affine DMVR step can be used to determine whether CPMV refinement is required for a certain CPMV. Since the affine mode is a sub-block mode and each of the sub-blocks has a different MV, motion compensation can be performed on each of the sub-blocks, and therefore bilateral matching cost calculation is usually also performed on a sub-block basis. In this case, the bilateral matching cost of each of the sub-blocks will be determined.

[0185] A CPMV search is performed on a square block centered at the CPMV coordinates and having a block size similar to the affine sub-block size. Therefore, the bilateral matching cost of a sub-block close to the CPMV coordinates may be a good candidate for determining early termination. Let the threshold for early termination of the CPMV search be called "Thred". In one instance, the bilateral matching cost of the upper left sub-block is compared with Thred, and if it is less than Thred, the CPMV search for the upper left CPMV can be skipped. The same rules can be applied to the upper right and lower left CPMV. In a second instance, the average bilateral matching cost of several sub-blocks close to the CPMV position is used instead of a single sub-block. For example, the upper left sub-block and the sub-block on the right can be used.

[0186] The CPMV position in the process is not necessarily a corner of the current block 330. In some examples, the CPMV position may be a position inside the current block 330. The output CPMV of the process may be mapped to a desired position (eg, a corner of the current block).

[0187] Some size constraints may be applied to the current block. For example, these techniques may only apply if the area of ​​the current block is greater than a predefined threshold, such as 256 samples. In another example, the techniques may only apply if the width and height of the current block are greater than a predefined threshold, such as 8 or 16 samples. In yet another example, the techniques may only apply if the area of ​​the current block is within a certain range, such as 256 to 4096 samples. In yet another example, the techniques may only apply if the width and height of the current block are within a certain range, such as 16 to 64 samples.

[0188] In some examples, the video encoder 200 and the video decoder 300 may be configured to iteratively refine the CPMV of the current block 330 to minimize the DMVR cost (e.g., bilateral matching cost) of the current block 330. In one iteration, the video encoder 200 and the video decoder 300 may perform a process that loops over all CPMVs and refines the current CPMV while keeping other CPMVs unchanged. The input and output of the process are the CPMVs of the two prediction directions, denoted as mvAffineInit[i][k] and mvAffineBest[i][k], where i=0,1 represents the index of the prediction direction, and k=0,1,2 represents the index of the CPMV.

[0189] The following pseudocode is an example of one iteration of all CPMVs:

[0190] mvOffset[j], j=0, 1, ..., N-1 is an array of possible offsets of CPMV.

[0191] Set minCost equal to a predefined maximum value (e.g., the largest 64-bit integer value)

[0192] For i=0,1,k=0,1,2, mvAffineBest[i][k] is set equal to mvAffineInit[i][k].

[0193] Cyclic CPMV k, k = 0, 1, 2:

[0194] cpmvTemp[i][0] is set equal to mvAffineBest[i][0][cpmvIdx0] for i=0,1

[0195] cpmvTemp[i][1] is set equal to mvAffineBest[i][1][cpmvIdx1] for i=0,1

[0196] cpmvTemp[i][2] is set equal to mvAffineBest[i][2][cpmvIdx1] for i=0,1

[0197] Loop over mvOffset[j], 0, 1, ..., N-1:

[0198] cpmvTemp[0][k] is set equal to cpmvTemp[0][k] + mvOffset[j]

[0199] cpmvTemp[1][k] is set equal to cpmvTemp[1][k] - mvOffset[j]

[0200] Use cpmvTemp[i][k] (i=0, 1, k=0, 1, 2) to calculate the bilateral matching cost of the current block, denoted as costTemp.

[0201] If costTemp is less than minCost:

[0202] minCost is set equal to costTemp, and mvAffineBest[i][k] is equal to cpmvTemp[i][k], i = 0, 1, k = 0, 1, 2

[0203] The video encoder 200 and the video decoder 300 may be configured to apply early termination to the above process. For example, if minCost is less than a threshold, the process may stop. As another example, if all CPMVs do not change during an iteration, the process may stop.

[0204] In some examples, these techniques may be performed only for certain sizes of current block 330 (eg, based on the width, height, and / or area of ​​current block 330 ).

[0205] Adding an offset to one CPMV may change the entire sub-block motion vector field. Subsequently, motion compensation may be performed for each of the sub-blocks to recalculate the bilateral matching cost. However, it may be sufficient to use a sub-sampled version of the sub-block as a representation of the entire CU when calculating the bilateral matching cost of the entire CU, and the sub-sampled bilateral matching cost is used to determine the optimal MV offset for the current CPMV being refined. For example, the method disclosed in U.S. Provisional Application No. 63 / 377,659, entitled “METHODS OF SUBBLOCK SKIPPING FOR AFFINE MOTION SEARCH FOR VIDEO CODING,” filed on September 29, 2022, may be used. Only after the optimal MV offset is determined, the bilateral matching cost for the entire CU is calculated and compared with the previous best bilateral matching cost. In another example, the sub-sampled sub-block bilateral matching cost may be used throughout the process, where the bilateral matching cost of the initial input candidate is also derived from the sub-sampled version of the sub-block.

[0206] The various techniques described above may be performed individually or in combination. In one example, the output of one technique may be used as the input of another technique, and vice versa.

[0207] In some examples, the techniques described above may be further used to perform decoder-side motion vector refinement for affine merging with motion vector difference (MMVD) mode. Affine MMVD mode is similar to MMVD mode because in affine MMVD mode, a merge index is signaled to indicate a basic affine merge candidate, followed by a signaled MVD information. Video encoder 200 and video decoder 300 may add an MVD to each of the control point motion vectors of the basic affine merge candidate to generate a final affine motion vector. The MVD information may include an index to specify a motion magnitude and an index to indicate a motion direction. A distance index may specify motion magnitude information and indicate a predefined offset from a starting point. A direction index may represent the direction of the MVD relative to the starting point.

[0208] After adding MVD to each control point motion vector of the basic affine merge candidate, the control point motion vector refinement method can be applied. The use of MVD in affine MMVD mode can be similar to the application of offsets, for example, the same value can be added to each of the control point motion vectors, and only the e and f parameters of the affine motion model are refined. Therefore, performing the refinement method as proposed in JVET-AA0144 may not be efficient for affine MMVD mode because the offsets are already signaled in the bitstream by the encoder. However, the control point motion vector refinement method can further improve the affine motion by refining other parameters (a, b, c, and d) of the affine model.

[0209] In yet another example, the affine MMVD-based candidates are first refined before the MVD is added to the refined candidates. Various versions of the affine DMVR refinement technique may be applied to the affine MMVD-based candidates individually or in cascade. For example, any of the affine DMVR methods introduced above and any possible combination of these methods may be performed.

[0210] Figure 6 is a block diagram illustrating an example video encoder 200 that may perform the techniques of this disclosure. Figure 6 The above description is provided for the purpose of explanation and should not be construed as limiting the techniques as broadly illustrated and described in this disclosure. For the purpose of explanation, this disclosure describes a video encoder 200 in accordance with techniques of VVC (ITU-T H.266, under development) and HEVC (ITU-T H.265). However, the techniques of this disclosure may be performed by video encoding devices configured for other video coding standards and video coding formats, such as AV1 and successors to the AV1 video coding format.

[0211] exist Figure 6 In the example of , the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any one or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy coding unit 220 can be implemented in one or more processors or in processing circuits. For example, the units of the video encoder 200 can be implemented as one or more circuits or logic elements, as part of a hardware circuit, or as part of a processor, ASIC, or FPGA. In addition, the video encoder 200 may include additional or alternative processors or processing circuits to perform these and other functions.

[0212] Video data memory 230 may store video data to be encoded by components of video encoder 200. Video encoder 200 may receive video data from, for example, video source 104 ( Figure 1) receives video data stored in video data memory 230. DPB 218 can act as a reference picture memory that stores reference video data for predicting subsequent video data by video encoder 200. Video data memory 230 and DPB 218 can be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM) or other types of memory devices. Video data memory 230 and DPB 218 can be provided by the same memory device or a separate memory device. In various examples, video data memory 230 can be on-chip with other components of video encoder 200, as illustrated, or off-chip relative to those components.

[0213] In the present disclosure, references to the video data memory 230 should not be interpreted as limited to memory inside the video encoder 200, unless specifically described as such, or should not be interpreted as limited to memory outside the video encoder 200, unless specifically described as such. Instead, references to the video data memory 230 should be understood as a reference memory that stores video data received by the video encoder 200 for encoding (e.g., video data of a current block to be encoded). Figure 1 The memory 106 may also provide temporary storage of outputs from the various units of the video encoder 200 .

[0214] Figure 6 Various units are shown to help understand the operations performed by the video encoder 200. The units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide specific functionality and are preset in the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions in the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive parameters or output parameters), but the type of operation performed by the fixed-function circuit is generally immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0215] The video encoder 200 may include an arithmetic logic unit (ALU), an elementary function unit (EFU), a digital circuit, an analog circuit, and / or a programmable core formed by a programmable circuit. In an example where the operation of the video encoder 200 is performed using software executed by a programmable circuit, the memory 106 ( Figure 1) may store instructions (eg, object code) for software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.

[0216] The video data memory 230 is configured to store received video data. The video encoder 200 may retrieve pictures of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 may be original video data to be encoded.

[0217] Mode selection unit 202 includes motion estimation unit 222, motion compensation unit 224, and intra prediction unit 226. Mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes. As examples, mode selection unit 202 may include a palette unit, an intra block copy unit (which may be part of motion estimation unit 222 and / or motion compensation unit 224), an affine unit, a linear model (LM) unit, or the like. In some examples, motion compensation unit 224 may be configured to perform affine motion compensation.

[0218] The mode selection unit 202 typically coordinates multiple encoding stages to test combinations of encoding parameters and the resulting rate-distortion values ​​of these combinations. The encoding parameters may include the partitioning of CTUs into CUs, the prediction mode of the CU, the transform type of the residual data of the CU, the quantization parameter of the residual data of the CU, etc. The mode selection unit 202 may ultimately select a combination of encoding parameters that has a better rate-distortion value than other tested combinations.

[0219] The video encoder 200 may partition a picture retrieved from the video data memory 230 into a series of CTUs and pack one or more CTUs into a slice. The mode selection unit 202 may partition the CTUs of the picture according to a tree structure (e.g., the MTT structure, QTBT structure, super block structure, or quadtree structure described above). As described above, the video encoder 200 may form one or more CUs by partitioning the CTUs according to the tree structure. This CU may also be generally referred to as a "video block" or "block".

[0220] In general, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for a current block (e.g., a current CU, or in HEVC, the overlapping portions of a PU and a TU). For inter prediction of the current block, the motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously coded pictures stored in the DPB 218). In particular, the motion estimation unit 222 may calculate values ​​representing how similar a potential reference block is to the current block, such as based on a sum of absolute differences (SAD), a sum of squared differences (SSD), a mean absolute difference (MAD), a mean squared difference (MSD), or the like. The motion estimation unit 222 may generally perform these calculations using the sample-by-sample differences between the current block and the reference block under consideration. The motion estimation unit 222 may identify the reference block having the lowest value resulting from these calculations, indicating the reference block that most closely matches the current block.

[0221] Motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to the position of a current block in a current picture. Motion estimation unit 222 may then provide the motion vectors to motion compensation unit 224. For example, for unidirectional inter prediction, motion estimation unit 222 may provide a single motion vector, while for bidirectional inter prediction, motion estimation unit 222 may provide two motion vectors. Motion compensation unit 224 may then use the motion vectors to generate a prediction block. For example, motion compensation unit 224 may use the motion vectors to retrieve data for the reference block. As another example, if the motion vector has fractional sample precision, motion compensation unit 224 may interpolate values ​​for the prediction block according to one or more interpolation filters. Furthermore, for bidirectional inter prediction, motion compensation unit 224 may retrieve data for two reference blocks identified by respective motion vectors and combine the retrieved data (e.g., by sample-by-sample averaging or weighted averaging).

[0222] When performing affine motion compensation, according to the techniques of this disclosure, motion compensation unit 224 may receive (e.g., from motion estimation unit 222) an actual CPMV to be used to predict the current block. Motion compensation unit 224 may determine a merge candidate for a CPMV that closely matches the actual CPMV. Motion compensation unit 224 may then determine a merge candidate for a CPMV that closely matches the actual CPMV. Figure 5The process described above performs an independent and different DMVR process on each of the CPMV merge candidates to generate a CPMV for the intermediate refinement. The motion compensation unit 224 may then calculate a motion vector difference (MVD) value representing the difference between the actual CPMV and the corresponding intermediate refinement CPMV. The motion compensation unit 224 may use the actual CPMV to generate a prediction block for the current block and provide the merge index and the MVD to the entropy encoding unit 220. The MVD value may represent the length and direction of the offset to be applied to the intermediate refinement CPMV.

[0223] When operating according to the AV1 video coding format, motion estimation unit 222 and motion compensation unit 224 may be configured to encode coding blocks of video data (e.g., both luma and chroma coding blocks) using translational motion compensation, affine motion compensation, overlapped block motion compensation (OBMC), and / or composite inter-intra prediction. When performing affine mode prediction, motion compensation unit 224 may refine each control point motion vector individually according to any of the various techniques of this disclosure, either alone or in any combination.

[0224] As another example, for intra prediction or intra prediction coding, the intra prediction unit 226 may generate a prediction block from samples adjacent to the current block. For example, for directional mode, the intra prediction unit 226 may generally mathematically combine the values ​​of the adjacent samples and pad these calculated values ​​in a defined direction on the current block to generate a prediction block. As another example, for DC mode, the intra prediction unit 226 may calculate an average value of adjacent samples of the current block and generate a prediction block to include this resulting average value for each sample of the prediction block.

[0225] When operating according to the AV1 video coding format, the intra prediction unit 226 may be configured to encode coding blocks of video data (e.g., luma and chroma coding blocks) using directional intra prediction, non-directional intra prediction, recursive filter intra prediction, luma chroma (CFL) prediction, intra block copy (IBC), and / or palette mode. The mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes.

[0226] The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives the original, uncoded version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines a residual block for the current block. In some examples, the residual generation unit 204 may also determine the difference between the sample values ​​in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, the residual generation unit 204 may be formed using one or more subtractor circuits that perform binary subtraction.

[0227] In an example where the mode selection unit 202 partitions a CU into PUs, each PU may be associated with a luma prediction unit and a corresponding chroma prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As indicated above, the size of a CU may refer to the size of the luma coding block of the CU, and the size of a PU may refer to the size of the luma prediction unit of the PU. Assuming that the size of a particular CU is 2N×2N, the video encoder 200 may support PU sizes of 2N×2N or N×N for intra-prediction, and symmetric PU sizes of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitioning of PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-prediction.

[0228] In an example where mode select unit 202 does not further partition a CU into PUs, each CU may be associated with a luma coding block and a corresponding chroma coding block. As above, the size of a CU may refer to the size of the luma coding block of the CU. Video encoder 200 and video decoder 300 may support CU sizes of 2N×2N, 2N×N, or N×2N.

[0229] For other video coding techniques, such as intra block copy mode coding, affine mode coding, and linear model (LM) mode coding, as some examples, mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the coding technique. In some examples, such as palette mode coding, mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how to reconstruct the block based on the selected palette. In such modes, mode selection unit 202 may provide these syntax elements to entropy encoding unit 220 for encoding.

[0230] As described above, the residual generation unit 204 receives video data of a current block and a corresponding prediction block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.

[0231] Transform processing unit 206 applies one or more transforms to the residual block to produce a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transforms to the residual block to form a transform coefficient block. For example, transform processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, transform processing unit 206 may perform multiple transforms on the residual block, such as a primary transform and a secondary transform, such as a rotation transform. In some examples, transform processing unit 206 does not apply a transform to the residual block.

[0232] When operating according to AV1, transform processing unit 206 may apply one or more transforms to the residual block to produce a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transforms to the residual block to form a transform coefficient block. For example, transform processing unit 206 may apply a horizontal / vertical transform combination, which may include a discrete cosine transform (DCT), an asymmetric discrete sine transform (ADST), a flipped ADST (e.g., an ADST in reverse order), and an identity transform (IDTX). When an identity transform is used, a transform is skipped in one of the vertical or horizontal directions. In some examples, transform processing may be skipped.

[0233] Quantization unit 208 may quantize transform coefficients in a transform coefficient block to produce a quantized transform coefficient block. Quantization unit 208 may quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) may adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and therefore, the quantized transform coefficients may have lower precision than the original transform coefficients generated by transform processing unit 206.

[0234] The inverse quantization unit 210 and the inverse transform processing unit 212 may apply inverse quantization and inverse transform to the quantized transform coefficient block, respectively, to reconstruct a residual block from the transform coefficient block. The reconstruction unit 214 may generate a reconstructed block corresponding to the current block (although possibly with a certain degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add samples of the reconstructed residual block to corresponding samples from the prediction block generated by the mode selection unit 202 to generate a reconstructed block.

[0235] Filter unit 216 may perform one or more filter operations on the reconstructed block. For example, filter unit 216 may perform a deblocking operation to reduce blocking artifacts along the edges of the CU. In some examples, the operations of filter unit 216 may be skipped.

[0236] When operating according to AV1, filter unit 216 may perform one or more filter operations on the reconstructed block. For example, filter unit 216 may perform a deblocking operation to reduce blocking artifacts along the edges of the CU. In other examples, filter unit 216 may apply a constrained directional enhancement filter (CDEF), which may be applied after deblocking and may include applying a non-separable, non-linear, low-pass directional filter based on an estimated edge direction. Filter unit 216 may also include a loop recovery filter applied after CDEF, and may include a separable, symmetric, normalized Wiener filter or a dual self-steering filter.

[0237] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in an instance where the operation of the filter unit 216 is not performed, the reconstruction unit 214 may store the reconstructed blocks to the DPB 218. In an instance where the operation of the filter unit 216 is performed, the filter unit 216 may store the filtered reconstructed blocks to the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may retrieve a reference picture formed by the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter-frame prediction on blocks of subsequently encoded pictures. In addition, the intra-frame prediction unit 226 may use the reconstructed blocks in the DPB 218 of the current picture to perform intra-frame prediction on other blocks in the current picture.

[0238] In general, entropy coding unit 220 may entropy encode syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 may entropy encode quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 may entropy encode prediction syntax elements (e.g., motion information for inter-prediction or intra-mode information for intra-prediction) from mode selection unit 202. Entropy coding unit 220 may perform one or more entropy encoding operations on syntax elements (which is another example of video data) to generate entropy-encoded data. For example, entropy coding unit 220 may perform context-adaptive variable length coding (CAVLC) operations, CABAC operations, variable-to-variable (V2V) length coding operations, syntax-based context-adaptive binary arithmetic coding (SBAC) operations, probability interval partitioning entropy (PIPE) coding operations, exponential Golomb coding operations, or another type of entropy coding operations on the data. In some examples, entropy coding unit 220 may operate in a bypass mode, in which syntax elements are not entropy encoded.

[0239] The video encoder 200 may output a bitstream including entropy-coded syntax elements required to reconstruct blocks of a slice or picture. In particular, the entropy coding unit 220 may output a bitstream.

[0240] According to AV1, the entropy coding unit 220 may be configured as a symbol-to-symbol adaptive multi-symbol arithmetic decoder. The syntax elements in AV1 include an alphabet of N elements, and the context (e.g., a probability model) includes a set of N probabilities. The entropy coding unit 220 may store the probabilities as an n-bit (e.g., 15-bit) cumulative distribution function (CDF). The entropy coding unit 220 may perform recursive scaling to update the context with an update factor based on the size of the alphabet.

[0241] The operations described above are described with respect to blocks. Such descriptions should be understood as operations for luma coding blocks and / or chroma coding blocks. As described above, in some examples, the luma coding blocks and chroma coding blocks are the luma components and chroma components of a CU. In some examples, the luma coding blocks and chroma coding blocks are the luma components and chroma components of a PU.

[0242] In some examples, operations performed with respect to luma coding blocks need not be repeated for chroma coding blocks. As one example, operations to identify a motion vector (MV) and reference picture for a luma coding block need not be repeated to identify the MV and reference picture for a chroma block. Instead, the MV for the luma coding block may be scaled to determine the MV for the chroma block, and the reference picture may be the same. As another example, the intra prediction process may be the same for luma coding blocks and chroma coding blocks.

[0243] Figure 7 3 is a block diagram illustrating an example video decoder 300 that may perform the techniques of this disclosure. Figure 7 ,and Figure 7 The techniques as broadly illustrated and described in this disclosure are not limited. For purposes of explanation, this disclosure describes a video decoder 300 in accordance with techniques of VVC (ITU-T H.266, under development) and HEVC (ITU-T H.265). However, the techniques of this disclosure may be performed by video coding devices configured for other video coding standards.

[0244] exist Figure 7In the example of , the video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any or all of the CPB memory 320, the entropy decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 314 can be implemented in one or more processors or in processing circuits. For example, the units of the video decoder 300 can be implemented as one or more circuits or logic elements, as part of a hardware circuit, or as part of a processor, ASIC, or FPGA. In addition, the video decoder 300 may include additional or alternative processors or processing circuits to perform these and other functions.

[0245] Prediction processing unit 304 includes motion compensation unit 316 and intra prediction unit 318. Prediction processing unit 304 may include additional units that perform prediction according to other prediction modes. As an example, prediction processing unit 304 may include a palette unit, an intra block copy unit (which may form part of motion compensation unit 316), an affine unit, a linear model (LM) unit, or the like. In other examples, video decoder 300 may include more, fewer, or different functional components. When performing affine mode prediction, motion compensation unit 316 may refine each control point motion vector individually according to any of the various techniques of this disclosure, either alone or in any combination.

[0246] When performing affine motion compensation, according to the techniques of this disclosure, motion compensation unit 316 may receive merge candidates and MVD values ​​for each of the CPMVs used to form the prediction block. Motion compensation unit 316 may determine the merge candidate for the CPMV according to the merge index received from entropy decoding unit 302. Motion compensation unit 316 may then determine the merge candidate for the CPMV according to, for example, the merge indexes described above with respect to Figure 5 The described process performs an independent and different DMVR process on each of the CPMV merge candidates to generate an intermediate refined CPMV. The motion compensation unit 316 may receive an MVD value representing the difference between the actual CPMVs from the entropy decoding unit 302. The MVD value may represent the length and direction of the offset to be applied to the intermediate refined CPMV. The motion compensation unit 316 may apply the MVD value to the corresponding intermediate refined CPMV to reconstruct the actual CPMV. The motion compensation unit 316 may use the actual CPMV to generate a prediction block for the current block.

[0247] When operating according to AV1, the motion compensation unit 316 may be configured to decode coding blocks of video data (e.g., both luma and chroma coding blocks) using translational motion compensation, affine motion compensation, OBMC, and / or composite inter-intra prediction, as described above. The intra-prediction unit 318 may be configured to decode coding blocks of video data (e.g., both luma and chroma coding blocks) using directional intra prediction, non-directional intra prediction, recursive filter intra prediction, CFL, intra block copy (IBC), and / or color palette mode, as described above.

[0248] CPB memory 320 may store video data, such as an encoded video bitstream, to be decoded by components of video decoder 300. The video bitstream may be, for example, accessed from computer readable medium 110 ( Figure 1 ) obtains video data stored in CPB memory 320. CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from an encoded video bitstream. In addition, CPB memory 320 may store video data other than syntax elements of decoded pictures, such as temporary data representing outputs from various units of video decoder 300. DPB 314 typically stores decoded pictures, and video decoder 300 may output and / or use decoded pictures as reference video data when decoding subsequent data or pictures of the encoded video bitstream. CPB memory 320 and DPB 314 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. CPB memory 320 and DPB 314 may be provided by the same memory device or a separate memory device. In various examples, CPB memory 320 may be on-chip with other components of video decoder 300, or off-chip relative to those components.

[0249] Additionally or alternatively, in some examples, video decoder 300 may retrieve the video from memory 120 ( Figure 1 ) retrieves decoded video data. That is, memory 120 may store data, as discussed above with respect to CPB memory 320. Likewise, memory 120 may store instructions to be executed by video decoder 300 when some or all of the functionality of video decoder 300 is implemented in software to be executed by processing circuitry of video decoder 300.

[0250] Figure 7 The various units shown in FIG. 3 are shown to aid in understanding the operations performed by the video decoder 300. The units may be implemented as fixed function circuits, programmable circuits, or a combination thereof. Figure 6, fixed-function circuits refer to circuits that provide specific functionality and are preset in the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions in the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (for example, to receive parameters or output parameters), but the type of operations performed by the fixed-function circuits is generally immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0251] The video decoder 300 may include an ALU, an EFU, a digital circuit, an analog circuit, and / or a programmable core formed by a programmable circuit. In an example where the operation of the video decoder 300 is performed by software executed on a programmable circuit, an on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.

[0252] Entropy decoding unit 302 may receive the encoded video data from the CPB and entropy decode the video data to reproduce syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 may generate decoded video data based on syntax elements extracted from the bitstream.

[0253] Typically, the video decoder 300 reconstructs a picture on a block-by-block basis. The video decoder 300 may perform a reconstruction operation on each block individually (where a block currently being reconstructed (ie, decoded) may be referred to as a "current block").

[0254] The entropy decoding unit 302 may entropy decode syntax elements defining quantized transform coefficients of the quantized transform coefficient block, as well as transform information, such as a quantization parameter (QP) and / or a transform mode indication. The inverse quantization unit 306 may use the QP associated with the quantized transform coefficient block to determine a degree of quantization, and likewise, determine a degree of inverse quantization to be applied by the inverse quantization unit 306. The inverse quantization unit 306 may, for example, perform a bitwise left shift operation to inverse quantize the quantized transform coefficients. The inverse quantization unit 306 may thereby form a transform coefficient block including the transform coefficients.

[0255] After inverse quantization unit 306 forms the transform coefficient block, inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to produce a residual block associated with the current block. For example, inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotational transform, an inverse directional transform, or another inverse transform to the transform coefficient block.

[0256] In addition, prediction processing unit 304 generates a prediction block based on the prediction information syntax elements entropy decoded by entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is inter-predicted, then motion compensation unit 316 may generate a prediction block. In this case, the prediction information syntax elements may indicate a reference picture in DPB 314 from which to retrieve the reference block, and a motion vector that identifies the position of the reference block in the reference picture relative to the position of the current block in the current picture. Motion compensation unit 316 may generally generate a prediction block in a manner generally similar to that described with respect to motion compensation unit 224 ( Figure 6 )The inter-frame prediction process is performed in a manner similar to that described in FIG.

[0257] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, then the intra-prediction unit 318 may generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, the intra-prediction unit 318 may generally generate a prediction block in a manner substantially similar to that described with respect to the intra-prediction unit 226 ( Figure 6 The intra prediction process is performed in a manner similar to that described in the preceding paragraph. The intra prediction unit 318 may retrieve data of neighboring samples of the current block from the DPB 314.

[0258] The reconstruction unit 310 may reconstruct the current block using the prediction block and the residual block. For example, the reconstruction unit 310 may add samples of the residual block to corresponding samples of the prediction block to reconstruct the current block.

[0259] The filter unit 312 may perform one or more filter operations on the reconstructed block. For example, the filter unit 312 may perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. The operations of the filter unit 312 may not necessarily be performed in all examples.

[0260] The video decoder 300 may store the reconstructed block in the DPB 314. For example, in instances where the operations of the filter unit 312 are not performed, the reconstruction unit 310 may store the reconstructed block to the DPB 314. In instances where the operations of the filter unit 312 are performed, the filter unit 312 may store the filtered reconstructed block to the DPB 314. As discussed above, the DPB 314 may provide reference information (e.g., samples of the current picture for intra-frame prediction and samples of previously decoded pictures for subsequent motion compensation) to the prediction processing unit 304. In addition, the video decoder 300 may output a decoded picture (e.g., a decoded video) from the DPB 314 for subsequent display on a display device (such as a video device). Figure 1 is presented on a display device 118).

[0261] Figure 81 is a flowchart illustrating an example method for encoding a current block according to the techniques of the present invention. The current block may include a current CU. Although relative to the video encoder 200 ( Figure 1 and 6 ), but it should be understood that other devices may be configured to perform the same Figure 8 A similar approach to the one used in this paper.

[0262] In this example, the video encoder 200 initially predicts a current block (350). For example, the video encoder 200 may form a prediction block for the current block using an affine mode. When forming a prediction block using an affine mode, the video encoder 200 may refine each control point motion vector individually or in any combination according to any of the various techniques of the present invention. The video encoder 200 may then calculate a residual block for the current block (352). To calculate the residual block, the video encoder 200 may calculate the difference between the original, uncoded block and the prediction block for the current block. The video encoder 200 may then transform the residual block and quantize the transform coefficients of the residual block (354). Next, the video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or after the scan, the video encoder 200 may entropy encode the transform coefficients (358). For example, the video encoder 200 may encode the transform coefficients using CAVLC or CABAC. The video encoder 200 may then output entropy encoded data for the block (360).

[0263] The video encoder 200 may also decode the current block after encoding the current block to use the decoded version of the current block as reference data for subsequently decoded data (e.g., in an inter-frame or intra-frame prediction mode). Accordingly, the video encoder 200 may inverse quantize and inverse transform the coefficients to reproduce a residual block (362). The video encoder 200 may combine the residual block with the prediction block to form a decoded block (364). The video encoder 200 may then store the decoded block in the DPB 218 (366).

[0264] Fig. 9 1 is a flowchart illustrating an example method for decoding a current block of video data according to the techniques of the present invention. The current block may include a current CU. Although relative to the video decoder 300 ( Figure 1 and 7 ), but it should be understood that other devices may be configured to perform the same Fig. 9 A similar approach to the one used in this paper.

[0265] The video decoder 300 may receive entropy-encoded data of a current block, such as entropy-encoded prediction information and entropy-encoded data of transform coefficients of a residual block corresponding to the current block (370). The video decoder 300 may entropy decode the entropy-encoded data to determine prediction information of the current block and reproduce transform coefficients of the residual block (372). The video decoder 300 may predict the current block (374), for example, using an affine prediction mode as indicated by the prediction information of the current block to calculate a prediction block of the current block. When forming a prediction block using an affine mode, the video decoder 300 may refine each control point motion vector individually or in any combination according to any of the various techniques of the present disclosure. The video decoder 300 may then reverse scan the reproduced transform coefficients (376) to create a quantized transform coefficient block. The video decoder 300 may then inverse quantize the transform coefficients and apply an inverse transform to the transform coefficients to produce a residual block (378). The video decoder 300 may finally decode the current block (380) by combining the prediction block and the residual block.

[0266] Fig.10 1 is a flowchart illustrating an example method for encoding a current block using refined control point motion vectors (CPMVs) when performing affine motion compensation according to the techniques of this disclosure. Figure 1 and 6 The video encoder 200 is described Fig.10 In other examples, other devices may perform this method or similar methods.

[0267] Initially, video encoder 200 determines an actual CPMV to be used to generate a prediction block (400). Video encoder 200 may then determine, for each of the CPMVs, a merge candidate that includes a motion vector that most closely matches the CPMV (402). For example, video encoder 200 may construct a merge candidate list and test each merge candidate in the merge candidate list to determine which merge candidates include a motion vector that is closest to the actual CPMV.

[0268] Then, the video encoder 200 may, for example, perform the above Figure 5 The discussed technique refines the determined merge candidates according to the DMVR (404) to generate an intermediate-refined CPMV. The video encoder 200 may then calculate a motion vector difference (MVD) value representing the difference between the actual CPMV and the intermediate-refined CPMV (406). The MVD value may include a magnitude and a direction value, wherein the magnitude indicates an offset to be applied to the intermediate-refined CPMV and the direction indicates a direction in which the magnitude is applied to reconstruct the actual CPMV.

[0269] The video encoder 200 may then encode the merge candidate (408) and encode the MVD (410). For example, the video encoder 200 may encode a merge index identifying the selected merge candidate in the merge candidate list to encode the merge candidate. To encode the MVD value, the video encoder 200 may encode a distance index representing the magnitude of the MVD value and a direction index representing the direction of the MVD value.

[0270] In addition, video encoder 200 may use the actual CPMV to generate a prediction block (412). For example, video encoder 200 may use the actual CPMV to perform affine motion compensation. Video encoder 200 may also use the prediction block to encode the current block, for example, as described above with respect to Figure 8 Similarly, the video encoder 200 can use the prediction block to reconstruct (ie, decode) the current block.

[0271] In this way, Fig.10 The method represents an example of a method for encoding a block of video data, the method comprising: refining a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; refining a second CPMV of the current block of video data using a second DMVR process, independent of the first DMVR process, to form a second refined CPMV of the current block; forming a prediction block for the current block using the first refined CPMV and the second refined CPMV; and encoding the current block using the prediction block.

[0272] Fig.11 1 is a flowchart illustrating an example method for decoding a current block using a refined CPMV when performing affine motion compensation according to the techniques of this disclosure. Figure 1 and 7 The video decoder 300 is described Fig.11 In other examples, other devices may perform this method or similar methods.

[0273] Initially, the video decoder 300 may decode a merge index (420). The video decoder 300 may also construct a merge candidate list for a corresponding CPMV of a current block of video data. The video decoder 300 may determine a merge candidate in the merge candidate list at a position indicated by the decoded merge index (422). The video decoder 300 may then refine the merge candidate according to the DMVR (424) to construct a refined intermediate CPMV.

[0274] The video decoder 300 may also decode an MVD for the intermediate-refined CPMV (426). The MVD may include an amplitude component represented by an amplitude index and a direction component represented by a direction index. The video decoder 300 may determine the amplitude and direction from the amplitude index and the direction index, respectively. The video decoder 300 may then apply the MVD to the intermediate-refined CPMV (428). For example, the video decoder 300 may add the amplitude in the direction indicated by the direction component to the intermediate-refined CPMV to reconstruct the CPMV. The video decoder 300 may then generate a prediction block using the reconstructed CPMV.

[0275] In this way, Fig.11 The method represents an example of a method for decoding a block of video data, the method comprising: refining a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; forming a prediction block for the current block using the first refined CPMV and the second refined CPMV; and decoding the current block using the prediction block.

[0276] Various examples of the present invention's techniques are summarized in the following clauses:

[0277] Clause 1: A method for decoding video data, the method comprising: refining a first control point motion vector CPMV of a current block of video data using a first decoder-side motion vector refinement DMVR process to form a first refined CPMV of the current block; refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; and decoding the current block using the prediction block.

[0278] Clause 2: The method of clause 1, wherein the first DMVR process comprises a first two-sided matching process, and wherein the second DMVR process comprises a second, different two-sided matching process.

[0279] Clause 3: The method according to any one of clauses 1 and 2 further includes refining a third CPMV of the current block of video data using a third DMVR process independently of the first DMVR process and the second DMVR process to form a third refined CPMV for the current block, wherein forming the prediction block of the current block includes forming the prediction block using the first refined CPMV, the second refined CPMV and the third refined CPMV.

[0280] Clause 4: A method according to any one of clauses 1-3, wherein the first CPMV of the current block corresponds to the upper left corner sample of the current block, and wherein the first DMVR process includes performing a search on a first representative block including the upper left corner sample of the current block, and wherein the second CPMV of the current block corresponds to the upper right corner sample of the current block, and wherein the second DMVR process includes performing a search on a second representative block including the upper right corner sample of the current block.

[0281] Clause 5: A method according to clause 4, wherein the first representative block includes the upper left sample of the current block at the center of the first representative block, and wherein the second representative block includes the upper right sample of the current block at the center of the second representative block.

[0282] Clause 6: A method as recited in any one of clauses 4 and 5, wherein the first representative block is a first 4x4 block of samples, and wherein the second representative block is a second 4x4 block of samples.

[0283] Clause 7: A method according to any of clauses 1-6, wherein a first DMVR process generates a first CPMV offset, and wherein refining the first CPMV includes applying the first CPMV offset to the first CPMV, and wherein a second DMVR process generates a second CPMV offset different from the first CPMV offset, and wherein refining the second CPMV includes applying the second CPMV offset to the second CPMV.

[0284] Clause 8: The method according to any one of clauses 1-7 further includes decoding data indicating that the first CPMV and the second CPMV are to be refined, wherein the data forms part of at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header or a block header.

[0285] Clause 9: A method according to any one of clauses 1-8, wherein the first DMVR process includes: analyzing a first reference block set of a first reference picture in a first reference picture list of a current picture including a current block; analyzing a second reference block set of a second reference picture in a second reference picture list of the current picture; determining whether a first best performance reference block in the first reference block set performs better than a second best performance reference block in the second reference block set; and based on whether the first best performance reference block performs better than the second best performance reference block, using the first best performance reference block or the second best performance reference block to refine the first CPMV, and wherein the second DMVR process includes: analyzing a third reference block set of a third reference picture in the first reference picture list of the current picture; analyzing a fourth reference block set of a fourth reference picture in the second reference picture list of the current picture; determining whether a third best performance reference block in the third reference block set performs better than a fourth best performance reference block in the fourth reference block set; and based on whether the third best performance reference block performs better than the fourth best performance reference block, using the third best performance reference block or the fourth best performance reference block to refine the second CPMV.

[0286] Clause 10: The method of any of clauses 1-9, wherein refining the second CPMV comprises, after refining the first CPMV, refining the second CPMV based on the first refined CPMV.

[0287] Clause 11: The method of any of clauses 1-10 and 30-32, further comprising encoding the current block before decoding the current block.

[0288] Clause 12: An apparatus for decoding video data, the apparatus comprising one or more means for performing the method of any of clauses 1-11 and 30-32.

[0289] Clause 13: The apparatus of clause 12, wherein the one or more means comprise one or more processors implemented in circuitry.

[0290] Clause 14: The apparatus of any of clauses 12 and 13, further comprising a display configured to display the decoded video data.

[0291] Clause 15: The apparatus of any of clauses 12-14, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0292] Clause 16: The apparatus of clauses 12-15, further comprising a memory configured to store the video data.

[0293] Clause 17: A computer-readable storage medium having stored thereon instructions which, when executed, cause a processor of an apparatus for decoding video data to perform the method of any of clauses 1-11 and 30-32.

[0294] Item 18: A device for decoding video data, the device comprising: a device for refining a first control point motion vector CPMV of a current block of video data using a first decoder-side motion vector refinement DMVR process to form a first refined CPMV of the current block; a device for refining a second CPMV of the current block of video data using a second DMVR process independently of the first DMVR process to form a second refined CPMV of the current block; a device for forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; and a device for decoding the current block using the prediction block.

[0295] Clause 19: A method for decoding video data, the method comprising: refining a first control point motion vector CPMV of a current block of video data using a first decoder-side motion vector refinement DMVR process to form a first refined CPMV of the current block; refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; and decoding the current block using the prediction block.

[0296] Clause 20: The method of clause 19, wherein the first DMVR process comprises a first two-sided matching process, and wherein the second DMVR process comprises a second, different two-sided matching process.

[0297] Clause 21: The method according to Clause 19 further includes using a third DMVR process to refine the third CPMV of the current block of video data independently of the first DMVR process and the second DMVR process to form a third refined CPMV of the current block, wherein forming the prediction block of the current block includes forming the prediction block using the first refined CPMV, the second refined CPMV and the third refined CPMV.

[0298] Clause 22: A method according to Clause 19, wherein the first CPMV of the current block corresponds to the upper left corner sample of the current block, and wherein the first DMVR process includes performing a search on a first representative block including the upper left corner sample of the current block, and wherein the second CPMV of the current block corresponds to the upper right corner sample of the current block, and wherein the second DMVR process includes performing a search on a second representative block including the upper right corner sample of the current block.

[0299] Clause 23: A method according to clause 22, wherein the first representative block includes the upper left sample of the current block at the center of the first representative block, and wherein the second representative block includes the upper right sample of the current block at the center of the second representative block.

[0300] Clause 24: The method of clause 22, wherein the first representative block is a first 4x4 block of samples, and wherein the second representative block is a second 4x4 block of samples.

[0301] Clause 25: A method according to clause 19, wherein the first DMVR process generates a first CPMV offset, and wherein refining the first CPMV includes applying the first CPMV offset to the first CPMV, and wherein the second DMVR process generates a second CPMV offset different from the first CPMV offset, and wherein refining the second CPMV includes applying the second CPMV offset to the second CPMV.

[0302] Clause 26: The method according to Clause 19 further includes decoding data indicating the first CPMV and the second CPMV to be refined, wherein the data forms part of at least one of a video parameter set VPS, a sequence parameter set SPS, a picture parameter set PPS, a picture header, a slice header or a block header.

[0303] Clause 27: A method according to clause 19, wherein the first DMVR process includes: analyzing a first reference block set of a first reference picture in a first reference picture list of a current picture including a current block; analyzing a second reference block set of a second reference picture in a second reference picture list of the current picture; determining whether a first best performance reference block in the first reference block set performs better than a second best performance reference block in the second reference block set; and based on whether the first best performance reference block performs better than the second best performance reference block, using the first best performance reference block or the second best performance reference block to refine the first CPMV, and wherein the second DMVR process includes: analyzing a third reference block set of a third reference picture in the first reference picture list of the current picture; analyzing a fourth reference block set of a fourth reference picture in the second reference picture list of the current picture; determining whether a third best performance reference block in the third reference block set performs better than a fourth best performance reference block in the fourth reference block set; and based on whether the performance of the third best performance reference block performs better than the fourth best performance reference block, using the third best performance reference block or the fourth best performance reference block to refine the second CPMV.

[0304] Clause 28: The method of clause 19, wherein refining the second CPMV comprises, after refining the first CPMV, refining the second CPMV based on the first refined CPMV.

[0305] Clause 29: The method of clause 19, further comprising encoding the current block before decoding the current block.

[0306] Clause 30: The method of any of clauses 1-10, further comprising decoding the first CPMV and the second CPMV using affine merging with a motion vector difference (MMVD) mode.

[0307] Item 31. A method according to Item 30, wherein decoding the first CPMV and the second CPMV using the affine MMVD mode includes: decoding a first merge index of the first CPMV; decoding a first motion vector difference MVD value of the first CPMV; decoding a second merge index of the second CPMV; and decoding a second MVD value of the second CPMV.

[0308] Item 32. A method according to Item 31, wherein decoding the first MVD and the second MVD includes: decoding a first distance index representing a first motion amplitude of the first MVD; decoding a first direction index representing a direction of the first MVD; decoding a second distance index representing a second motion amplitude of the second MVD; and decoding a second direction index representing a direction of the second MVD.

[0309] Clause 33: The method of clause 19, further comprising decoding the first CPMV and the second CPMV using an affine merge MMVD mode with motion vector differences.

[0310] Clause 34. A method according to Clause 33, wherein decoding the first CPMV and the second CPMV using the affine MMVD mode includes: decoding the first merge index of the first CPMV; decoding the first motion vector difference MVD value of the first CPMV; decoding the second merge index of the second CPMV; and decoding the second MVD value of the second CPMV.

[0311] Item 35. A method according to Item 34, wherein decoding the first MVD and the second MVD includes: decoding a first distance index representing a first motion amplitude of the first MVD; decoding a first direction index representing a direction of the first MVD; decoding a second distance index representing a second motion amplitude of the second MVD; and decoding a second direction index representing a direction of the second MVD.

[0312] Clause 36: A method for decoding video data, the method comprising: refining a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; and decoding the current block using the prediction block.

[0313] Clause 37: A method according to clause 36, wherein refining the first CPMV includes: determining a first predicted CPMV based on a first merge index value; refining the first predicted CPMV using the first DMVR process to form a first intermediate refined CPMV; decoding first motion vector difference MVD data; and adding the first MVD data to the first intermediate refined CPMV to form the first refined CPMV; and wherein refining the second CPMV includes: determining a second predicted CPMV based on a second merge index value; refining the second predicted CPMV using the second DMVR process to form a second intermediate refined CPMV; decoding second motion vector difference MVD data; and adding the second MVD data to the second intermediate refined CPMV to form the second refined CPMV.

[0314] Clause 38: The method of clause 36, wherein the first DMVR process comprises a first two-sided matching process, and wherein the second DMVR process comprises a second, different two-sided matching process.

[0315] Clause 39: The method according to Clause 36 further includes using a third DMVR process to refine the third CPMV of the current block of video data independently of the first DMVR process and the second DMVR process to form a third refined CPMV of the current block, wherein forming the prediction block of the current block includes forming the prediction block using the first refined CPMV, the second refined CPMV and the third refined CPMV.

[0316] Clause 40: A method according to clause 36, wherein the first CPMV of the current block corresponds to the upper left corner sample of the current block, and wherein the first DMVR process includes performing a search on a first representative block including the upper left corner sample of the current block, and wherein the second CPMV of the current block corresponds to the upper right corner sample of the current block, and wherein the second DMVR process includes performing a search on a second representative block including the upper right corner sample of the current block.

[0317] Clause 41: A method according to clause 40, wherein the first representative block includes the upper left sample of the current block at the center of the first representative block, and wherein the second representative block includes the upper right sample of the current block at the center of the second representative block.

[0318] Clause 42: The method of clause 40, wherein the first representative block is a first 4x4 block of samples, and wherein the second representative block is a second 4x4 block of samples.

[0319] Clause 43: A method according to clause 36, wherein the first DMVR process generates a first CPMV offset, and wherein refining the first CPMV includes applying the first CPMV offset to the first CPMV, and wherein the second DMVR process generates a second CPMV offset different from the first CPMV offset, and wherein refining the second CPMV includes applying the second CPMV offset to the second CPMV.

[0320] Clause 44: The method according to Clause 36 further includes decoding data indicating that the first CPMV and the second CPMV will be refined, and the data forms part of at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header or a block header.

[0321] Clause 45: A method according to clause 36, wherein the first DMVR process includes: analyzing a first reference block set of a first reference picture in a first reference picture list of a current picture including a current block; analyzing a second reference block set of a second reference picture in a second reference picture list of the current picture; determining whether a first best performance reference block in the first reference block set performs better than a second best performance reference block in the second reference block set; and based on whether the first best performance reference block performs better than the second best performance reference block, using the first best performance reference block or the second best performance reference block to refine the first CPMV, and wherein the second DMVR process includes: analyzing a third reference block set of a third reference picture in the first reference picture list of the current picture; analyzing a fourth reference block set of a fourth reference picture in the second reference picture list of the current picture; determining whether a third best performance reference block in the third reference block set performs better than a fourth best performance reference block in the fourth reference block set; and based on whether the performance of the third best performance reference block performs better than the fourth best performance reference block, using the third best performance reference block or the fourth best performance reference block to refine the second CPMV.

[0322] Clause 46: The method of clause 36, wherein refining the second CPMV comprises, after refining the first CPMV, refining the second CPMV based on the first refined CPMV.

[0323] Clause 47: A method according to Clause 46, wherein refining the first CPMV and the second CPMV includes: refining the first CPMV to minimize the first DMVR cost of the current block represented by the intermediate refined first CPMV and the second CPMV until the first refined CPMV produces the minimized first DMVR cost; and refining the second CPMV to minimize the second DMVR cost of the current block represented by the first refined CPMV and the intermediate refined second CPMV until the second refined CPMV produces the minimized second DMVR cost.

[0324] Clause 48: The method of clause 36, further comprising decoding the first CPMV and the second CPMV using an affine merge MMVD mode with motion vector differences.

[0325] Clause 49: A method according to clause 48, wherein decoding the first CPMV and the second CPMV using the affine MMVD mode includes: decoding the first merge index of the first CPMV; decoding the first motion vector difference MVD value of the first CPMV; decoding the second merge index of the second CPMV; and decoding the second MVD value of the second CPMV.

[0326] Item 50: A method according to Item 49, wherein decoding the first MVD and the second MVD includes: decoding a first distance index representing a first motion amplitude of the first MVD; decoding a first direction index representing a direction of the first MVD; decoding a second distance index representing a second motion amplitude of the second MVD; and decoding a second direction index representing a direction of the second MVD.

[0327] Clause 51: The method of clause 36, further comprising encoding the current block prior to decoding the current block.

[0328] Item 52: A device for decoding video data, the device comprising: a memory configured to store video data; and a processing system comprising one or more processors implemented in a circuit, the processing system being configured to: refine a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; refine a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; form a prediction block of the current block using the first refined CPMV and the second refined CPMV; and decode the current block using the prediction block.

[0329] Clause 53: An apparatus according to clause 52, wherein, to refine the first CPMV, the processing system is configured to: determine a first predicted CPMV based on a first merge index value; refine the first predicted CPMV using the first DMVR process to form a first intermediate refined CPMV; decode first motion vector difference MVD data; and add the first MVD data to the first intermediate refined CPMV to form the first refined CPMV; and wherein, to refine the second CPMV, the processing system is configured to: determine a second predicted CPMV based on a second merge index value; refine the second predicted CPMV using the second DMVR process to form a second intermediate refined CPMV; decode second motion vector difference MVD data; and add the second MVD data to the second intermediate refined CPMV to form the second refined CPMV.

[0330] Clause 54: The apparatus of clause 52, wherein the first DMVR process comprises a first two-sided matching process, and wherein the second DMVR process comprises a second, different two-sided matching process.

[0331] Clause 55: An apparatus according to clause 52, wherein the processing system is further configured to refine the third CPMV of the current block of video data using a third DMVR process independently of the first DMVR process and the second DMVR process to form a third refined CPMV of the current block, wherein in order to form the prediction block of the current block, the processing system is configured to form the prediction block using the first refined CPMV, the second refined CPMV and the third refined CPMV.

[0332] Clause 56: The apparatus of clause 55, wherein the first representative block comprises the top left sample of the current block at the center of the first representative block, and wherein the second representative block comprises the top right sample of the current block at the center of the second representative block.

[0333] Clause 57: The apparatus of clause 55, wherein the first representative block is a first 4x4 block of samples, and wherein the second representative block is a second 4x4 block of samples.

[0334] Clause 58: An apparatus according to clause 52, wherein the first DMVR process generates a first CPMV offset, and wherein, to refine the first CPMV, the processing system is configured to apply the first CPMV offset to the first CPMV, and wherein the second DMVR process generates a second CPMV offset different from the first CPMV offset, and wherein, to refine the second CPMV, the processing system is configured to apply the second CPMV offset to the second CPMV.

[0335] Clause 59: An apparatus according to clause 52, wherein the processing system is further configured to decode data indicating that the first CPMV and the second CPMV will be refined, and the data forms part of at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header or a block header.

[0336] Clause 60: An apparatus according to clause 52, wherein, in order to perform the first DMVR process, the processing system is configured to: analyze a first reference block set of a first reference picture in a first reference picture list of a current picture including the current block; analyze a second reference block set of a second reference picture in a second reference picture list of the current picture; determine whether a first best performance reference block in the first reference block set performs better than a second best performance reference block in the second reference block set; and based on whether the first best performance reference block performs better than the second best performance reference block, use the first best performance reference block or the second best performance reference block to refine the A first CPMV, and wherein, in order to perform the second DMVR process, the processing system is configured to: analyze a third reference block set of a third reference picture in the first reference picture list of the current picture; analyze a fourth reference block set of a fourth reference picture in the second reference picture list of the current picture; determine whether a third best performance reference block in the third reference block set performs better than a fourth best performance reference block in the fourth reference block set; and refine the second CPMV using the third best performance reference block or the fourth best performance reference block based on whether the third best performance reference block performs better than the fourth best performance reference block.

[0337] Clause 61: The apparatus of clause 52, wherein to refine the second CPMV, the processing system is configured to refine the second CPMV based on the first refined CPMV after refining the first CPMV.

[0338] Clause 62: An apparatus according to clause 61, wherein, in order to refine the first CPMV and the second CPMV, the processing system is configured to: refine the first CPMV to minimize the first DMVR cost of the current block represented by the intermediate refined first CPMV and the second CPMV until the first refined CPMV produces the minimized first DMVR cost; and refine the second CPMV to minimize the second DMVR cost of the current block represented by the first refined CPMV and the intermediate refined second CPMV until the second refined CPMV produces the minimized second DMVR cost.

[0339] Clause 63: The apparatus of clause 52, wherein the processing system is further configured to decode the first CPMV and the second CPMV using affine merging with a motion vector difference (MMVD) mode.

[0340] Clause 64: An apparatus according to clause 63, wherein, in order to decode the first CPMV, the processing system is configured to: decode a first merge index of the first CPMV; and decode a first motion vector difference (MVD) value of the first CPMV, and wherein, in order to decode the second CPMV, the processing system is configured to: decode a second merge index of the second CPMV; and decode a second MVD value of the second CPMV.

[0341] Clause 65: An apparatus according to clause 64, wherein, in order to decode the first MVD, the processing system is configured to: decode a first distance index representing a first motion amplitude of the first MVD; and decode a first direction index representing a direction of the first MVD, and wherein, in order to decode the second MVD, the processing system is configured to: decode a second distance index representing a second motion amplitude of the second MVD; and decode a second direction index representing the direction of the second MVD.

[0342] Clause 66: The apparatus of clause 52, wherein the processing system is further configured to encode the current block prior to decoding the current block.

[0343] Clause 67: The apparatus of clause 52, further comprising a display configured to display the decoded video data.

[0344] Clause 68: The device of clause 52, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0345] Clause 69: The apparatus of clause 52, further comprising a memory configured to store the video data.

[0346] Item 70: A computer-readable storage medium having instructions stored thereon, which, when executed, cause a processor to: refine a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; refine a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; form a prediction block for the current block using the first refined CPMV and the second refined CPMV; and decode the current block using the prediction block.

[0347] Item 71: An apparatus for decoding video data, the apparatus comprising: an apparatus for refining a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; an apparatus for refining a second CPMV of the current block of video data using a second DMVR process independently of the first DMVR process to form a second refined CPMV of the current block; an apparatus for forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; and an apparatus for decoding the current block using the prediction block.

[0348] Clause 72: A method for decoding video data, the method comprising: refining a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV for the current block; refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; and decoding the current block using the prediction block.

[0349] Clause 73: A method according to clause 72, wherein refining the first CPMV includes: determining a first predicted CPMV based on a first merge index value; refining the first predicted CPMV using the first DMVR process to form a first intermediate refined CPMV; decoding first motion vector difference MVD data; and adding the first MVD data to the first intermediate refined CPMV to form the first refined CPMV; and wherein refining the second CPMV includes: determining a second predicted CPMV based on a second merge index value; refining the second predicted CPMV using the second DMVR process to form a second intermediate refined CPMV; decoding second motion vector difference MVD data; and adding the second MVD data to the second intermediate refined CPMV to form the second refined CPMV.

[0350] Clause 74: The method of any of clauses 72 and 73, wherein the first DMVR process comprises a first two-sided matching process, and wherein the second DMVR process comprises a second, different two-sided matching process.

[0351] Clause 75: The method according to any one of clauses 72-74 further includes refining the third CPMV of the current block of video data using a third DMVR process independently of the first DMVR process and the second DMVR process to form a third refined CPMV of the current block, wherein forming the prediction block of the current block includes forming the prediction block using the first refined CPMV, the second refined CPMV and the third refined CPMV.

[0352] Clause 76: A method according to any of clauses 72-75, wherein the first CPMV of the current block corresponds to the upper left corner sample of the current block, and wherein the first DMVR process includes performing a search on a first representative block including the upper left corner sample of the current block, and wherein the second CPMV of the current block corresponds to the upper right corner sample of the current block, and wherein the second DMVR process includes performing a search on a second representative block including the upper right corner sample of the current block.

[0353] Clause 77: A method according to clause 76, wherein the first representative block includes the upper left sample of the current block at the center of the first representative block, and wherein the second representative block includes the upper right sample of the current block at the center of the second representative block.

[0354] Clause 78: The method of any of clauses 76 and 77, wherein the first representative block is a first 4x4 block of samples, and wherein the second representative block is a second 4x4 block of samples.

[0355] Clause 79: A method according to any of clauses 72-78, wherein the first DMVR process generates a first CPMV offset, and wherein refining the first CPMV includes applying the first CPMV offset to the first CPMV, and wherein the second DMVR process generates a second CPMV offset different from the first CPMV offset, and wherein refining the second CPMV includes applying the second CPMV offset to the second CPMV.

[0356] Clause 80: The method according to any of clauses 72-79 also includes decoding data indicating that the first and second CPMVs are to be refined, the data forming part of at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header or a block header.

[0357] Clause 81: A method according to any one of clauses 72-80, wherein the first DMVR process includes: analyzing a first reference block set of a first reference picture in a first reference picture list of a current picture including a current block; analyzing a second reference block set of a second reference picture in a second reference picture list of the current picture; determining whether a first best performance reference block in the first reference block set performs better than a second best performance reference block in the second reference block set; and based on whether the first best performance reference block performs better than the second best performance reference block, using the first best performance reference block or the second best performance reference block to refine the first CPMV, and wherein the second DMVR process includes: analyzing a third reference block set of a third reference picture in the first reference picture list of the current picture; analyzing a fourth reference block set of a fourth reference picture in the second reference picture list of the current picture; determining whether a third best performance reference block in the third reference block set performs better than a fourth best performance reference block in the fourth reference block set; and based on whether the performance of the third best performance reference block performs better than the fourth best performance reference block, using the third best performance reference block or the fourth best performance reference block to refine the second CPMV.

[0358] Clause 82: The method of any of clauses 72-81, wherein refining the second CPMV comprises, after refining the first CPMV, refining the second CPMV based on the first refined CPMV.

[0359] Clause 83: A method according to clause 82, wherein refining the first CPMV and the second CPMV includes: refining the first CPMV to minimize the first DMVR cost of the current block represented by the intermediate refined first CPMV and the second CPMV until the first refined CPMV produces the minimized first DMVR cost; and refining the second CPMV to minimize the second DMVR cost of the current block represented by the first refined CPMV and the intermediate refined second CPMV until the second refined CPMV produces the minimized second DMVR cost.

[0360] Clause 84: The method of any of clauses 72-83, further comprising decoding the first CPMV and the second CPMV using affine merging with a motion vector difference (MMVD) mode.

[0361] Clause 85: A method according to clause 84, wherein decoding the first CPMV and the second CPMV using the affine MMVD mode includes: decoding the first merge index of the first CPMV; decoding the first motion vector difference MVD value of the first CPMV; decoding the second merge index of the second CPMV; and decoding the second MVD value of the second CPMV.

[0362] Item 86: A method according to Item 85, wherein decoding the first MVD and the second MVD includes: decoding a first distance index representing a first motion amplitude of the first MVD; decoding a first direction index representing a direction of the first MVD; decoding a second distance index representing a second motion amplitude of the second MVD; and decoding a second direction index representing the direction of the second MVD.

[0363] Clause 87: The method of any of clauses 72-86, further comprising encoding the current block before decoding the current block.

[0364] Item 88: A device for decoding video data, the device comprising: a memory configured to store video data; and a processing system comprising one or more processors implemented in a circuit, the processing system configured to: refine a first control point motion vector (CPMV) of a current block of video data using a first decoder-side motion vector refinement (DMVR) process to form a first refined CPMV of the current block; refine a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; form a prediction block of the current block using the first refined CPMV and the second refined CPMV; and decode the current block using the prediction block.

[0365] Clause 89: An apparatus according to clause 88, wherein, to refine the first CPMV, the processing system is configured to: determine a first predicted CPMV based on a first merge index value; refine the first predicted CPMV using the first DMVR process to form a first intermediate refined CPMV; decode first motion vector difference MVD data; and add the first MVD data to the first intermediate refined CPMV to form the first refined CPMV; and wherein, to refine the second CPMV, the processing system is configured to: determine a second predicted CPMV based on a second merge index value; refine the second predicted CPMV using the second DMVR process to form a second intermediate refined CPMV; decode second motion vector difference MVD data; and add the second MVD data to the second intermediate refined CPMV to form the second refined CPMV.

[0366] Clause 90: The apparatus of any of clauses 88 and 89, wherein the first DMVR process comprises a first two-sided matching process, and wherein the second DMVR process comprises a second, different two-sided matching process.

[0367] Clause 91: An apparatus according to any one of clauses 88 to 90, wherein the processing system is further configured to use a third DMVR process independently of the first DMVR process and the second DMVR process to refine the third CPMV of the current block of video data to form a third refined CPMV of the current block, wherein in order to form the prediction block for the current block, the processing system is configured to use the first refined CPMV, the second refined CPMV and the third refined CPMV to form the prediction block.

[0368] Clause 92: The apparatus of clause 91, wherein the first representative block comprises the top left sample of the current block at the center of the first representative block, and wherein the second representative block comprises the top right sample of the current block at the center of the second representative block.

[0369] Clause 93: Apparatus according to any of clauses 91 and 92, wherein the first representative block is a first 4x4 block of samples, and wherein the second representative block is a second 4x4 block of samples.

[0370] Clause 94: An apparatus according to any of clauses 88-93, wherein the first DMVR process generates a first CPMV offset, and wherein, to refine the first CPMV, the processing system is configured to apply the first CPMV offset to the first CPMV, and wherein the second DMVR process generates a second CPMV offset different from the first CPMV offset, and wherein, to refine the second CPMV, the processing system is configured to apply the second CPMV offset to the second CPMV.

[0371] Clause 95: An apparatus according to any one of clauses 88-94, wherein the processing system is also configured to decode data indicating that the first CPMV and the second CPMV are to be refined, and the data forms part of at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header or a block header.

[0372] Clause 96: An apparatus according to any one of clauses 88-95, wherein, in order to perform the first DMVR process, the processing system is configured to: analyze a first reference block set of a first reference picture in a first reference picture list of a current picture including the current block; analyze a second reference block set of a second reference picture in a second reference picture list of the current picture; determine whether a first best performance reference block in the first reference block set performs better than a second best performance reference block in the second reference block set; and use the first best performance reference block or the second best performance reference block based on whether the first best performance reference block performs better than the second best performance reference block. to refine the first CPMV, and wherein, in order to perform the second DMVR process, the processing system is configured to: analyze a third reference block set of a third reference picture in the first reference picture list of the current picture; analyze a fourth reference block set of a fourth reference picture in the second reference picture list of the current picture; determine whether a third best performance reference block in the third reference block set performs better than a fourth best performance reference block in the fourth reference block set; and, based on whether the third best performance reference block performs better than the fourth best performance reference block, use the third best performance reference block or the fourth best performance reference block to refine the second CPMV.

[0373] Clause 97: The apparatus of any of clauses 88-96, wherein to refine the second CPMV, the processing system is configured to refine the second CPMV based on the first refined CPMV after refining the first CPMV.

[0374] Clause 98: An apparatus according to clause 97, wherein, in order to refine the first CPMV and the second CPMV, the processing system is configured to: refine the first CPMV to minimize the first DMVR cost of the current block represented by the intermediate refined first CPMV and the second CPMV until the first refined CPMV produces the minimized first DMVR cost; and refine the second CPMV to minimize the second DMVR cost of the current block represented by the first refined CPMV and the intermediate refined second CPMV until the second refined CPMV produces the minimized second DMVR cost.

[0375] Clause 99: The apparatus of any of clauses 88-98, wherein the processing system is further configured to decode the first CPMV and the second CPMV using affine merging with a motion vector difference (MMVD) mode.

[0376] Clause 100: An apparatus according to clause 99, wherein, in order to decode the first CPMV, the processing system is configured to: decode a first merge index of the first CPMV; and decode a first motion vector difference (MVD) value of the first CPMV, and wherein, in order to decode the second CPMV, the processing system is configured to: decode a second merge index of the second CPMV; and decode a second MVD value of the second CPMV.

[0377] Clause 101: An apparatus according to clause 100, wherein, in order to decode the first MVD, the processing system is configured to: decode a first distance index representing a first motion amplitude of the first MVD; and decode a first direction index representing a direction of the first MVD, and wherein, in order to decode the second MVD, the processing system is configured to: decode a second distance index representing a second motion amplitude of the second MVD; and decode a second direction index representing the direction of the second MVD.

[0378] Clause 102: The apparatus of any of clauses 88-101, wherein the processing system is further configured to encode the current block before decoding the current block.

[0379] Clause 103: The apparatus of any of clauses 88-102, further comprising a display configured to display the decoded video data.

[0380] Clause 104: The device of any of clauses 88-103, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0381] Clause 105: The apparatus of any of clauses 88-104, further comprising a memory configured to store the video data.

[0382] Clause 106: A computer-readable storage medium having stored thereon instructions which, when executed, cause a processor to perform the method of any of clauses 72-87.

[0383] Clause 107: An apparatus for decoding video data, the apparatus comprising means for performing the method of any of clauses 72-87.

[0384] It should be recognized that, depending on the example, certain actions or events of any of the techniques described herein may be performed in a different order, may be added, combined, or omitted entirely (e.g., not all described actions or events are required to practice the techniques). Furthermore, in some examples, actions or events may be performed simultaneously rather than sequentially, such as through multithreading, interrupt processing, or multiple processors.

[0385] In one or more instances, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or sent via a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates, for example, the transfer of a computer program from one place to another according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementation of the techniques described in the present invention. A computer program product may include a computer-readable medium.

[0386] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other temporary media, but are actually directed to non-temporary tangible storage media. As used herein, disks and optical disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, and Blu-ray disks, where disks typically reproduce data magnetically, while optical disks reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0387] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuitry" as used herein may refer to any of the aforementioned structures or any other structure suitable for implementation in the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Furthermore, these techniques may be fully implemented in one or more circuits or logic elements.

[0388] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require implementation by different hardware units. In particular, as described above, the various units may be combined in a codec hardware unit in conjunction with appropriate software and / or firmware, or provided by a collection of interoperating hardware units (including one or more processors as described above).

[0389] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method for decoding video data, the method comprising: refining a first control point motion vector CPMV of a current block of video data using a first decoder-side motion vector refinement DMVR process to form a first refined CPMV of the current block; refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; as well as The current block is decoded using the prediction block.

2. The method according to claim 1, The refinement of the first CPMV includes: Determine a first predicted CPMV according to the first merge index value; refining the first predicted CPMV using the first DMVR process to form a first intermediate refined CPMV; Decoding first motion vector difference MVD data; as well as adding the first MVD data to the first intermediate refined CPMV to form the first refined CPMV; as well as The refinement of the second CPMV includes: Determine a second predicted CPMV according to the second merge index value; refining the second predicted CPMV using the second DMVR process to form a second intermediate refined CPMV; decoding second motion vector difference MVD data; and The second MVD data is added to the second intermediate refined CPMV to form the second refined CPMV.

3. The method of claim 1, wherein the first DMVR process comprises a first double-sided matching process, and wherein the second DMVR process comprises a second, different double-sided matching process.

4. The method of claim 1 , further comprising refining a third CPMV of the current block of video data using a third DMVR process, independent of the first DMVR process and the second DMVR process, to form a third refined CPMV of the current block, wherein forming the prediction block of the current block comprises forming the prediction block using the first refined CPMV, the second refined CPMV, and the third refined CPMV.

5. The method according to claim 1, wherein the first CPMV of the current block corresponds to an upper left corner sample of the current block, and wherein the first DMVR process includes performing a search for a first representative block including the upper left corner sample of the current block, and The second CPMV of the current block corresponds to an upper right corner sample of the current block, and the second DMVR process includes performing a search for a second representative block including the upper right corner sample of the current block.

6. The method according to claim 5, wherein: The first representative block includes an upper left sample of the current block located at a center of the first representative block, and wherein the second representative block includes an upper right sample of the current block located at a center of the second representative block.

7. The method according to claim 5, wherein: The first representative block is a first 4x4 block of samples, and wherein the second representative block is a second 4x4 block of samples.

8. The method according to claim 1, in, The first DMVR process generates a first CPMV offset, and wherein refining the first CPMV comprises applying the first CPMV offset to the first CPMV, and Wherein the second DMVR process generates a second CPMV offset different from the first CPMV offset, and wherein refining the second CPMV includes applying the second CPMV offset to the second CPMV.

9. The method according to claim 1 further includes decoding data indicating that the first CPMV and the second CPMV are to be refined, and the data forms part of at least one of a video parameter set VPS, a sequence parameter set SPS, a picture parameter set PPS, a picture header, a slice header or a block header.

10. The method according to claim 1, in, The first DMVR process includes: Analyzing a first reference block set of a first reference picture in a first reference picture list of a current picture including the current block; Analyze a second reference block set of a second reference picture in a second reference picture list of the current picture; determining whether a first best performing reference block in the first set of reference blocks performs better than a second best performing reference block in the second set of reference blocks; and refining the first CPMV using the first best performance reference block or the second best performance reference block according to whether the first best performance reference block performs better than the second best performance reference block, and The second DMVR process includes: analyzing a third reference block set of a third reference picture in the first reference picture list of the current picture; analyzing a fourth reference block set of a fourth reference picture in the second reference picture list of the current picture; determining whether a third best performing reference block in the third set of reference blocks performs better than a fourth best performing reference block in the fourth set of reference blocks; and Depending on whether the third best performing reference block performs better than the fourth best performing reference block, the second CPMV is refined using the third best performing reference block or the fourth best performing reference block.

11. The method according to claim 1, wherein: Refining the second CPMV includes: after refining the first CPMV, refining the second CPMV based on the first refined CPMV.

12. The method according to claim 11, wherein: Refining the first CPMV and the second CPMV includes: Refining the first CPMV to minimize a first DMVR cost of the current block represented by the intermediate refined first CPMV and the second CPMV until a minimized first DMVR cost is produced by the first refined CPMV; and The second CPMV is refined to minimize a second DMVR cost of the current block represented by the first refined CPMV and the intermediate refined second CPMV until a minimized second DMVR cost is produced by the second refined CPMV.

13. The method of claim 1, further comprising decoding the first CPMV and the second CPMV using an affine merge MMVD mode with motion vector differences.

14. The method of claim 13, wherein coding the first CPMV and the second CPMV using an affine MMVD mode comprises: Decoding a first merge index of the first CPMV; Decoding a first motion vector difference MVD value of the first CPMV; Decoding a second merge index of the second CPMV; as well as A second MVD value of the second CPMV is decoded.

15. The method of claim 14, wherein decoding the first MVD and the second MVD comprises: decoding a first distance index representing a first motion magnitude of the first MVD; decoding a first direction index representing a direction of the first MVD; decoding a second distance index representing a second motion magnitude of the second MVD; and A second direction index representing a direction of the second MVD is decoded.

16. The method of claim 1, further comprising encoding the current block before decoding the current block.

17. An apparatus for decoding video data, the apparatus comprising: a memory configured to store video data; as well as A processing system, comprising one or more processors implemented in circuitry, the processing system being configured to: refining a first control point motion vector CPMV of a current block of the video data using a first decoder-side motion vector refinement DMVR process to form a first refined CPMV of the current block; refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV for the current block; forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; as well as The current block is decoded using the prediction block.

18. The device according to claim 17, in, To refine the first CPMV, the processing system is configured to: Determine a first predicted CPMV according to the first merge index value; refining the first predicted CPMV using the first DMVR process to form a first intermediate refined CPMV; Decoding first motion vector difference MVD data; as well as adding the first MVD data to the first intermediate refined CPMV to form the first refined CPMV; and In order to refine the second CPMV, the processing system is configured as follows: Determine a second predicted CPMV according to the second merge index value; refining the second predicted CPMV using the second DMVR process to form a second intermediate refined CPMV; decoding second motion vector difference MVD data; and The second MVD data is added to the second intermediate refined CPMV to form the second refined CPMV.

19. The apparatus of claim 17, wherein the first DMVR process comprises a first two-sided matching process, and wherein the second DMVR process comprises a second, different two-sided matching process.

20. The apparatus of claim 17, wherein the processing system is further configured to refine a third CPMV of the current block of video data using a third DMVR process, independent of the first DMVR process and the second DMVR process, to form a third refined CPMV of the current block, wherein to form the prediction block of the current block, the processing system is configured to form the prediction block using the first refined CPMV, the second refined CPMV, and the third refined CPMV.

21. The device according to claim 17, in, The first DMVR process generates a first CPMV offset, and wherein, to refine the first CPMV, the processing system is configured to apply the first CPMV offset to the first CPMV, and Wherein the second DMVR process generates a second CPMV offset different from the first CPMV offset, and wherein, to refine the second CPMV, the processing system is configured to apply the second CPMV offset to the second CPMV.

22. An apparatus according to claim 17, wherein the processing system is further configured to decode data indicating that the first CPMV and the second CPMV are to be refined, the data forming part of at least one of a video parameter set VPS, a sequence parameter set SPS, a picture parameter set PPS, a picture header, a slice header or a block header.

23. The device according to claim 17, in, To perform the first DMVR process, the processing system is configured to: Analyzing a first reference block set of a first reference picture in a first reference picture list of a current picture including the current block; Analyze a second reference block set of a second reference picture in a second reference picture list of the current picture; determining whether a first best performing reference block in the first set of reference blocks performs better than a second best performing reference block in the second set of reference blocks; as well as refining the first CPMV using the first best performance reference block or the second best performance reference block according to whether the first best performance reference block performs better than the second best performance reference block, and Wherein, in order to perform the second DMVR process, the processing system is configured as follows: analyzing a third reference block set of a third reference picture in the first reference picture list of the current picture; analyzing a fourth reference block set of a fourth reference picture in the second reference picture list of the current picture; determining whether a third best performing reference block in the third set of reference blocks performs better than a fourth best performing reference block in the fourth set of reference blocks; and Depending on whether the third best performing reference block performs better than the fourth best performing reference block, the second CPMV is refined using the third best performing reference block or the fourth best performing reference block.

24. The apparatus of claim 17, wherein to refine the second CPMV, the processing system is configured to refine the second CPMV based on the first refined CPMV after refining the first CPMV.

25. The device according to claim 24, wherein: In order to refine the first CPMV and the second CPMV, the processing system is configured to: Refining the first CPMV to minimize a first DMVR cost of the current block represented by the intermediate refined first CPMV and the second CPMV until a minimized first DMVR cost is produced by the first refined CPMV; and The second CPMV is refined to minimize a second DMVR cost of the current block represented by the first refined CPMV and the intermediate refined second CPMV until a minimized second DMVR cost is produced by the second refined CPMV.

26. The apparatus of claim 17, wherein the processing system is further configured to decode the first CPMV and the second CPMV using an affine merge MMVD mode with motion vector differences.

27. The device according to claim 26, in, To encode the first CPMV, the processing system is configured to: Decoding a first merge index of the first CPMV; and decoding a first motion vector difference MVD value of the first CPMV, and wherein to decode the second CPMV, the processing system is configured to: Decoding a second merge index of the second CPMV; and The second MVD value of the second CPMV is decoded.

28. The device according to claim 27, In order to decode the first MVD, the processing system is configured to: Decoding a first distance index representing a first motion magnitude of the first MVD; and decoding a first direction index indicating the direction of the first MVD, and In order to decode the second MVD, the processing system is configured to: decoding a second distance index representing a second motion magnitude of the second MVD; and A second direction index indicating the direction of the second MVD is decoded.

29. The apparatus of claim 17, wherein the processing system is further configured to encode the current block before decoding the current block.

30. The apparatus of claim 17, further comprising a display configured to display the decoded video data.

31. The device of claim 17, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

32. The apparatus of claim 17, further comprising a memory configured to store the video data.

33. A computer readable storage medium having stored thereon instructions which, when executed, cause a processor to: refining a first control point motion vector CPMV of a current block of video data using a first decoder-side motion vector refinement DMVR process to form a first refined CPMV of the current block; refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; forming a prediction block of the current block using the first refined CPMV and the second refined CPMV; as well as The current block is decoded using the prediction block.

34. An apparatus for decoding video data, the apparatus comprising: means for refining a first control point motion vector CPMV of a current block of video data using a first decoder-side motion vector refinement DMVR process to form a first refined CPMV of the current block; means for refining a second CPMV of the current block of video data using a second DMVR process independent of the first DMVR process to form a second refined CPMV of the current block; means for forming a prediction block for the current block using the first refined CPMV and the second refined CPMV; as well as Means for decoding the current block using the prediction block.