Adaptive bilateral matching for decoder-side affine motion vector refinement
By adopting adaptive affine DMVR technology in video encoding and decoding technology, setting a certain motion vector difference in the reference picture list is zero, and only refining the predictor of another reference list is solved, which solves the problem of inefficient inter prediction decoding in the prior art, and achieving more efficient video data decoding.
Patent Information
- Application Number
- CN202380079280.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2023-11-15
- Publication Date
- 2025-06-24
AI Technical Summary
The existing video encoding and decoding techniques have problems with inefficient decoding in inter-frame prediction, especially when one of the predictors already has high accuracy.
Adaptive affine bilateral matching side affine motion vector refinement (DMVR) technology is used to set the difference between one motion vector in the reference picture list to zero, and only the predictor of the other reference list is refined to improve the decoding efficiency.
The decoding efficiency of video data is improved, and the motion vectors are further refined through flexible signaling, which enhances the accuracy of inter prediction.
Smart Images

Figure CN120202659A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority to U.S. Patent Application No. 18 / 507,544, filed on November 13, 2023, and U.S. Provisional Patent Application No. 63 / 384,979, filed on November 25, 2022, the entire contents of which are incorporated herein by reference. U.S. Patent Application No. 18 / 507,544, filed on November 13, 2023, claims the benefit of U.S. Provisional Patent Application No. 63 / 384,979, filed on November 25, 2022. Technical Field
[0003] This disclosure relates to video encoding and video decoding. Background Art
[0004] Digital video capabilities can be incorporated into a variety of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e - book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, so - called "smartphones", video teleconferencing devices, video streaming devices, and the like. Digital video devices implement video decoding technologies, such as those described in the standards defined by MPEG - 2, MPEG - 4, ITU - T H.263, ITU - T H.264 / MPEG - 4, Part 10, Advanced Video Coding (AVC), ITU - T H.265 / High Efficiency Video Coding (HEVC), ITU - T H.266 / Versatile Video Coding (VVC) and extensions of such standards, as well as proprietary video codecs / formats such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media. By implementing such video decoding technologies, video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information.
[0005] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or eliminate redundancy inherent in a video sequence. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction relative to reference samples in neighboring blocks within the same picture. Video blocks in an inter-coded (P or B) slice of a picture may be encoded using spatial prediction relative to reference samples in neighboring blocks within the same picture or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. SUMMARY OF THE DISCLOSURE
[0006] Generally, the present disclosure describes techniques for encoding and decoding video data, including techniques for inter-prediction. More specifically, the present disclosure describes techniques for adaptive affine decoder-side motion vector refinement (DMVR) using bilateral matching. In some examples of affine DMVR, two predictors are refined simultaneously. However, in some cases, one of the predictors may already have an acceptable level of accuracy for coding efficiency, and coding efficiency may be improved by refining the other predictor.
[0007] The present disclosure describes techniques for adaptive affine DMVR that allow for further refinement flexibility with additional signaling. Compared to other affine DMVR techniques, the adaptive affine DMVR of the present disclosure may include setting a motion vector difference for one of the reference picture lists to zero (0,0). In this way, instead of refining motion vectors from both reference lists simultaneously, only one of the predictors from a given reference list is refined. Thus, coding efficiency may be increased.
[0008] In one example, the present disclosure describes a method of decoding video data that includes: receiving a first video data block to be decoded using adaptive affine DMVR; determining to set a first motion vector difference (MVD) for a first reference picture list to zero; refining a control point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV; and decoding the first video data block using the refined CPMV.
[0009] In another example, the present disclosure describes an apparatus configured to decode video data, the apparatus including a memory and one or more processors in communication with the memory, the one or more processors being configured to: receive a first video data block to be decoded using adaptive affine DMVR; determine to set a first MVD for a first reference picture list to zero; refine a CPMV associated with a second reference picture list to generate a refined CPMV; and decode the first video data block using the refined CPMV.
[0010] In another example, the present disclosure describes an apparatus configured to decode video data, the apparatus including: a unit for receiving a first video data block to be decoded using adaptive affine DMVR; a unit for determining to set an MVD for a first reference picture list to zero; a unit for refining a CPMV associated with a second reference picture list to generate a refined CPMV; and a unit for decoding the first video data block using the refined CPMV.
[0011] In another example, the present disclosure describes a non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors to: receive a first video data block to be decoded using adaptive affine DMVR; determine to set a first MVD for a first reference picture list to zero; refine a CPMV associated with a second reference picture list to generate a refined CPMV; and decode the first video data block using the refined CPMV.
[0012] In another example, the present disclosure describes a method for encoding video data, the method including: receiving a first video data block to be encoded using adaptive affine DMVR; determining to set a first MVD for a first reference picture list to zero; refining a CPMV associated with a second reference picture list to generate a refined CPMV; and encoding the first video data block using the refined CPMV.
[0013] In another example, the present disclosure describes an apparatus configured to encode video data, the apparatus including a memory and one or more processors in communication with the memory, the one or more processors being configured to: receive a first video data block to be encoded using adaptive affine DMVR; determine to set a first MVD for a first reference picture list to zero; refine a CPMV associated with a second reference picture list to generate a refined CPMV; and encode the first video data block using the refined CPMV.
[0014] In another example, the present disclosure describes an apparatus configured to encode video data, the apparatus including: a unit for receiving a first video data block to be encoded using adaptive affine DMVR; a unit for determining to set an MVD for a first reference picture list to zero; a unit for refining a CPMV associated with a second reference picture list to generate a refined CPMV; and a unit for encoding the first video data block using the refined CPMV.
[0015] In another example, the present disclosure describes a non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform the following operations: receive a first video data block to be encoded using adaptive affine DMVR; determine to set a first MVD for a first reference picture list to zero; refine a CPMV associated with a second reference picture list to generate a refined CPMV; and encode the first video data block using the refined CPMV.
[0016] One or more example details are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the specification, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a block diagram illustrating an example video encoding and decoding system that may execute the techniques of the present disclosure.
[0018] Figure 2 is a conceptual diagram illustrating an example of bilateral matching.
[0019] Figure 3 is a conceptual diagram illustrating an example of an independent bilateral matching search for a control point motion vector.
[0020] Figure 4 is a conceptual diagram illustrating an example template and reference samples of the template in a reference picture.
[0021] Figure 5 is a conceptual diagram illustrating an example template and reference samples of a template for a block with sub-block motion.
[0022] Figure 6 is a block diagram illustrating an example video encoder that may execute the techniques of the present disclosure.
[0023] Figure 7 is a block diagram illustrating an example video decoder that may execute the techniques of the present disclosure.
[0024] Figure 8 is a flowchart illustrating an example method of encoding a current block according to the techniques of the present disclosure.
[0025] Figure 9 is a flowchart showing an example method for decoding a current block according to the techniques of the present disclosure.
[0026] Figure 10 is a flowchart showing another example method for encoding a current block according to the techniques of the present disclosure.
[0027] Figure 11 is a flowchart showing another example method for decoding a current block according to the techniques of the present disclosure. Detailed Description
[0028] Generally, the present disclosure describes techniques for encoding and decoding video data, including techniques for inter prediction. More specifically, the present disclosure describes techniques for adaptive affine decoder-side motion vector refinement (DMVR) using bilateral matching. In some examples of affine DMVR, two predictors are refined simultaneously. However, in some cases, one of the predictors may already have an acceptable level of accuracy for decoding efficiency, and the decoding efficiency can be improved by refining the other predictor.
[0029] The present disclosure describes techniques for adaptive affine DMVR that allow for further refinement flexibility with additional signaling. Compared to other affine DMVR techniques, the adaptive affine DMVR of the present disclosure can include setting the motion vector difference of one of the reference picture lists to (0,0). In this way, instead of refining the motion vectors from both reference lists simultaneously, only one of the predictors from a given reference list is refined. Thus, the decoding efficiency can be increased.
[0030] Figure 1 is a block diagram showing an example video encoding and decoding system 100 in which the techniques of the present disclosure can be implemented. The techniques of the present disclosure are generally directed to decoding (encoding and / or decoding) video data. Generally, video data includes any data for processing video. Thus, video data can include raw unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling data.
[0031] As Figure 1As shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. In particular, source device 102 provides the video data to destination device 116 via a computer-readable medium 110. Source device 102 and destination device 116 can be or include any of a variety of devices, such as a desktop computer, a notebook (i.e., laptop) computer, a mobile device, a tablet, a set-top box, a cellular phone (e.g., a smart phone), a television, a camera, a display device, a digital media player, a video game console, a video streaming device, a broadcast receiver device, and so on. In some cases, source device 102 and destination device 116 can be equipped for wireless communication and can thus be referred to as wireless communication devices.
[0032] In Figure 1 the example of, source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. In accordance with the present disclosure, video encoder 200 of source device 102 and video decoder 300 of destination device 116 can be configured to apply techniques for affine DMVR. Thus, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, source devices and destination devices can include other components or arrangements. For example, source device 102 can receive video data from an external video source such as an external camera. Similarly, destination device 116 can interface with an external display device rather than include an integrated display device.
[0033] As Figure 1 shown in system 100 is merely one example. In general, any digital video encoding and / or decoding device can perform techniques for affine DMVR. Source device 102 and destination device 116 are merely examples of such encoding devices, where source device 102 generates encoded video data for transmission to destination device 116. The present disclosure refers to an "encoding" device as a device that performs encoding (encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of encoding devices, specifically, a video encoder and a video decoder, respectively. In some examples, source device 102 and destination device 116 can operate in a substantially symmetric manner such that each of source device 102 and destination device 116 includes video encoding and decoding components. Thus, system 100 can support unidirectional or bidirectional video transmission between source device 102 and destination device 116, e.g., for video streaming, video playback, video broadcast, or video telephony.
[0034] Typically, video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a sequence of pictures (also referred to as “frames”) of the video data to video encoder 200, which encodes the data of the pictures. The video source 104 of source device 102 may include a video capture device, such as a camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As another alternative, the video source 104 may generate computer graphics-based data as the source video, or be a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 encodes the captured, pre-captured, or computer-generated video data. The video encoder 200 may reorder the pictures from the received order (sometimes referred to as “display order”) to a decoding order for decoding. The video encoder 200 may generate a bitstream including the encoded video data. Then, the source device 102 may output the encoded video data to a computer-readable medium 110 via output interface 108 for reception and / or retrieval by, for example, an input interface 122 of destination device 116.
[0035] The memory 106 of source device 102 and the memory 120 of destination device 116 represent general-purpose memories. In some examples, the memories 106, 120 may store raw video data, e.g., raw video from video source 104 and raw decoded video data from video decoder 300. Additionally or alternatively, the memories 106, 120 may store software instructions that may be executed, respectively, by, for example, video encoder 200 and video decoder 300. Although the memory 106 and the memory 120 are shown separately from the video encoder 200 and the video decoder 300 in this example, it should be understood that the video encoder 200 and the video decoder 300 may also include internal memories for functionally similar or equivalent purposes. Additionally, the memories 106, 120 may store, for example, the encoded video data output from the video encoder 200 and input to the video decoder 300. In some examples, portions of the memories 106, 120 may be allocated as one or more video buffers, e.g., for storing raw video data, decoded video data, and / or encoded video data.
[0036] Computer-readable medium 110 may represent any type of medium or device capable of transmitting the encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium for enabling source device 102 to directly send the encoded video data to destination device 116 in real time (e.g., via a radio frequency network or a computer-based network). According to a communication standard such as a wireless communication protocol, output interface 108 may modulate the transmission signal including the encoded video data, and input interface 122 may demodulate the received transmission signal. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other device that may be used to facilitate communication from source device 102 to destination device 116.
[0037] In some examples, source device 102 may output the encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include any data storage medium among various distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data.
[0038] In some examples, source device 102 may output the encoded video data to file server 114 or another intermediate storage device, which may store the encoded video data generated by source device 102. Destination device 116 may access the stored video data from file server 114 by means of streaming or downloading.
[0039] The file server 114 can be any type of server device capable of storing the encoded video data and sending the encoded video data to the destination device 116. The file server 114 can represent a web server (e.g., for a website), a server configured to provide a file transfer protocol service (such as File Transfer Protocol (FTP) or File Delivery over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Service (MBMS) or Enhanced MBMS (eMBMS) server, and / or a Network Attached Storage (NAS) device. The file server 114 can additionally or alternatively implement one or more HTTP streaming protocols, such as Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), Real-Time Streaming Protocol (RTSP), HTTP Dynamic Streaming, etc.
[0040] The destination device 116 can access the encoded video data from the file server 114 via any standard data connection, including an Internet connection. This can include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both that is suitable for accessing the encoded video data stored on the file server 114. The input interface 122 can be configured to operate according to any one or more of the various protocols discussed above for retrieving or receiving media data from the file server 114 or other such protocols for retrieving media data.
[0041] The output interface 108 and the input interface 122 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any one of the various IEEE 802.11 standards, or other physical components. In an example where the output interface 108 and the input interface 122 include wireless components, the output interface 108 and the input interface 122 can be configured to forward data such as the encoded video data according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, etc. In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 can be configured to operate according to other wireless standards (e.g., IEEE802.11 specifications, IEEE 802.15 specifications (such as ZigBee TM ) Bluetooth TMforward data such as encoded video data according to standards, etc. In some examples, the source device 102 and / or the destination device 116 may include respective system-on-a-chip (SoC) devices. For example, the source device 102 may include an SoC device to perform functions belonging to the video encoder 200 and / or the output interface 108, and the destination device 116 may include an SoC device to perform functions belonging to the video decoder 300 and / or the input interface 122.
[0042] The techniques of the present disclosure can be applied to support video coding for any of a variety of multimedia applications, such as: over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions (e.g., Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding digital video stored on a data storage medium, or other applications.
[0043] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200, which is also used by the video decoder 300, such as syntax elements having values that describe the characteristics and / or processing of video blocks or other coding units (e.g., slices, pictures, groups of pictures, sequences, etc.). The display device 118 displays the decoded pictures of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.
[0044] Although not shown in Figure 1 In some examples, the video encoder 200 and the video decoder 300 may each be integrated with an audio encoder and / or an audio decoder, and may include appropriate MUX-DEMUX units or other hardware and / or software to process a multiplexed stream including both audio and video in a common data stream.
[0045] Video encoder 200 and video decoder 300 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. When the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in a corresponding device. Devices including video encoder 200 and / or video decoder 300 may implement video encoder 200 and / or video decoder 300 in a processing circuit such as an integrated circuit and / or a microprocessor. Such devices may be wireless communication devices, such as cellular phones, or any other type of device described herein.
[0046] Video encoder 200 and video decoder 300 may operate in accordance with a video coding standard, such as ITU-T H.265, also known as High Efficiency Video Coding (HEVC) or extensions thereof, such as multi-view and / or scalable video coding extensions). Alternatively, video encoder 200 and video decoder 300 may operate in accordance with other proprietary or industry standards, such as ITU-T H.266, also known as Versatile Video Coding (VVC). In other examples, video encoder 200 and video decoder 300 may operate in accordance with a proprietary video codec / format, such as AOMedia Video 1 (AV1), extensions of AV1, and / or subsequent versions of AV1 (e.g., AV2). In other examples, video encoder 200 and video decoder 300 may operate in accordance with other proprietary formats or industry standards. However, the techniques of the present disclosure are not limited to any particular coding standard or format. Generally, video encoder 200 and video decoder 300 may be configured to perform the techniques of the present disclosure using any video coding technique that incorporates RD-MVR.
[0047] Generally, video encoder 200 and video decoder 300 may perform block-based coding of pictures. The term "block" generally refers to a structure that includes data to be processed (e.g., data to be encoded, decoded, or otherwise used during an encoding and / or decoding process). For example, a block may include a two-dimensional matrix of luminance and / or chrominance data samples. Generally, video encoder 200 and video decoder 300 may code video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, rather than coding the red, green, and blue (RGB) data of the samples of a picture, video encoder 200 and video decoder 300 may code a luminance component and a chrominance component, where the chrominance component may include both a red chrominance component and a blue chrominance component. In some examples, video encoder 200 converts the received data in RGB format to a YUV representation before encoding, and video decoder 300 converts the YUV representation to RGB format. Alternatively, a preprocessing unit and a postprocessing unit (not shown) may perform these conversions.
[0048] The present disclosure may generally relate to coding (e.g., encoding and decoding) of pictures to include a process of encoding or decoding data of a picture. Similarly, the present disclosure may relate to coding of blocks of a picture to include a process of encoding or decoding data of a block, e.g., prediction and / or residual coding. An encoded video bitstream generally includes a series of values for syntax elements that represent coding decisions (e.g., coding modes) and a partitioning of a picture into blocks. Thus, a reference to coding a picture or a block should generally be understood as coding values of syntax elements used to form the picture or the block.
[0049] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder (e.g., video encoder 200) divides a coding tree unit (CTU) into CUs according to a quadtree structure. That is, the video coder divides the CTU and the CUs into four equal and non-overlapping squares, and each node of the quadtree has zero or four child nodes. A node without child nodes may be referred to as a "leaf node", and the CU of such a leaf node may include one or more PUs and / or one or more TUs. The video coder may further divide the PUs and the TUs. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of TUs. In HEVC, a PU represents inter-prediction data, while a TU represents residual data. A CU predicted intra-frame includes intra-frame prediction information, such as an intra-mode indication.
[0050] As another example, video encoder 200 and video decoder 300 may be configured to operate according to VVC. According to VVC, a video coder (such as video encoder 200) divides a picture into multiple CTUs. Video encoder 200 may divide a CTU according to a tree structure such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple partitioning types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level divided according to quadtree partitioning, and a second level divided according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to CUs.
[0051] In the MTT partitioning structure, a block may be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) (also referred to as a trinary tree (TT)) partitioning. Ternary tree or trinary tree partitioning is a partitioning that divides a block into three sub-blocks. In some examples, the ternary tree or trinary tree partitioning divides a block into three sub-blocks without dividing the original block through the center. The partitioning types in the MTT (e.g., QT, BT, and TT) may be symmetric or asymmetric.
[0052] When operating according to the AV1 codec, video encoder 200 and video decoder 300 may be configured to decode video data in a block. In AV1, the largest decoding block that can be processed is referred to as a superblock. In AV1, a superblock may be 128x128 luma samples or 64x64 luma samples. However, in a subsequent video coding format (e.g., AV2), a superblock may be defined by a different (e.g., larger) luma sample size. In some examples, a superblock is the top level of a block quadtree. Video encoder 200 may further divide a superblock into smaller decoding blocks. Video encoder 200 may divide a superblock and other decoding blocks into smaller blocks using square or non-square partitioning. Non-square blocks may include N / 2xN, NxN / 2, N / 4xN, and NxN / 4 blocks. Video encoder 200 and video decoder 300 may perform separate prediction and transform processing on each decoding block.
[0053] AV1 also defines tiles of video data. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, video encoder 200 and video decoder 300 may encode and decode the decoding blocks within a tile without using video data from other tiles. However, video encoder 200 and video decoder 300 may perform filtering across tile boundaries. The size of a tile may be uniform or non-uniform. Tile-based decoding may enable parallel processing and / or multi-threading for encoder and decoder implementations.
[0054] In some examples, video encoder 200 and video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, video encoder 200 and video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for the two chrominance components (or two QTBT / MTT structures for the respective chrominance components).
[0055] Video encoder 200 and video decoder 300 may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning, or other partitioning structures.
[0056] In some examples, a CTU includes: a coding tree block (CTB) of the luminance samples of a picture having three sample arrays, two corresponding CTBs of the chrominance samples, or a CTB of the samples of a monochrome picture or a picture decoded using three separate color planes; and a syntax structure for coding the samples. A CTB may be an N×N block of samples for some N value such that partitioning the component into CTBs is a "partitioning". A component is an array or a single sample from one of the three arrays (luminance and two chrominances) that make up a picture in a 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample of an array that makes up a picture in a monochrome format. In some examples, a coding block is an M×N sample block for some M and N values such that partitioning the CTB into coding blocks is a "partitioning".
[0057] Blocks (e.g., CTUs or CUs) may be grouped in various ways in a picture. As an example, a brick may refer to a rectangular region of CTU rows within a particular tile in a picture. A tile may be a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile column refers to a rectangular region of CTUs having a height equal to the height of the picture and a width specified by a syntax element (e.g., such as in a picture parameter set). A tile row refers to a rectangular region of CTUs having a height specified by a syntax element (e.g., such as in a picture parameter set) and a width equal to the width of the picture.
[0058] In some examples, a tile may be divided into multiple bricks, each of which may include one or more CTU rows within the tile. A tile that is not divided into multiple bricks may also be referred to as a brick. However, a brick that is a proper subset of a tile may not be referred to as a tile. Bricks in a picture may also be arranged in slices. A slice may be an integer number of bricks of a picture that can be exclusively included in a single network abstraction layer (NAL) unit. In some examples, a slice includes a contiguous sequence of multiple complete tiles or complete bricks of only one tile.
[0059] The present disclosure may interchangeably use "NxN" and "N by N" to refer to the sample size of a block (such as a CU or other video block) in terms of the vertical and horizontal dimensions. For example, 16x16 samples or 16 by 16 samples. Generally, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N×N CU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non - negative integer value. The samples in a CU can be arranged in rows and columns. Additionally, a CU does not necessarily have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU can include N×M samples, where M does not necessarily equal N.
[0060] The video encoder 200 encodes video data representing prediction information and / or residual information of a CU, as well as other information. The prediction information indicates how to predict the CU in order to form a prediction block of the CU. The residual information generally represents the sample - by - sample difference between the samples of the CU before encoding and the prediction block.
[0061] To predict a CU, the video encoder 200 generally forms a prediction block for the CU through inter - frame prediction or intra - frame prediction. Inter - frame prediction generally refers to predicting the CU based on the data of previously decoded pictures, while intra - frame prediction generally refers to predicting the CU based on the previously decoded data of the same picture. To perform inter - frame prediction, the video encoder 200 can use one or more motion vectors to generate the prediction block. The video encoder 200 can generally perform a motion search, for example, according to the difference between the CU and a reference block, to identify a reference block that closely matches the CU. The video encoder 200 can use the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to calculate a difference metric to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use uni - directional prediction or bi - directional prediction to predict the current CU.
[0062] Some examples of VVC also provide an affine motion compensation mode, which can be regarded as an inter - frame prediction mode. In the affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non - translational motion (such as zooming in or out, rotation, perspective motion, or other irregular motion types).
[0063] To perform intra prediction, video encoder 200 may select an intra prediction mode to generate a prediction block. Some examples of VVC provide 67 intra prediction modes, including various directional modes, as well as planar mode and DC mode. Generally, video encoder 200 selects an intra prediction mode that describes the neighboring samples of the current block according to which the samples of the current block (e.g., the block of a CU) are predicted. Assuming that video encoder 200 decodes CTUs and CUs in raster scan order (from left to right, top to bottom), such samples are typically above, top-left, or to the left of the current block in the same picture as the current block.
[0064] Video encoder 200 encodes data representing the prediction mode of the current block. For example, for an inter prediction mode, video encoder 200 may encode data indicating which one of various available inter prediction modes is used and the motion information of the corresponding mode. For uni-directional or bi-directional inter prediction, for example, video encoder 200 may use advanced motion vector prediction (AMVP) or merge mode to encode the motion vectors. Video encoder 200 may use a similar mode to encode the motion vectors for affine motion compensation mode.
[0065] AV1 includes two general techniques for encoding and decoding decoded blocks of video data. These two general techniques are intra prediction (e.g., intra prediction or spatial prediction) and inter prediction (e.g., inter prediction or temporal prediction). In the context of AV1, when using an intra prediction decoding mode to predict a block of the current frame of video data, video encoder 200 and video decoder 300 do not use video data from other frames of the video data. For most intra prediction modes, video encoder 200 encodes the block of the current frame based on the difference between the sample values in the current block and the predicted values generated from reference samples in the same frame. Video encoder 200 determines the predicted values generated from the reference samples based on the intra prediction decoding mode.
[0066] After prediction such as intra prediction or inter prediction of a block, video encoder 200 may compute residual data for the block. The residual data (e.g., residual block) represents the per-sample difference between the block and a predicted block of the block formed using the corresponding prediction mode. Video encoder 200 may apply one or more transforms to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, video encoder 200 may apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, video encoder 200 may apply a secondary transform after the first transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), etc. Video encoder 200 produces transform coefficients after applying one or more transforms.
[0067] As described above, after any transform used to produce transform coefficients, video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to the process of quantizing the transform coefficients to possibly reduce the amount of data used to represent the transform coefficients, thereby providing further compression. By performing the quantization process, video encoder 200 may reduce the bit depth associated with some or all of the transform coefficients. For example, video encoder 200 may round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, video encoder 200 may perform a bitwise right shift on the value to be quantized.
[0068] After quantization, video encoder 200 may scan the transform coefficients to produce a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place the higher energy (and thus lower frequency) transform coefficients at the front of the vector and the lower energy (and thus higher frequency) transform coefficients at the back of the vector. In some examples, video encoder 200 may use a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy encode the quantized transform coefficients of the vector. In other examples, video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 200 may entropy encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). Video encoder 200 may also entropy encode the values of syntax elements that describe metadata associated with the encoded video data for use by video decoder 300 when decoding the video data.
[0069] To perform CABAC, video encoder 200 may assign a context within a context model to the symbol to be sent. The context may relate to, for example, whether the neighboring values of the symbol are zero values. Probability determination may be based on the context assigned to the symbol.
[0070] The video encoder 200 can also generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, or other syntax data, such as sequence parameter sets (SPS), picture parameter sets (PPS), or video parameter sets (VPS), for example, in a picture header, a block header, or a slice header, to the video decoder 300. The video decoder 300 can similarly decode such syntax data to determine how to decode the corresponding video data.
[0071] In this way, the video encoder 200 can generate a bitstream including the encoded video data, for example, describing syntax elements that divide a picture into blocks (e.g., CUs) and prediction information and / or residual information for the blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.
[0072] Generally, the video decoder 300 performs a reciprocal process of the process performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 can use CABAC to decode the values of the syntax elements of the bitstream in a way that is reciprocal but substantially similar to the CABAC encoding process of the video encoder 200. The syntax elements can define partitioning information for dividing a picture into CTUs and for dividing each CTU according to a corresponding partitioning structure (e.g., QTBT structure) to define the CUs of the CTU. The syntax elements can further define prediction information and residual information for blocks (e.g., CUs) of the video data.
[0073] The residual information can be represented by, for example, quantized transform coefficients. The video decoder 300 can inverse-quantize and inverse-transform the quantized transform coefficients of the block to reproduce the residual block for the block. The video decoder 300 uses the signaling sent prediction mode (intra prediction or inter prediction) and the associated prediction information (e.g., motion information for inter prediction) to form a prediction block for the block. The video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. The video decoder 300 can perform additional processing, such as performing deblocking processing to reduce visual artifacts along the boundaries of the blocks.
[0074] The present disclosure may generally refer to "signaling" certain information, such as syntax elements. The term "signaling" may generally refer to the communication of the values of syntax elements and / or other data values used to decode the encoded video data. That is, the video encoder 200 may signal the values of syntax elements in a bitstream. Generally, signaling refers to generating values in a bitstream. As described above, the source device 102 may transmit the bitstream to the destination device 116 substantially in real time or non-real time (such as may occur when storing the syntax elements in the storage device 112 for later retrieval by the destination device 116).
[0075] According to the techniques of the present disclosure, as will be explained in more detail below, the video encoder 200 and the video decoder 300 may be configured to decode video data using an adaptive affine DMVR mode. For example, the video encoder 200 and the video decoder 300 may be configured to: receive a first video data block to be decoded using adaptive affine DMVR; determine to set a first motion vector difference (MVD) for a first reference picture list to zero; refine a control point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV; and decode the first video data block using the refined CPMV.
[0076] Bilateral matching
[0077] Bilateral matching (BM) is a technique in which the video encoder 200 and the video decoder 300 may be configured to refine a pair of two initial motion vectors: MV0 and MV1. Generally, the BM technique includes searching around the regions of the reference pictures pointed to by MV0 and MV1 to derive refined MVs, MV0' and MV1', that minimize a block matching cost. The block matching cost measures the similarity between two motion compensated predictors generated by the two MVs. Some typical criteria for the block matching cost are the sum of absolute differences (SAD), the sum of absolute transform differences (SATD), the sum of squared errors (SSE), etc. The bilateral matching cost may also include a regularization term derived based on the MV difference between the current MV pair and the initial MV pair. Some specific constraints may also be applied to the MV difference (MVD) between MVD0 (MV0' - MV0) and MVD1 (MV1' - MV1). Generally, BM is applied using the assumption that the use of MVD0 and MVD1 should be proportional to the temporal distance (TD) between the current picture and the reference pictures pointed to by the two MVs. However, in some applications, BM is applied using the assumption that MVD0 is equal to MVD1.
[0078] Affine mode
[0079] The affine motion model may be described by the following equation:
[0080]
[0081] where (v x , v y ) is the motion vector at coordinates (x, y), and a, b, c, d, e, and f are six affine parameters. This disclosure refers to the above affine motion model as a 6-parameter affine motion model.
[0082] In a typical video decoder, a picture is partitioned into blocks for block-based decoding. The affine motion model for a block can also be described by three motion vectors (MVs) at three different positions that are not in the same line and . These three positions are typically referred to as control points, and the three motion vectors are referred to as control point motion vectors (CPMVs). In the case where the three control points are at the three corners of the block, the affine motion can be described as follows:
[0083]
[0084] where blkW and blkH are the width and height of the block.
[0085] In the affine mode, different motion vectors can be derived for each pixel in the block according to the associated affine motion model. Thus, motion compensation can be performed in a pixel-by-pixel manner. However, to reduce the complexity of the affine mode, sub-block-based motion compensation can be used, where the block is partitioned into multiple sub-blocks (where each sub-block has a smaller block size than the original block), and each sub-block is associated with one motion vector for block-based motion compensation. The motion vector for each sub-block is derived using the representative coordinates of the sub-block. Typically, the center position is used as the representative coordinate.
[0086] In one example, the block is partitioned into non-overlapping sub-blocks. The block width is blkW, the block height is blkH, the sub-block width is sbW, and the sub-block height is sbH. In this example, there are blkH / sbH rows of sub-blocks and blkW / sbW sub-blocks in each row. For the 6-parameter affine motion model, the motion vector for the sub-block (referred to as sub-block MV) in the i-th row (0 <= i < blkW / sbW) and the j-th column (0 <= j < blkH / sbH) is derived as follows:
[0087]
[0088] The sub-block motion vector (MV) is rounded to a predefined precision and stored in a motion buffer for motion compensation and motion vector prediction.
[0089] The simplified 4-parameter affine model (for scaling and rotation motion) is described as follows:
[0090]
[0091] Similar to the 6-parameter affine model, the 4-parameter affine model for a block can be described by two CPMVs: and at two corners of the block (usually the upper left corner and the upper right corner) and Then the motion field is described as follows:
[0092]
[0093] The motion vector (MV) of the sub-block at the i-th row and j-th column is derived as follows:
[0094]
[0095] DMVR in VVC
[0096] In VVC, the video encoder 200 and the video decoder 300 can use BM-based DMVR to increase the accuracy of the MVs of the bi-prediction merge candidates. The BM-based DMVR method calculates the SAD between two candidate blocks in the reference picture lists L0 and L1. Figure 2 is a conceptual diagram illustrating an example of bilateral matching. As Figure 2 shown, the video encoder 200 and the video decoder 300 can be configured to calculate the SAD between block 500 and block 502 based on each MV candidate around the initial MV. The SAD can be referred to as distortion cost calculation or bilateral matching cost calculation. The MV candidate with the lowest SAD becomes the refined MV and is used to generate a bi-predicted signal. In some examples, 1 / 4 of the SAD value is subtracted from the SAD of the initial MV to be used as a regularization term. In Figure 2 example, the temporal distance from the two reference pictures to the current picture (e.g., the picture order count (POC) difference) should be the same, so the motion vector difference (MVD0) of MV0 is only the opposite sign of the motion vector difference (MVD1) of MV1.
[0097] The refined search range is two integer luminance samples from the initial MV. The search includes an integer sample offset search stage and a fractional sample refinement stage. A 25-point full search is applied to the integer sample offset search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample stage of DMVR terminates. Otherwise, the SADs of the remaining 24 points are calculated and checked in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search stage.
[0098] After integer sample search, fractional sample refinement is performed. To reduce the computational complexity, fractional sample refinement can be derived by using the parametric error surface equation instead of an additional search with SAD comparison. In one example, fractional sample refinement is conditionally invoked based on the output of the integer sample search stage. When the integer sample search stage terminates at the center with the minimum SAD in the first iterative search or the second iterative search, fractional sample refinement is further applied.
[0099] In the sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at four adjacent positions from the center are used to fit a 2-D parabolic error surface equation of the following form:
[0100] E(x,y) = A(x - x min ) 2 + B(y - y min ) 2 + C, (1)
[0101] where (x min , y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. The cost values of five search points are used to solve the above equation, and (x min , y min ) is calculated as:
[0102] x min = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0))), (2)
[0103] y min = (E(0,-1) - E(0,1)) / (2((E(0,-1) + E(0,1) - 2E(0,0))), (3)
[0104] Since all cost values are positive and the minimum value is E(0,0), the values of x min and y min are automatically constrained between -8 and 8. This corresponds to a half-pixel offset with 1 / 16-pel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer distance refinement MV to obtain a sub-pixel accurate refinement increment MV.
[0105] In VVC, the resolution of the MV is 1 per 16 luma samples. An 8-tap interpolation filter is used to interpolate samples at fractional positions. In DMVR, the search points are centered around the initial fractional pixel MV with integer sample offsets, so the samples at those fractional positions are interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another effect is that by using a bilinear filter with a 2-sample search range, the DMVR process does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV using the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To not access more reference samples compared to the normal motion compensation process, samples that are not required for the interpolation process based on the original MV but are required for the interpolation process based on the refined MV can be filled from those available samples.
[0106] When the width and / or height of the CU is greater than 16 luma samples, the CU can be further split into sub-blocks with width and / or height equal to 16 luma samples for the DMVR process.
[0107] In VVC, DMVR can be applied to CUs decoded using the following modes and characteristics:
[0108] - CU-level merge mode with dual-prediction MVs
[0109] - One reference picture is in the past and the other reference picture is in the future relative to the current picture
[0110] - The distances from the two reference pictures to the current picture (i.e., POC differences) are the same
[0111] - Both reference pictures are short-term reference pictures
[0112] - The CU has more than 64 luma samples
[0113] - Both the CU height and the CU width are greater than or equal to 8 luma samples
[0114] - The dual-prediction indication with CU-level weight (BCW) weight index indicates equal weights
[0115] - Weighted prediction (WP) is not enabled for the current block
[0116] - Combined inter-intra prediction (CIIP) mode is not used for the current block
[0117] Adaptive DMVR
[0118] The general idea of adaptive DMVR is to configure the video encoder 200 and the video decoder 300 to use different search strategies and / or methods for different decoding blocks for bilateral matching. The selected search strategy for a block is signaled as one or more syntax elements decoded in the bitstream. The search strategy includes the constraints / relationships between MVD0 and MVD1 imposed during the bilateral matching search process.
[0119] For each bilateral matching block, one of the following constraints between MVD0 and MVD1 is selected:
[0120] 1) Mirrored MVD: MVD0 and MVD1 have the same magnitude but opposite signs, i.e., MVD0 = -MVD1 (conventional DMVR);
[0121] 2) MVD0 is zero (both x and y components are zero), i.e., when performing a search around MV1 to derive the refined MV1', MV0 is fixed and MV0' is equal to MV0 (adaptive DMVR);
[0122] 3) MVD1 is zero, i.e., when performing a search around MV0 to derive the refined MV0', MV1 is fixed and MV1' is equal to MV1 (adaptive DMVR).
[0123] The video encoder 200 can signal a first syntax element representing mode information (e.g., whether conventional DMVR or adaptive DMVR should be applied). The above three options are classified by the first syntax element. Option 1) applies conventional DMVR to the decoding block when the conventional merge candidate satisfies the DVMR condition, and options 2) or 3) are applied when the decoding block uses a specified new merge mode, where all candidates should also satisfy the specified DMVR condition. Constraints 2) and 3) are further distinguished by a mode flag or a merge index.
[0124] DMVR for affine merge mode
[0125] The affine DMVR design can be summarized in the following steps:
[0126] 1) Split the current block into sub-blocks.
[0127] 2) Generate initial motion vectors (for two prediction directions) for each sub-block (sub-block motion field) according to the initial affine motion model.
[0128] 3) Loop over each sub-block, calculating the sub-block bilateral matching costs for all possible offsets.
[0129] 4) For each possible offset, accumulate the sub-block bilateral matching costs to generate the bilateral matching cost corresponding to the entire block.
[0130] 5) The optimal offset is determined by selecting the offset with the minimum bilateral matching cost corresponding to the entire block.
[0131] In this way, the sub-block motion field is generated only once, rather than once for each candidate offset.
[0132] The sub-block size in the above process is determined based on the affine parameters. The affine parameters reflect the change of the per-pixel motion vector in the affine decoded block. Generally, when the affine parameters are small, a larger sub-block size is used, and vice versa.
[0133] Given the offset and the initial motion vector (generated in step 2 above), the candidate motion vectors can be derived. Pre-interpolation is applied in one step to generate predictors for all possible offsets, which reduces the complexity. Bilinear interpolation is used instead of the 8-tap (6-tap or 12-tap) interpolation filter usually used for final motion compensation.
[0134] After step 5), sub-pixel offset estimation based on the parametric error surface is also applied to generate the sub-pixel offset.
[0135] In the Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, the 28th meeting, Mainz, DE, October 20 - 28, 2022 (hereinafter referred to as "JVET-AB0112"), the "EE2-2.6: DMVR for affine merge coded block" by Jie Chen et al., an affine DMVR technique following the above process was proposed.
[0136] In the Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, the 28th meeting, Mainz, DE, October 20 - 28, 2022 (hereinafter referred to as "JVET-AB0177"), the "EE2-related Sub-block processing for affine DMVR" by Han Huang et al., some simplifications of affine DMVR were described. For example, instead of using each sub-block of the sub-blocks in affine DMVR, only a subset of the sub-blocks is used. In addition, a regression-based affine merge candidate derivation method can be applied based on the affine DMVR search results.
[0137] In the Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, the twenty-eighth meeting, Mainz, DE, October 20-28, 2022 (hereinafter referred to as "JVET-AB0178"), the affine DMVR search method based on CPMV was described in "EE2-related: Control-point motion vector refinement for Affine DMVR" by Han Huang et al. Given the initial control-point motion vector initCpMvLX[cpIdx], where cpIdx = 0 ··· numCpMv-1, where numCpMv is the number of CPMVs of the current affine decoding block. The proposed method can be described in the following steps:
[0138] 1) As Figure 3 shown, for each control point, bilateral matching is performed on the block 510 centered on the control point in order to derive the refined CPMV bmRefinedCpMvLx[cpIdx]. Figure 3 is a conceptual diagram showing an example of an independent bilateral matching search for control-point motion vectors.
[0139] 2) Loop through the combination of initCpMvLX[cpIdx] and bmRefinedCpMvLx[cpIdx], and derive the best set of CPMVs that minimizes the bilateral matching cost of the current block.
[0140] 3) Iteratively further refine the CPMV to minimize the bilateral matching cost of the current block. In each iteration, one CPMV is refined while the other CPMVs are fixed.
[0141] Adaptive reordering of merge candidates (ARMC)
[0142] In the currently studied Enhanced Compression Model (ECM), template matching (TM) is used to adaptively reorder merge candidates. The reordering method is applied to the regular merge candidate list, the TM merge candidate list, and the affine merge candidate list (excluding the sub-block merge candidate list for sub-block temporal motion vector predictors (SbTMVP) candidates). In some examples of ECM, multiple SbTMVP candidates can be included in the candidate list, and these multiple SbTMVP candidates can be reordered using the ARMC process. For the TM merge mode, the merge candidates are reordered before the TM refinement process.
[0143] After constructing the merge candidate list, the TM cost of a merge candidate is measured by the sum of absolute differences (SAD) between samples of the template of the current block and their corresponding reference samples. The template of the current block includes a set of reconstructed samples adjacent to the current block. The reference samples of the template are located by the motion information of the merge candidate.
[0144] When a merge candidate utilizes bi-prediction, the reference samples of the template for generating the merge candidate are also generated by bi-prediction, as Figure 4 shown. Figure 4 is a conceptual diagram illustrating an example template and reference samples of the template in the reference picture. The current block 600 of the current picture 602 has an upper template 604 and a left template 606. The motion vectors of the merge candidates in reference list 0 are used to identify the reference block 610 of the reference picture 612 in reference list 0. The reference samples (RT0) of the upper template 614 and the left template 616 are obtained. Similarly, the motion vectors of the merge candidates in reference list 1 are used to identify the reference block 620 of the reference picture 622 in reference list 1. The reference samples (RT1) of the upper template 624 and the left template 626 are obtained. The RT0 reference samples and the RT1 references are combined and averaged to obtain a single template. The sum of absolute differences (SAD) between the current template T and the single template formed by the RT0 and RT1 samples can be calculated. The video encoder 200 and the video decoder 300 can use the SAD value to determine whether the template is well-matched.
[0145] For sub-block based merge candidates with sub-block size equal to Wsub×Hsub, the upper template includes a number of sub-templates with size Wsub×1, and the left template includes a number of sub-templates with size 1×Hsub. Figure 5 is a conceptual diagram illustrating an example template and reference samples for sub-block motion. As Figure 5 shown, the motion information of the sub-blocks in the first row (A, B, C, D) and the first column (A, E, F, G) of the current block 700 of the current picture 702 is used to derive the reference samples of each sub-template 710 and 712. The sub-template 710 is represented by the black box above the reference blocks A ref, B ref, C ref, and D ref in the reference picture 714. The reference blocks of the sub-template 710 are identified by the motion vectors of the sub-blocks A, B, C, D relative to the juxtaposed block 716. The sub-template 712 is represented by the black box to the left of the reference blocks A ref, Eref, F ref, and G ref in the reference picture 714. The reference blocks of the sub-template 712 are identified by the motion vectors of the sub-blocks A, E, F, G relative to the juxtaposed block 716.
[0146] Based on the TM cost of each merge candidate, the merge candidate list is re-sorted in ascending order.
[0147] General example of adaptive affine DMVR
[0148] In the present disclosure, the video encoder 200 and the video decoder 300 may be configured to decode video data according to an adaptive affine DMVR process. The adaptive affine DMVR of the present disclosure allows for further refinement flexibility for affine DMVR with additional signaling. Compared with other affine DMVR techniques, the general idea of the adaptive affine DMVR of the present disclosure is to set the MVD of one of the reference lists to (0, 0). In this way, instead of refining both reference lists simultaneously, only one of the predictors from a given reference list is refined.
[0149] The adaptive affine DMVR may be an extension of the above-described conventional adaptive DMVR method. However, when refining the MVD to determine the refined CPMV, the BM cost is derived for each sub-block in the sub-block rather than for the entire CU, and the BM cost is accumulated for each sub-block for the final cost to determine the optimal MVD. Additional decoding benefits can be observed because in some cases, one of the predictors may already be accurate and only the other predictor needs to be refined. Together with the original affine decoder-side motion vector refinement, for each affine merge candidate that satisfies the DMVR condition, in one example, a total of three different refinement options may be provided:
[0150] 1) The original affine DMVR - MVD is mirrored: MVD0 and MVD1 have the same magnitude but opposite signs, i.e., MVD0 = -MVD1;
[0151] 2) Adaptive affine DMVR 1 - MVD0 is zero (both x and y components are zero). For each sub-block in the affine CU, MV0 is fixed when deriving the BM cost within the search range of MV1. The accumulated BM cost is used to determine the final MVD1 for each CPMV of reference list 1;
[0152] 3) Adaptive affine DMVR 2 - MVD1 is zero (both x and y components are zero). For each sub-block in the affine CU, MV1 is fixed when deriving the BM cost within the search range of MV0. The accumulated BM cost is used to determine the final MVD0 for each CPMV of reference list 0;
[0153] By introducing multiple refinement options, signaling using additional syntax can be used. In one example, the video encoder 200 may encode a first syntax element to indicate whether to use option 1) (e.g., affine DMVR with mirrored MVD). If the first syntax indicates not to use option 1), the video encoder 200 may encode a second syntax element to indicate whether to use option 2) or 3) (MVD0 is zero or MVD1 is zero). The video decoder 300 may decode and parse the first syntax element and the second syntax element to determine the type of affine DMVR decoding to perform.
[0154] In a general example of the present disclosure, the video encoder 200 and the video decoder 300 may receive a video data block to be decoded (e.g., encoded or decoded) using adaptive affine DMVR. The video encoder 200 and the video decoder 300 may determine to set a first MVD for a first reference picture list to zero. In one example, the first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1. In another example, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
[0155] The video encoder 200 and the video decoder 300 may be configured to decode a first syntax element indicating to use adaptive affine DMVR (e.g., according to options 2) and 3) above) to decode a first video data block. In one example, to determine to set the first MVD for the first reference picture list to zero, the video decoder 300 may decode a second syntax element indicating to set the first MVD for the first reference picture list to zero. In this example, the second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1.
[0156] The video encoder 200 and the video decoder 300 may then refine the CPMV associated with the second reference picture list to generate a refined CPVM. To perform the refinement, the video encoder 200 and the video decoder 300 may use the affine DMVR techniques described above or any affine DMVR techniques described below to refine the second MVD for the second reference picture list. For example, the video encoder 200 and the video decoder 300 may determine the bilateral matching (BM) cost within the search range of the motion vector for each sub-block in the first video data block, accumulate the BM costs for multiple sub-blocks of the first video data block to generate an accumulated BM cost, determine the second MVD of the CPMV for the second reference picture list based on the accumulated BM cost, and determine the refined CPMV based on the second MVD.
[0157] The video encoder 200 and the video decoder 300 can then decode the first video data block using the refined CPMV. In this context, the decoding is bidirectional decoding. The refined CPMV is for the second reference picture list, while the original CPMV (e.g., because the first MVD is set to zero) is for the first reference picture list.
[0158] Syntax signaling
[0159] In one example, the video encoder 200 can be configured to signal the above-mentioned adaptive affine DMVR mode (adaptive_aff_bm_mode) as an additional affine merge mode to a regular affine merge mode. Various signaling methods can be applied. In one example, adaptive_aff_bm_mode is considered a variant of the regular affine merge mode. The video encoder 200 can first signal a syntax element to indicate the regular merge mode, and then signal an additional flag to indicate whether the merge mode is adaptive_aff_bm_mode.
[0160] In another example, the video encoder 200 can signal a first flag to indicate whether to use the regular affine merge mode. The video encoder 200 can also signal a second flag to indicate whether the affine merge candidate is an affine MMVD candidate. In this example, adaptive_aff_bm_mode is signaled after the affine MMVD flag to indicate whether a regular affine merge candidate or an adaptive affine merge candidate should be derived.
[0161] In another example, adaptive_aff_bm_mode is indicated by a flag before the indication of the regular affine merge mode. If the syntax indicates that the current block does not use adaptive_aff_bm_mode, other syntax elements are signaled to indicate that the merge mode is the regular affine merge mode or the affine MMVD mode.
[0162] The video encoder 200 can signal the merge index in adaptive_aff_bm_mode using the same signaling method as in the regular affine merge mode. In one example, the same context model is used to decode the merge index. In another example, if the mode is adaptive_aff_bm_mode, a separate context model is used. The maximum number of merge candidates can be different for adaptive_aff_bm_mode and the regular affine merge mode.
[0163] Some high-level syntax elements can be used to indicate whether the adaptive_aff_bm_mode can be applied. In one example, the same high-level syntax that controls the enabling / disabling of the regular affine DMVR is also used to control the enabling / disabling of the adaptive_aff_bm_mode. In another example, a separate high-level syntax is used to control the enabling / disabling of the adaptive_aff_bm_mode. In another example, a separate high-level syntax is used to control the enabling / disabling of the adaptive_aff_bm_mode, but the high-level syntax exists only when the regular affine DMVR is enabled. If the high-level syntax of the regular DMVR indicates that the regular affine DMVR is disabled, the high-level syntax for controlling the enabling / disabling of the adaptive_aff_bm_mode is not signaled and is inferred to be disabled. The signaling for the corresponding high-level syntax can be in the SPS, PPS, picture header, slice header, or other syntax structures.
[0164] Syntax signaling reduction
[0165] In some examples, the signaling for the above syntax can be saved (e.g., not signaled). Instead, the TM cost information or BM cost information can be used to determine the use of the adaptive affine DMVR. To save the signaling for the second syntax, in one example, two TM costs are derived before performing the affine DMVR or the adaptive affine DMVR. The TM cost between the template of the current block and the template of the predictor from reference list 0 is defined as TM0, and the TM cost between the template of the current block and the template of the predictor from reference list 1 is defined as TM0. If TM1 < TM0, only options 1) and 2) are checked, and only one syntax is needed to indicate which option is used. Otherwise, only options 1) and 3) are checked, and similarly only one syntax is needed.
[0166] In a second example, all three options are first checked. Then, using the refined affine motion vectors, a dual-prediction reference template is derived, and the TM cost between the current template and the dual-reference template is calculated for both options 2) and 3). By comparing the TM costs of options 2) and 3), one of the options is selected, and thus only one syntax element needs to be signaled.
[0167] In a third example, instead of using the TM cost as in the second example, the minimum BM cost in the adaptive affine DMVR refinement search process is used and compared between options 2) and 3) to decide which option will be used. In the case of excluding one of the options, the number of required syntaxes is reduced from 2 to 1.
[0168] The first syntax signaling can also be saved by using the TM cost or the BM cost. In one example, the minimum BM cost of all three options during the DMVR search process is recorded. After checking all three options, only the option with the minimum BM cost is retained. In another example, instead of using the BM cost, the TM cost is used to determine which of the three options will be retained.
[0169] Alternative affine merge list for adaptive affine DMVR
[0170] For adaptive affine DMVR, since there already exists a first syntax element for indicating whether affine DMVR or adaptive affine DMVR will be applied, a separate affine merge list including only affine merge candidates that meet the affine DMVR conditions can be constructed. In one example, after constructing the affine merge list, those candidates that meet the adaptive affine DMVR conditions can be added to a separate list dedicated only to the adaptive affine DMVR process. By excluding single-predicted or those candidates that do not meet the adaptive affine DMVR conditions, the merge index to be signaled will potentially be smaller, and thus the signaling overhead can be reduced. ARMC can be further applied to the adaptive affine merge list before refining each candidate in the list to further reduce the signaling overhead.
[0171] In a second example, alternatively, the adaptive affine merge candidates are added to the adaptive bilateral matching regular merge candidate list. When scanning to construct the adaptive bilateral matching regular merge candidate list, the affine flag is additionally checked. If the adjacent decoded block is decoded in the affine mode, then instead of adding translational inter-frame merge candidates, affine merge candidates are added. Thus, depending on whether the merge candidate is a translational inter-frame merge candidate or an affine merge candidate, the adaptive DMVR or the affine adaptive DMVR process can be applied accordingly.
[0172] When an alternative merge list is used for adaptive affine DMVR, the same method as described above can also be used to save the second syntax element by using, for example, the TM cost. Additionally, the merge index to be signaled can also be modified to be smaller based on the TM or BM cost. For example, in one example, an adaptive affine merge list with a list size of M is constructed. Using the two adaptive affine DMVR options 2) and 3), a total of 2M refined candidates can be generated. These 2M candidates can be collected into a single list, and ARMC reordering is performed using the TM cost, and only the first M candidates in the list are retained. In this way, we can skip the signaling of the second syntax element and signal a smaller merge index.
[0173] Alternative search pattern
[0174] In affine DMVR, a square search pattern is used for integer search, and a parametric error surface is followed for sub-pixel search. For adaptive affine DMVR, in one example, the same search pattern is used. In a second example, alternatively, an exhaustive search is used. For sub-pixel search, in one example, the same parametric error surface method is used. In a second example, a diamond search pattern is used. In yet a third example, sub-pixel search is skipped.
[0175] Alternative search range
[0176] In one example of affine DMVR, the search range is set to 3 pixels. In one example, the same search range used for affine DMVR is used for the adaptive affine DMVR technology of the present disclosure. In another example, an alternative search range is used only for adaptive affine DMVR, such as 4 pixels.
[0177] Alternative cost metric
[0178] In affine DMVR, depending on the block size, SAD or mean-removed SAD (MRSAD) is used. For adaptive affine DMVR, alternatively, an alternative cost metric can be used. In one example, SATD is alternatively used. In a second example, SSE is used. In other examples, the adaptive DMVR technology uses the same cost metrics as affine DMVR, including SAD, MRSAD, SATD, or SSE.
[0179] Combination with affine DMVR and further refinement
[0180] As described above, different affine DMVR designs can be applied to pictures of video data. Adaptive affine DMVR generally follows the same process when refining affine merge candidates, except that an MVD is added to calculate the BM cost. Thus, in one example, adaptive affine DMVR uses the same design as affine DMVR. For example, both use the design from JVET-AB0177. In yet another example, affine DMVR uses the design from JVET-AB0112, while adaptive affine DMVR uses the design from JVET-AB0177.
[0181] Furthermore, as described above, further improvements after affine DMVR can be possible. Similar further refinement schemes can also be applied to adaptive affine DMVR. In one example, a regression-based affine merge candidate derivation method is applied to the adaptive affine DMVR output. In yet another example, a CPMV-based affine DMVR search method is applied to the adaptive affine DMVR result. In another example, two further refinement methods are sequentially applied to adaptive affine DMVR.
[0182] Figure 6 is a block diagram showing an example video encoder 200 that can implement the techniques of the present disclosure. The provision Figure 6 is for illustrative purposes and should not be considered a limitation on the techniques broadly illustrated and described in the present disclosure. For purposes of explanation, the present disclosure describes the video encoder 200 in terms of the techniques of VVC and HEVC. However, the techniques of the present disclosure may be performed by video coding devices configured for other video coding standards and video coding formats, such as successors to the AV1 and AV1 video coding formats.
[0183] In Figure 6 the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any one or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy coding unit 220 may be implemented in one or more processors or in processing circuitry. For example, the units of the video encoder 200 may be implemented as one or more circuits or logic elements that are part of a hardware circuit or as part of a processor, ASIC, or FPGA. Additionally, the video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.
[0184] The video data memory 230 may store video data to be encoded by components of the video encoder 200. The video encoder 200 may receive the video data stored in the video data memory 230 from, for example, a video source 104 ( Figure 1 ). The DPB 218 may be used as a reference picture memory that stores reference video data for use by the video encoder 200 in predicting subsequent video data. The video data memory 230 and the DPB 218 may be formed of any of a variety of storage devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The video data memory 230 and the DPB 218 may be provided by the same storage device or by separate storage devices. In various examples, the video data memory 230 may be on-chip with other components of the video encoder 200 (as shown) or off-chip relative to those components.
[0185] In the present disclosure, a reference to the video data memory 230 should not be construed as being limited to a memory internal to the video encoder 200, unless so specifically described, or a memory external to the video encoder 200, unless so specifically described. Rather, a reference to the video data memory 230 should be understood as a reference memory that stores video data received by the video encoder 200 for encoding (e.g., video data of a current block to be encoded). Figure 1 The memory 106 may also provide temporary storage of the outputs of the various units of the video encoder 200.
[0186] illustrates Figure 6 the various units to assist in understanding the operations performed by the video encoder 200. These units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit is a circuit that provides a specific function and is pre-set in terms of the operations it can perform. A programmable circuit is a circuit that can be programmed to perform various tasks and provides flexible functionality in terms of the operations it can execute. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit may execute software instructions (e.g., to receive parameters or output parameters), but the type of operations performed by the fixed-function circuit is generally immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be integrated circuits.
[0187] The video encoder 200 may include an arithmetic logic unit (ALU), a basic function unit (EFU), digital circuits, analog circuits, and / or a programmable core formed by programmable circuits. In an example where software executed by programmable circuits is used to perform the operations of the video encoder 200, the memory 106( Figure 1 ) may store the instructions (e.g., object code) of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.
[0188] The video data memory 230 is configured to store received video data. The video encoder 200 may extract pictures of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 may be the original video data to be encoded.
[0189] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra prediction unit 226. The mode selection unit 202 may include other functional units to perform video prediction according to other prediction modes. As an example, the mode selection unit 202 may include a palette unit, a block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, and so on.
[0190] The mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and the resulting rate-distortion values for such combinations. The encoding parameters may include: the partitioning of CTUs into CUs, the prediction mode for the CUs, the transform type for the residual data of the CUs, the quantization parameter for the residual data of the CUs, and so on. The mode selection unit 202 may ultimately select the combination of encoding parameters with a rate-distortion value superior to other tested combinations.
[0191] The video encoder 200 may divide a picture retrieved from the video data memory 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 may divide the CTUs of the picture according to a tree structure (such as the above-mentioned MTT structure, QTBT structure, superblock structure, or quadtree structure). As described above, the video encoder 200 may form one or more CUs by dividing CTUs according to a tree structure. Such CUs are also commonly referred to as "video blocks" or "blocks".
[0192] Generally, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for the current block (e.g., the current CU, or the overlapping portion of the PU and TU in HEVC). For inter prediction of the current block, the motion estimation unit 222 may perform a motion search to identify one or more reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in the DPB 218) that closely match. Specifically, the motion estimation unit 222 may calculate, for example, values representing the similarity between a potential reference block and the current block according to the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), and so on. The motion estimation unit 222 typically performs these calculations using the per-sample differences between the current block and the considered reference block. The motion estimation unit 222 may identify the reference block having the lowest value resulting from these calculations, which indicates the reference block that most closely matches the current block.
[0193] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to the position of a current block in the current picture. Then, the motion estimation unit 222 may provide the motion vectors to the motion compensation unit 224. For example, for uni-directional inter prediction, the motion estimation unit 222 may provide a single motion vector, while for bi-directional inter prediction, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then use the motion vectors to generate a prediction block. For example, the motion compensation unit 224 may use the motion vectors to retrieve data of the reference block. As another example, if the motion vectors have fractional sample precision, the motion compensation unit 224 may interpolate values of the prediction block according to one or more interpolation filters. Additionally, for bi-directional inter prediction, the motion compensation unit 224 may retrieve data of two reference blocks identified by the respective motion vectors and combine the retrieved data, such as by per-sample averaging or weighted averaging.
[0194] When operating according to the AV1 video coding format, the motion estimation unit 222 and the motion compensation unit 224 may be configured to use translational motion compensation, affine motion compensation, overlapped block motion compensation (OBMC), and / or combined inter-intra prediction to encode coded blocks (e.g., both luminance coded blocks and chrominance coded blocks) of video data.
[0195] The motion estimation unit 222 and the motion compensation unit 224 may also be configured to perform one or more techniques of the present disclosure related to adaptive affine DMVR. For example, the motion estimation unit 222 and the motion compensation unit 224 may be configured to: receive a first video data block to be encoded using adaptive affine DMVR; determine to set a first MVD for a first reference picture list to zero; refine a CPMV associated with a second reference picture list to generate a refined CPMV; and use the refined CPMV to encode the first video data block.
[0196] As another example, for intra prediction or intra prediction coding, the intra prediction unit 226 may generate a prediction block according to samples adjacent to the current block. For example, for a directional mode, the intra prediction unit 226 may generally mathematically combine values of adjacent samples and fill the calculated values in the defined direction across the current block to produce a prediction block. As another example, for a DC mode, the intra prediction unit 226 may calculate an average value of adjacent samples of the current block and generate a prediction block to include the obtained average value for each sample of the prediction block.
[0197] When operating according to the AV1 video coding format, the intra prediction unit 226 may be configured to encode decoded blocks of video data (e.g., both luma decoded blocks and chroma decoded blocks) using directional intra prediction, non-directional intra prediction, recursive filter intra prediction, chroma from luma (CFL) prediction, intra block copy (IBC), and / or palette mode. The mode selection unit 202 may include other functional units to perform video prediction according to other prediction modes.
[0198] The mode selection unit 202 provides the predicted block to the residual generation unit 204. The residual generation unit 204 receives the original uncoded version of the current block from the video data memory 230 and the predicted block from the mode selection unit 202. The residual generation unit 204 calculates the per-sample difference between the current block and the predicted block. The resulting per-sample difference defines the residual block of the current block. In some examples, the residual generation unit 204 may also determine the differences between the sample values in the residual block to generate the residual block using residual differential pulse coding modulation (RDPCM). In some examples, one or more subtractor circuits performing binary subtraction may be used to form the residual generation unit 204.
[0199] In an example where the mode selection unit 202 divides a CU into PUs, each PU may be associated with a luma prediction unit and a corresponding chroma prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As described above, the size of a CU may refer to the size of the luma decoded block of the CU, and the size of a PU may refer to the size of the luma prediction unit of the PU. Assuming the size of a particular CU is 2Nx2N, the video encoder 200 may support PU sizes of 2Nx2N or NxN for intra prediction, and 2Nx2N, 2NxN, Nx2N, NxN, or similar symmetric PU sizes for inter prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitions for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter prediction.
[0200] In an example where the mode selection unit 202 does not further divide a CU into PUs, each CU may be associated with a luma decoded block and a corresponding chroma decoded block. As described above, the size of a CU may refer to the size of the luma decoded block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2Nx2N, 2NxN, or Nx2N.
[0201] For other video decoding techniques (such as intra block copy mode decoding, affine mode decoding, and linear model (LM) mode decoding, as some examples), the mode selection unit 202 generates a predicted block for the current block being encoded through corresponding units associated with the decoding technique. In some examples, such as palette mode decoding, the mode selection unit 202 may not generate a predicted block, but instead generate syntax elements indicating the manner in which the block is to be reconstructed based on the selected palette. In such a mode, the mode selection unit 202 may provide these syntax elements to the entropy encoding unit 220 for encoding.
[0202] As described above, the residual generation unit 204 receives the video data of the current block and the corresponding predicted block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the per-sample difference between the predicted block and the current block.
[0203] The transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, the transform processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transform processing unit 206 may perform multiple transforms on the residual block, e.g., a first transform and a second transform, such as a rotation transform. In some examples, the transform processing unit 206 does not apply a transform to the residual block.
[0204] When operating according to AV1, the transform processing unit 206 may apply one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, the transform processing unit 206 may apply a horizontal / vertical transform combination that may include a discrete cosine transform (DCT), an asymmetric discrete sine transform (ADST), a flipped ADST (e.g., ADST in reverse order), and an identity transform (IDTX). When using the identity transform, the transform is skipped in one of the vertical or horizontal directions. In some examples, the transform processing may be skipped.
[0205] Quantization unit 208 may quantize the transform coefficients in a transform coefficient block to produce a quantized transform coefficient block. Quantization unit 208 may quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) may adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and thus, the quantized transform coefficients may have lower precision than the original transform coefficients generated by transform processing unit 206.
[0206] Inverse quantization unit 210 and inverse transform processing unit 212 may apply inverse quantization and inverse transform to the quantized transform coefficient block, respectively, to reconstruct a residual block from the transform coefficient block. Reconstruction unit 214 may generate a reconstructed block corresponding to the current block (although there may be some degree of distortion) based on the reconstructed residual block and the prediction block generated by mode selection unit 202. For example, reconstruction unit 214 may add the samples of the reconstructed residual block to the corresponding samples from the prediction block generated by mode selection unit 202 to generate the reconstructed block.
[0207] Filter unit 216 may perform one or more filter operations on the reconstructed block. For example, filter unit 216 may perform a deblocking operation to reduce blocking artifacts along the edges of the CU. In some examples, the operation of filter unit 216 may be skipped.
[0208] When operating according to AV1, filter unit 216 may perform one or more filter operations on the reconstructed block. For example, filter unit 216 may perform a deblocking operation to reduce blocking artifacts along the edges of the CU. In other examples, filter unit 216 may apply a constrained directional enhancement filter (CDEF), which may be applied after deblocking and may include applying a non-separable non-linear low-pass directional filter based on the estimated edge direction. Filter unit 216 may also include a loop restoration filter applied after CDEF, and may include a separable symmetric normalized Wiener filter or a dual self-guiding filter.
[0209] Video encoder 200 stores the reconstructed blocks in DPB 218. For example, in an example where the operations of filter unit 216 are not performed, reconstruction unit 214 may store the reconstructed blocks into DPB 218. In an example where the operations of filter unit 216 are performed, filter unit 216 may store the filtered reconstructed blocks into DPB 218. Motion estimation unit 222 and motion compensation unit 224 may retrieve reference pictures from DPB 218, which are formed by reconstructed (and possibly filtered) blocks, for performing inter prediction on blocks of subsequently encoded pictures. Additionally, intra prediction unit 226 may use the reconstructed blocks of the current picture in DPB 218 to perform intra prediction on other blocks in the current picture.
[0210] Generally, entropy coding unit 220 may perform entropy coding on syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 may perform entropy coding on the quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 may perform entropy coding on prediction syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from mode selection unit 202. Entropy coding unit 220 may perform one or more entropy coding operations on the syntax elements (which are another example of video data) to generate entropy-coded data. For example, entropy coding unit 220 may perform context-adaptive variable length coding (CAVLC) operations, CABAC operations, variable-to-variable (V2V) length coding operations, syntax-based context-adaptive binary arithmetic coding (SBAC) operations, probability interval partitioning entropy (PIPE) coding operations, exponential Golomb coding operations, or another entropy coding operation on the data. In some examples, entropy coding unit 220 may operate in a bypass mode, in which the syntax elements are not entropy-coded.
[0211] Video encoder 200 may output a bitstream including the entropy-coded syntax elements required for the reconstructed blocks of a slice or picture. In particular, entropy coding unit 220 may output the bitstream.
[0212] According to AV1, entropy coding unit 220 may be configured as a symbol-to-symbol adaptive multi-symbol arithmetic coder. The syntax elements in AV1 include an alphabet of N elements, and the context (e.g., probability model) includes a set of N probabilities. Entropy coding unit 220 may store the probabilities as an n-bit (e.g., 15-bit) cumulative distribution function (CDF). Entropy coding unit 220 may perform recursive scaling using an update factor based on the alphabet size to update the context.
[0213] The above operations are described for blocks. Such descriptions should be understood as operations for a luminance decoding block and / or a chrominance decoding block. As described above, in some examples, the luminance decoding block and the chrominance decoding block are the luminance component and the chrominance component of a CU. In some examples, the luminance decoding block and the chrominance decoding block are the luminance component and the chrominance component of a PU.
[0214] In some examples, it is not necessary to repeat the operations performed for the luminance decoding block for the chrominance decoding block. As an example, it is not necessary to repeat the operation of identifying the motion vector (MV) and the reference picture for the luminance decoding block for identifying the MV and the reference picture for the chrominance block. Instead, the MV for the luminance decoding block can be scaled to determine the MV for the chrominance block, and the reference picture can be the same. As another example, for the luminance decoding block and the chrominance decoding block, the intra prediction process can be the same.
[0215] Video encoder 200 represents an example of a device configured to encode video data, the device including: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to: receive a first video data block to be encoded using adaptive affine DMVR; determine to set a first MVD for a first reference picture list to zero; refine a CPMV associated with a second reference picture list to generate a refined CPMV; and encode the first video data block using the refined CPMV.
[0216] Figure 7 is a block diagram showing an exemplary video decoder 300 that can execute the techniques of the present disclosure. Provided Figure 7 is for illustrative purposes only and is not a limitation on the techniques widely illustrated and described in the present disclosure. For purposes of explanation, the present disclosure describes the video decoder 300 in accordance with the techniques of VVC and HEVC. However, the techniques of the present disclosure can be performed by video decoding devices configured for other video coding standards.
[0217] In Figure 7In the example, video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a DPB 314. Any one or all of the CPB memory 320, the entropy decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 314 may be implemented in one or more processors or in processing circuitry. For example, the units of video decoder 300 may be implemented as one or more circuits or logic elements that are part of a hardware circuit or part of a processor, ASIC, or FPGA. Additionally, video decoder 300 may include additional or alternative processors or processing circuitry to perform these and other functions.
[0218] The prediction processing unit 304 includes a motion compensation unit 316 and an intra prediction unit 318. The prediction processing unit 304 may include additional units to perform prediction according to other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, a block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, and so on. In other examples, video decoder 300 may include more, fewer, or different functional components.
[0219] As described above, when operating according to AV1, the motion compensation unit 316 may be configured to decode coded blocks of video data (e.g., both luma coded blocks and chroma coded blocks) using translational motion compensation, affine motion compensation, OBMC, and / or composite inter-intra prediction. As described above, the intra prediction unit 318 may be configured to decode coded blocks of video data (e.g., both luma coded blocks and chroma coded blocks) using directional intra prediction, non-directional intra prediction, recursive filter intra prediction, CFL, IBC, and / or palette mode.
[0220] The motion compensation unit 316 may also be configured to perform one or more techniques of the present disclosure related to adaptive affine DMVR. For example, the motion compensation unit 316 may be configured to: receive a first video data block to be decoded using adaptive affine DMVR; determine to set a first MVD for a first reference picture list to zero; refine a CPMV associated with a second reference picture list to generate a refined CPMV; and use the refined CPMV to decode the first video data block.
[0221] The CPB memory 320 may store video data to be decoded by components of the video decoder 300, such as an encoded video bitstream. For example, the video data stored in the CPB memory 320 may be obtained from a computer-readable medium 110( Figure 1 ). The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Moreover, the CPB memory 320 may store video data other than the syntax elements of decoded pictures, such as temporary data representing the output of the respective units from the video decoder 300. The DPB 314 generally stores decoded pictures, and the video decoder 300 may output decoded pictures and / or use the decoded pictures as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 may be formed of any of a variety of storage devices, such as DRAM, including SDRAM, MRAM, RRAM, or other types of storage devices. The CPB memory 320 and the DPB 314 may be provided by the same storage device or separate storage devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300 or off-chip relative to those components.
[0222] Additionally or alternatively, in some examples, the video decoder 300 may retrieve decoded video data from a memory 120( Figure 1 ). That is, the memory 120 may store data together with the CPB memory 320 as described above. Similarly, when some or all of the functions of the video decoder 300 are implemented in software to be executed by the processing circuitry of the video decoder 300, the memory 120 may store instructions to be executed by the video decoder 300.
[0223] Shown in Figure 7 are the respective units shown to assist in understanding the operations performed by the video decoder 300. These units may be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Similar to Figure 6 , fixed-function circuitry refers to circuitry that provides a specific function and is pre-set in terms of the operations it can perform. Programmable circuitry refers to circuitry that can be programmed to perform various tasks and provides flexible functionality in terms of the operations it can perform. For example, programmable circuitry may execute software or firmware that causes the programmable circuitry to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuitry may execute software instructions (e.g., to receive parameters or output parameters), but the type of operations performed by the fixed-function circuitry is generally immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be an integrated circuit.
[0224] The video decoder 300 may include an ALU, an EFU, digital circuits, analog circuits, and / or a programmable core formed by programmable circuitry. In an example where the operations of the video decoder 300 are performed by software executed on the programmable circuitry, on-chip or off-chip memory may store the instructions (e.g., object code) of the software received and executed by the video decoder 300.
[0225] The entropy decoding unit 302 may receive the encoded video data from the CPB and perform entropy decoding on the video data to reproduce syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 may generate the decoded video data based on the syntax elements extracted from the bitstream.
[0226] Generally, the video decoder 300 reconstructs pictures on a block-by-block basis. The video decoder 300 may perform the reconstruction operation separately on each block (where the block currently being reconstructed (i.e., decoded) may be referred to as the "current block").
[0227] The entropy decoding unit 302 may perform entropy decoding on: the syntax elements that define the quantized transform coefficients in the quantized transform coefficient block, and transform information such as the quantization parameter (QP) and / or the transform mode indication. The inverse quantization unit 306 may use the QP associated with the quantized transform coefficient block to determine the quantization degree, and similarly, determine the inverse quantization degree applied by the inverse quantization unit 306. The inverse quantization unit 306 may, for example, perform a bitwise left shift operation to inverse-quantize the quantized transform coefficients. The inverse quantization unit 306 may thereby form a transform coefficient block including the transform coefficients.
[0228] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the transform coefficient block.
[0229] In addition, the prediction processing unit 304 generates a prediction block according to the prediction information syntax elements entropy decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is inter-predicted, the motion compensation unit 316 may generate the prediction block. In this case, the prediction information syntax elements may indicate the reference picture in the DPB 314 from which the reference block is retrieved, and the motion vector that identifies the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 may generally operate in a manner similar to that for the motion compensation unit 224 ( Figure 6)perform the inter-frame prediction process in a substantially similar manner as described.
[0230] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, the intra-prediction unit 318 may generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, the intra-prediction unit 318 generally may perform the intra-prediction process in a manner substantially similar to the manner described with respect to the intra-prediction unit 226( Figure 6 )perform the intra-prediction process in a manner substantially similar to the manner described. The intra-prediction unit 318 may retrieve data of adjacent samples of the current block from the DPB 314.
[0231] The reconstruction unit 310 may use the prediction block and the residual block to reconstruct the current block. For example, the reconstruction unit 310 may add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.
[0232] The filter unit 312 may perform one or more filter operations on the reconstructed block. For example, the filter unit 312 may perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. The operations of the filter unit 312 are not necessarily performed in all examples.
[0233] The video decoder 300 may store the reconstructed block in the DPB 314. For example, in an example where the operations of the filter unit 312 are not performed, the reconstruction unit 310 may store the reconstructed block into the DPB 314. In an example where the operations of the filter unit 312 are performed, the filter unit 312 may store the filtered reconstructed block into the DPB 314. As described above, the DPB 314 may provide reference information to the prediction processing unit 304, such as samples of the current picture for intra-prediction and previously decoded pictures for subsequent motion compensation. In addition, the video decoder 300 may output the decoded picture (e.g., decoded video) from the DPB 314 for subsequent presentation on a display device (e.g., Figure 1 display device 118).
[0234] In this way, the video decoder 300 represents an example of a device configured to decode video data, the device including: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to: receive a first video data block to be decoded using adaptive affine DMVR; determine to set the first MVD for the first reference picture list to zero; refine the CPMV associated with the second reference picture list to generate a refined CPMV; and decode the first video data block using the refined CPMV.
[0235] Figure 8is a flowchart illustrating an example method of encoding a current block according to techniques of the present disclosure. The current block may be or include a current CU. Although described with respect to video encoder 200 ( Figure 1 and 6 ), it should be understood that other devices may be configured to perform methods similar to the method of Figure 8 .
[0236] In this example, video encoder 200 initially predicts the current block (350). For example, video encoder 200 may form a predicted block of the current block. Video encoder 200 may then compute a residual block (352) of the current block. To compute the residual block, video encoder 200 may compute the difference between the original uncoded block and the predicted block of the current block. Video encoder 200 may then transform the residual block and quantize the transform coefficients of the residual block (354). Next, video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or after the scan, video encoder 200 may entropy code the transform coefficients (358). For example, video encoder 200 may use CAVLC or CABAC to code the transform coefficients. Video encoder 200 may then output the entropy coded data of the block (360).
[0237] Figure 9 is a flowchart illustrating an example method of decoding a current block according to techniques of the present disclosure. The current block may be or include a current CU. Although described with respect to video decoder 300 ( Figure 1 and 7 ), it should be understood that other devices may be configured to perform methods similar to the method of Figure 9 .
[0238] Video decoder 300 may receive entropy coded data for the current block, such as entropy coded prediction information and entropy coded data of transform coefficients for a residual block corresponding to the current block (370). Video decoder 300 may entropy decode the entropy coded data to determine prediction information for the current block and reproduce the transform coefficients of the residual block (372). Video decoder 300 may predict the current block (374), for example, using an intra or inter prediction mode as indicated by the prediction information of the current block, to compute a predicted block of the current block. Video decoder 300 may then inverse scan the reproduced transform coefficients (376) to create a block of quantized transform coefficients. Video decoder 300 may then inverse quantize the transform coefficients and apply an inverse transform to the transform coefficients to produce a residual block (378). Video decoder 300 may finally decode the current block by combining the predicted block and the residual block (380).
[0239] Figure 10FIG. is a flowchart illustrating another example method for encoding a current block according to the techniques of the present disclosure. Figure 10 The techniques may be performed by one or more structural components of the video encoder 200, including the motion estimation unit 222 and the motion compensation unit 224.
[0240] In one example, the video encoder 200 may receive a first video data block (1000) to be encoded using adaptive affine decoder-side motion vector refinement (DMVR). The video encoder 200 may be configured to encode a first syntax element indicating that the first video data block is to be decoded using adaptive affine DMVR.
[0241] The video encoder 200 may determine to set a first motion vector difference (MVD) for a first reference picture list to zero (1010). In one example, the first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1. In another example, the first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
[0242] In one example, the video encoder 200 may encode a second syntax element indicating that the first MVD for the first reference picture list is set to zero. The second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1. In another example, to determine to set the first MVD for the first reference picture list to zero, the video encoder 200 may determine to set the first MVD for the first reference picture list to zero based on a template matching cost or a bilateral matching cost.
[0243] The video encoder 200 may also be configured to refine a control point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV (1020). To refine the CPMV associated with the second reference picture list to generate a refined CPMV, the video encoder 200 may determine a bilateral matching (BM) cost within a search range of motion vectors for each sub-block in the first video data block, accumulate the BM costs for multiple sub-blocks of the first video data block to generate an accumulated BM cost, determine a second MVD of the CPMV for the second reference picture list based on the accumulated BM cost, and determine the refined CPMV based on the second MVD. In some examples, the video encoder 200 may determine one or more of a search pattern, a search range, or a cost metric for refining the CPMV.
[0244] The video encoder 200 can then encode (1030) the first video data block using the refined CPMV. The video encoder 200 can construct an adaptive affine merge candidate list for the first video data block, where the adaptive affine merge candidate list is different from the affine merge candidate list for the conventional affine DMVR mode. In one example, the adaptive affine merge candidate list includes only affine merge candidates.
[0245] Figure 11 is a flowchart illustrating another example method for decoding a current block according to the techniques of the present disclosure. Figure 11 The techniques of can be performed by one or more structural components of the video decoder 300, including the motion compensation unit 316.
[0246] In one example, the video decoder 300 can receive a first video data block to be decoded using adaptive affine decoder-side motion vector refinement (DMVR) (1100). The video decoder 300 can be configured to decode a first syntax element indicating that the first video data block is to be decoded using adaptive affine DMVR.
[0247] The video decoder 300 can determine to set a first motion vector difference (MVD) for the first reference picture list to zero (1110). In one example, the first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1. In another example, the first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
[0248] In one example, to determine to set the first MVD for the first reference picture list to zero, the video decoder 300 can decode a second syntax element indicating that the first MVD for the first reference picture list is to be set to zero. The second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1. In another example, to determine to set the first MVD for the first reference picture list to zero, the video decoder 300 can determine to set the first MVD for the first reference picture list to zero based on a template matching cost or a bilateral matching cost.
[0249] The video decoder 300 may also be configured to refine a control point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV (1120). To refine the CPMV associated with the second reference picture list to generate a refined CPMV, the video decoder 300 may determine a bilateral matching (BM) cost within a search range of motion vectors for each sub-block in a first video data block, accumulate the BM costs for the plurality of sub-blocks of the first video data block to generate an accumulated BM cost, determine a second MVD for the CPMV for the second reference picture list based on the accumulated BM cost, and determine the refined CPMV based on the second MVD. In some examples, the video decoder 300 may determine one or more of a search pattern, a search range, or a cost metric for refining the CPMV.
[0250] The video decoder 300 may then decode the first video data block using the refined CPMV (1130). The video decoder 300 may construct an adaptive affine merge candidate list for the first video data block, where the adaptive affine merge candidate list is different from the affine merge candidate list for a conventional affine DMVR mode. In one example, the adaptive affine merge candidate list includes only affine merge candidates.
[0251] The following numbered clauses illustrate one or more aspects of the devices and techniques described in this disclosure.
[0252] Aspect 1A - A method of decoding video data, the method comprising: receiving a first video data block to be decoded using an adaptive affine decoder-side motion vector refinement (DMVR); determining to set a first motion vector difference (MVD) for a first reference picture list to zero; refining a control point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV; and decoding the first video data block using the refined CPMV.
[0253] Aspect 2A - The method according to Aspect 1A, wherein the first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1.
[0254] Aspect 3A - The method according to Aspect 1A, wherein the first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
[0255] Aspect 4A - The method according to any one of Aspects 1A - 3A, wherein using the second MVD to refine the CPMV associated with the second reference picture list to generate a refined CPMV includes: for each sub - block in the first video data block, determining the bilateral matching (BM) cost within the search range of the motion vector; accumulating the BM cost; based on the accumulated BM cost, determining the second MVD of the CPMV for the second reference picture list; and determining the refined CMPV according to the second MVD.
[0256] Aspect 5A - The method according to any one of Aspects 1A - 4A, wherein determining to set the first MVD for the first reference picture list to zero includes: decoding a syntax element indicating to set the first MVD for the first reference picture list to zero.
[0257] Aspect 6A - The method according to any one of Aspects 1A - 4A, wherein determining to set the first MVD for the first reference picture list to zero includes: determining to set the first MVD for the first reference picture list to zero based on the template matching cost or the bilateral matching cost.
[0258] Aspect 7A - The method according to any one of Aspects 1A - 6A, further includes: constructing an affine merge candidate list based on the conditions of the affine DMVR.
[0259] Aspect 8A - The method according to any one of Aspects 1A - 7A, further includes: determining one or more of the search pattern, search range, or cost metric of the adaptive affine DMVR.
[0260] Aspect 9A - The method according to any one of Aspects 1A - 8A, wherein decoding includes decoding.
[0261] Aspect 10A - The method according to any one of Aspects 1A - 8A, wherein decoding includes encoding.
[0262] Aspect 11A - A device for decoding video data, the device includes one or more units for performing the method according to any one of Aspects 1A - 10A.
[0263] Aspect 12A - The device according to Aspect 11A, wherein the one or more units include one or more processors implemented in a circuit.
[0264] Aspect 13A - The device according to any one of Aspects 11A and 12A, further includes: a memory for storing the video data.
[0265] Aspect 14A - The device according to any one of Aspects 11A–13A, further includes: a display configured to display the decoded video data.
[0266] Aspect 15A - The apparatus according to any one of aspects 11A–14A, wherein the apparatus comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.
[0267] Aspect 16A - The apparatus according to any one of aspects 11A–15A, wherein the apparatus comprises a video decoder.
[0268] Aspect 17A - The apparatus according to any one of aspects 11A–16A, wherein the apparatus comprises a video encoder.
[0269] Aspect 18A - A computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to perform the method according to any one of aspects 1A-10A.
[0270] Aspect 1B - A method of decoding video data, the method comprising: receiving a first video data block to be decoded using adaptive affine decoder-side motion vector refinement (DMVR); determining to set a first motion vector difference (MVD) for a first reference picture list to zero; refining a control point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV; and decoding the first video data block using the refined CPMV.
[0271] Aspect 2B - The method according to aspect 1B, wherein refining the CPMV associated with the second reference picture list to generate a refined CPMV comprises: for each sub-block in the first video data block, determining a bilateral matching (BM) cost within a search range of a motion vector; accumulating the BM costs for a plurality of sub-blocks of the first video data block to generate an accumulated BM cost; determining a second MVD for the CPMV for the second reference picture list based on the accumulated BM cost; and determining the refined CPMV based on the second MVD.
[0272] Aspect 3B - The method according to aspect 2B, wherein the first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1.
[0273] Aspect 4B - The method according to aspect 2B, wherein the first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
[0274] Aspect 5B - The method according to any one of Aspects 1B - 4B further includes: decoding a first syntax element indicating that adaptive affine DMVR is to be used to decode a first video data block.
[0275] Aspect 6B - The method according to Aspect 5B, wherein determining to set a first motion vector difference (MVD) for a first reference picture list to zero includes: decoding a second syntax element indicating that the first MVD for the first reference picture list is to be set to zero.
[0276] Aspect 7B - The method according to Aspect 6B, wherein the second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1.
[0277] Aspect 8B - The method according to any one of Aspects 1B - 7B further includes: constructing an adaptive affine merge candidate list for the first video data block, wherein the adaptive affine merge candidate list is different from the affine merge candidate list for a conventional affine DMVR mode.
[0278] Aspect 9B - The method according to Aspect 8B, wherein the adaptive affine merge candidate list includes only affine merge candidates.
[0279] Aspect 10B - The method according to any one of Aspects 1B - 9B further includes: determining one or more of a search pattern, a search range, or a cost metric for refining a control point motion vector (CPMV).
[0280] Aspect 11B - The method according to Aspect 1B, wherein determining to set a first MVD for a first reference picture list to zero includes: determining to set the first MVD for the first reference picture list to zero based on a template matching cost or a bilateral matching cost.
[0281] Aspect 12B - The method according to any one of Aspects 1B - 11B further includes: displaying a picture including the first video data block.
[0282] Aspect 13B - An apparatus configured to decode video data, the apparatus includes: a memory; and one or more processors in communication with the memory, the one or more processors being configured to: receive a first video data block to be decoded using an adaptive affine decoder side motion vector refinement (DMVR); determine to set a first motion vector difference (MVD) for a first reference picture list to zero; refine a control point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV; and decode the first video data block using the refined CPMV.
[0283] Aspect 14B - The apparatus according to aspect 13B, wherein, in order to refine the CPMV associated with the second reference picture list to generate a refined CPMV, one or more processors are further configured to: for each sub - block in the first video data block, determine the bilateral matching (BM) cost within the search range of the motion vector; accumulate the BM costs for the multiple sub - blocks of the first video data block to generate an accumulated BM cost; determine a second MVD of the CPMV for the second reference picture list based on the accumulated BM cost; and determine the refined CPMV based on the second MVD.
[0284] Aspect 15B - The apparatus according to aspect 14B, wherein the first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1.
[0285] Aspect 16B - The apparatus according to aspect 14B, wherein the first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
[0286] Aspect 17B - The apparatus according to any one of aspects 13B - 16B, wherein the one or more processors are further configured to: decode a first syntax element indicating that adaptive affine DMVR is to be used to decode the first video data block.
[0287] Aspect 18B - The apparatus according to aspect 17B, wherein, in order to determine to set the first MVD for the first reference picture list to zero, one or more processors are further configured to: decode a second syntax element indicating that the first MVD for the first reference picture list is to be set to zero.
[0288] Aspect 19B - The apparatus according to aspect 18B, wherein the second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1.
[0289] Aspect 20B - The apparatus according to any one of aspects 13B - 19B, wherein one or more processors are further configured to: construct an adaptive affine merge candidate list for the first video data block, wherein the adaptive affine merge candidate list is different from the affine merge candidate list for the conventional affine DMVR mode.
[0290] Aspect 21B - The apparatus according to aspect 20B, wherein the adaptive affine merge candidate list includes only affine merge candidates.
[0291] Aspect 22B - The apparatus according to any one of Aspects 13B - 21B, wherein one or more processors are further configured to: determine one or more of a search pattern, a search range, or a cost metric for refining the CPMV.
[0292] Aspect 23B - The apparatus according to Aspect 13B, wherein, in order to determine to set a first motion vector difference (MVD) for a first reference picture list to zero, one or more processors are further configured to: determine to set the first MVD for the first reference picture list to zero based on a template matching cost or a bilateral matching cost.
[0293] Aspect 24B - The apparatus according to any one of Aspects 13B - 23B, further comprising: a display configured to display a picture including a first video data block.
[0294] Aspect 25B - A method for encoding video data, the method comprising: receiving a first video data block to be encoded using adaptive affine decoder - side motion vector refinement (DMVR); determining to set a first motion vector difference (MVD) for a first reference picture list to zero; refining a control - point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV; and encoding the first video data block using the refined CPMV.
[0295] Aspect 26B - The method according to Aspect 25B, wherein refining the CPMV associated with the second reference picture list to generate a refined CPMV comprises: for each sub - block in the first video data block, determining a bilateral matching (BM) cost within a search range of a motion vector; accumulating the BM costs for a plurality of sub - blocks of the first video data block to generate an accumulated BM cost; determining a second MVD of the CPMV for the second reference picture list based on the accumulated BM cost; and determining the refined CPMV based on the second MVD.
[0296] Aspect 27B - The method according to Aspect 26B, wherein the first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1.
[0297] Aspect 28B - The method according to Aspect 26B, wherein the first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
[0298] Aspect 29B - The method according to any one of Aspects 25B - 28B, further comprising: encoding a first syntax element indicating to decode the first video data block using adaptive affine DMVR.
[0299] Aspect 30B - The method according to aspect 29B further includes: encoding a second syntax element indicating that a first MVD for a first reference picture list is set to zero.
[0300] Aspect 31B - The method according to aspect 30B, wherein the second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1.
[0301] Aspect 32B - The method according to any one of aspects 25B - 31B further includes: constructing an adaptive affine merge candidate list for a first video data block, wherein the adaptive affine merge candidate list is different from the affine merge candidate list for a conventional affine DMVR mode.
[0302] Aspect 33B - The method according to aspect 32B, wherein the adaptive affine merge candidate list includes only affine merge candidates.
[0303] Aspect 34B - The method according to any one of aspects 25B - 33B further includes: determining one or more of a search pattern, a search range, or a cost metric for refining the CPMV.
[0304] Aspect 35B - The method according to aspect 25B, wherein determining to set a first MVD for a first reference picture list to zero includes: determining to set the first MVD for the first reference picture list to zero based on a template matching cost or a bilateral matching cost.
[0305] Aspect 36B - The method according to any one of aspects 25B - 35B further includes: capturing a picture including the first video data block.
[0306] Aspect 37B - An apparatus configured to encode video data, the apparatus including: a memory; and one or more processors in communication with the memory, the one or more processors being configured to: receive a first video data block to be encoded using an adaptive affine decoder - side motion vector refinement (DMVR); determine to set a first motion vector difference (MVD) for a first reference picture list to zero; refine a control - point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV; and encode the first video data block using the refined CPMV.
[0307] Aspect 38B - The apparatus according to aspect 37B, wherein, to refine the CPMV associated with the second reference picture list to generate a refined CPMV, one or more processors are further configured to: for each sub - block in the first video data block, determine the bilateral matching (BM) cost within the search range of the motion vector; accumulate the BM costs for the multiple sub - blocks of the first video data block to generate an accumulated BM cost; determine a second MVD of the CPMV for the second reference picture list based on the accumulated BM cost; and determine the refined CPMV based on the second MVD.
[0308] Aspect 39B - The apparatus according to aspect 38B, wherein the first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1.
[0309] Aspect 40B - The apparatus according to aspect 38B, wherein the first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
[0310] Aspect 41B - The apparatus according to any one of aspects 37B - 40B, wherein one or more processors are further configured to: encode a first syntax element indicating that adaptive affine DMVR is to be used to decode the first video data block.
[0311] Aspect 42B - The apparatus according to aspect 41B, wherein the one or more processors are further configured to: encode a second syntax element indicating that the first MVD for the first reference picture list is set to zero.
[0312] Aspect 43B - The apparatus according to aspect 42B, wherein the second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1.
[0313] Aspect 44B - The apparatus according to any one of aspects 37B - 43B, wherein one or more processors are further configured to: construct an adaptive affine merge candidate list for the first video data block, wherein the adaptive affine merge candidate list is different from the affine merge candidate list for the conventional affine DMVR mode.
[0314] Aspect 45B - The apparatus according to aspect 44B, wherein the adaptive affine merge candidate list includes only affine merge candidates.
[0315] Aspect 46B - The apparatus according to any one of Aspects 37B - 45B, wherein the one or more processors are further configured to: determine one or more of a search pattern, a search range, or a cost metric for refining the CPMV.
[0316] Aspect 47B - The apparatus according to Aspect 37B, wherein, in order to determine to set the first MVD for the first reference picture list to zero, the one or more processors are further configured to: determine to set the first MVD for the first reference picture list to zero based on a template matching cost or a bilateral matching cost.
[0317] Aspect 48B - The apparatus according to any one of Aspects 37B - 47B, further comprising: a camera configured to capture a picture including a first video data block.
[0318] It should be recognized that, according to an example, certain operations or events of any of the techniques described herein may be performed in a different order, may be added, combined, or completely omitted (e.g., not all described operations or events are necessary to implement the technique). Additionally, in certain examples, the operations or events may be performed, for example, concurrently rather than sequentially via multi - threading, interrupt processing, or multiple processors.
[0319] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored or transmitted as one or more instructions or code on a computer - readable medium and executed by a hardware - based processing unit. The computer - readable medium may include: a computer storage medium, which corresponds to a tangible medium such as a data storage medium; or a communication medium, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this way, the computer - readable medium generally may correspond to: (1) a non - transitory tangible computer - readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures to implement the techniques described in this disclosure. A computer program product may include a computer - readable medium.
[0320] For example, but not limited to, such a computer-readable storage medium may include one or more of the following: RAM, ROM, EEPROM, CD-ROM, or other optical disk storage devices, magnetic disk storage devices, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are also included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather are directed to non-transitory, tangible storage media. As used herein, disk and optical disks include compact disk (CD), laser disk, optical disk, digital versatile disk (DVD), floppy disk, and Blu-ray disk, where disks typically reproduce data magnetically, while optical disks typically reproduce data optically using a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0321] The instructions may be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuitry. Accordingly, as used herein, the terms "processor" and "processing circuitry" may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Similarly, the techniques may be fully implemented in one or more circuits or logic elements.
[0322] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a group of ICs (e.g., a chipset). In the present disclosure, various components, modules, or units are described to emphasize the functional aspects of the devices configured to perform the disclosed techniques, but do not necessarily need to be implemented by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit, or provided by a collection of interoperating hardware units, including one or more processors as described above in combination with suitable software and / or firmware.
[0323] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A method for decoding video data, the method comprising: Receiving a first video data block to be decoded using adaptive affine decoder side motion vector refinement (DMVR); Determining to set a first motion vector difference (MVD) for a first reference picture list to zero; Refining a control point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV; And Decoding the first video data block using the refined CPMV.
2. The method according to claim 1, wherein Refining the CPMV associated with the second reference picture list to generate the refined CPMV includes: For each sub-block in the first video data block, determining a bilateral matching (BM) cost within a search range of a motion vector; Accumulating the BM costs for multiple sub-blocks of the first video data block to generate an accumulated BM cost; Determining a second MVD for the CPMV for the second reference picture list based on the accumulated BM cost; and Determining the refined CPMV based on the second MVD.
3. The method according to claim 2, wherein, The first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1.
4. The method according to claim 2, wherein The first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
5. The method according to claim 1, further comprising: Decoding a first syntax element indicating that the first video data block is to be decoded using adaptive affine DMVR.
6. The method according to claim 5, wherein Determining to set the first MVD for the first reference picture list to zero includes: Decoding a second syntax element indicating that the first MVD for the first reference picture list is to be set to zero.
7. The method according to claim 6, wherein The second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1.
8. The method according to claim 1, further comprising: Constructing an adaptive affine merge candidate list for the first video data block, wherein the adaptive affine merge candidate list is different from an affine merge candidate list for a conventional affine DMVR mode.
9. The method according to claim 8, wherein, The adaptive affine merge candidate list includes only affine merge candidates.
10. The method according to claim 1, further comprising: Determining one or more of a search pattern, a search range, or a cost metric for refining the CPMV.
11. The method according to claim 1, wherein, Determining to set the first MVD for the first reference picture list to zero includes: Based on a template matching cost or a bilateral matching cost, determining to set the first MVD for the first reference image list to zero.
12. The method according to claim 1, further comprising: Displaying a picture including the first video data block.
13. An apparatus configured to decode video data, the apparatus comprising: A memory; And One or more processors in communication with the memory, the one or more processors being configured to: Receive a first video data block to be decoded using adaptive affine decoder - side motion vector refinement (DMVR); Determine to set a first motion vector difference (MVD) for a first reference picture list to zero; Refine a control - point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV; And Decode the first video data block using the refined CPMV.
14. The apparatus according to claim 13, wherein, To refine the CPMV associated with the second reference picture list to generate the refined CPMV, the one or more processors are further configured to: For each sub - block in the first video data block, determine a bilateral matching (BM) cost within a search range of a motion vector; Accumulate the BM costs for multiple sub - blocks of the first video data block to generate an accumulated BM cost; Determine a second MVD of the CPMV for the second reference picture list based on the accumulated BM cost; And Determine the refined CPMV based on the second MVD.
15. The apparatus according to claim 14, wherein, The first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1.
16. The device according to claim 14, wherein, The first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
17. The device according to claim 13, wherein, The one or more processors are further configured to: Decode a first syntax element indicating to use adaptive affine DMVR to decode the first video data block.
18. The device according to claim 17, wherein, To determine to set the first MVD for the first reference picture list to zero, the one or more processors are further configured to: Decode a second syntax element indicating to set the first MVD for the first reference picture list to zero.
19. The apparatus according to claim 18, wherein, The second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1.
20. The apparatus according to claim 13, wherein The one or more processors are further configured to: Construct an adaptive affine merge candidate list for the first video data block, wherein the adaptive affine merge candidate list is different from an affine merge candidate list for a conventional affine DMVR mode.
21. The apparatus according to claim 20, wherein, The adaptive affine merge candidate list includes only affine merge candidates.
22. The device according to claim 13, wherein The one or more processors are further configured to: Determine one or more of a search pattern, a search range, or a cost metric for refining the CPMV.
23. The apparatus according to claim 13, wherein To determine to set the first MVD for the first reference picture list to zero, the one or more processors are further configured to: Based on a template - matching cost or a bilateral - matching cost, determine to set the first MVD for the first reference image list to zero.
24. The apparatus according to claim 13, further comprising: A display configured to display a picture including the first video data block.
25. A method for encoding video data, the method comprising: Receive a first video data block to be encoded using adaptive affine decoder - side motion vector refinement (DMVR); Determine to set the first motion vector difference (MVD) for the first reference picture list to zero; Refine the control point motion vector (CPMV) associated with the second reference picture list to generate a refined CPMV; And Encode the first video data block using the refined CPMV.
26. The method according to claim 25, wherein, Refining the CPMV associated with the second reference picture list to generate the refined CPMV includes: For each sub-block in the first video data block, determine the bilateral matching (BM) cost within the search range of the motion vector; Accumulate the BM costs for multiple sub-blocks of the first video data block to generate an accumulated BM cost; Determine a second MVD of the CPMV for the second reference picture list based on the accumulated BM cost; and Determine the refined CPMV based on the second MVD.
27. The method according to claim 26, wherein, The first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1.
28. The method according to claim 26, wherein, The first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
29. The method according to claim 25, further comprising: Encode a first syntax element indicating that the first video data block is to be decoded using adaptive affine DMVR.
30. The method according to claim 29, further comprising: Encode a second syntax element indicating that the first MVD for the first reference picture list is set to zero.
31. The method according to claim 30, wherein, The second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1.
32. The method according to claim 25, further comprising: Construct an adaptive affine merge candidate list for the first video data block, wherein the adaptive affine merge candidate list is different from the affine merge candidate list for the conventional affine DMVR mode.
33. The method according to claim 32, wherein, The adaptive affine merge candidate list includes only affine merge candidates.
34. The method according to claim 25, further comprising: Determine one or more of a search pattern, a search range, or a cost metric for refining the CPMV.
35. The method according to claim 25, wherein, Determining to set the first MVD for the first reference picture list to zero includes: Based on the template matching cost or the bilateral matching cost, determine to set the first MVD for the first reference image list to zero.
36. The method according to claim 25, further comprising: Capture a picture including the first video data block.
37. An apparatus configured to encode video data, the apparatus comprising: A memory; And One or more processors in communication with the memory, the one or more processors being configured to: Receive a first video data block to be encoded using adaptive affine decoder-side motion vector refinement (DMVR); Determine to set the first motion vector difference (MVD) for the first reference picture list to zero; Refine a control point motion vector (CPMV) associated with a second reference picture list to generate a refined CPMV; and Encode the first video data block using the refined CPMV.
38. The apparatus according to claim 37, wherein, To refine the CPMV associated with the second reference picture list to generate the refined CPMV, the one or more processors are further configured to: For each sub-block in the first video data block, determine a bilateral matching (BM) cost within a search range of a motion vector; Accumulate the BM costs for multiple sub-blocks of the first video data block to generate an accumulated BM cost; Determine a second MVD of the CPMV for the second reference picture list based on the accumulated BM cost; and Determine the refined CPMV based on the second MVD.
39. The apparatus according to claim 38, wherein, The first MVD is MVD0, the first reference picture list is reference picture list 0, the second MVD is MVD1, and the second reference picture list is reference picture list 1.
40. The apparatus according to claim 38, wherein, The first MVD is MVD1, the first reference picture list is reference picture list 1, the second MVD is MVD0, and the second reference picture list is reference picture list 0.
41. The device according to claim 37, wherein, The one or more processors are further configured to: Encode a first syntax element indicating that the first video data block is to be decoded using adaptive affine DMVR.
42. The apparatus according to claim 41, wherein, The one or more processors are further configured to: Encode a second syntax element indicating that the first MVD for the first reference picture list is set to zero.
43. The apparatus according to claim 42, wherein, The second syntax element indicates whether the first reference picture list is reference picture list 0 or reference picture list 1.
44. The apparatus according to claim 37, wherein The one or more processors are further configured to: Construct an adaptive affine merge candidate list for the first video data block, wherein the adaptive affine merge candidate list is different from an affine merge candidate list for a conventional affine DMVR mode.
45. The apparatus according to claim 44, wherein, The adaptive affine merge candidate list includes only affine merge candidates.
46. The apparatus according to claim 37, wherein, The one or more processors are further configured to: Determine one or more of a search pattern, a search range, or a cost metric for refining the CPMV.
47. The apparatus according to claim 37, wherein To determine that the first MVD for the first reference picture list is set to zero, the one or more processors are further configured to: Determine that the first MVD for the first reference picture list is set to zero based on a template matching cost or a bilateral matching cost.
48. The apparatus according to claim 37, further comprising: A camera configured to capture a picture including the first video data block.