Gradient-Based Prediction Refinement for Video Coding

By rounding the displacement accuracy of the motion vector, it is consistent in different inter-frame prediction modes, and the problem of inconsistent displacement accuracy levels in the prior art is solved, and the simplification of hardware design and improvement of efficiency is achieved.

CN113950839BActive Publication Date: 2025-06-13QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080034608.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-14
Filing Date
2020-05-15
Publication Date
2025-06-13
Estimated Expiration
2040-05-15

AI Technical Summary

Technical Problem

The existing video encoding and decoding technologies have the problem of inconsistent displacement accuracy levels in inter-frame prediction, which leads to the need for different logic circuits to support different accuracy levels, increasing hardware complexity and power consumption.

Method used

By rounding the horizontal and vertical displacements of the motion vectors, they have the same level of accuracy in different inter prediction modes, thus using the same logic circuit for gradient-based prediction refinement.

Benefits of technology

The hardware design of video encoder and decoder is simplified, reducing the number and power consumption of logic circuits, while improving the overall efficiency of operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113950839B_ABST
    Figure CN113950839B_ABST
Patent Text Reader

Abstract

The present disclosure describes gradient-based prediction refinement. A video decoder (e.g., a video encoder or a video decoder) determines one or more prediction blocks for performing inter prediction on a current block (e.g., based on one or more motion vectors for the current block). In gradient-based prediction refinement, the video decoder modifies one or more samples of the prediction block based on various factors such as displacement in the horizontal direction, horizontal gradient, displacement in the vertical direction, and vertical gradient. The present disclosure provides gradient-based prediction refinement, in which for different prediction modes (e.g., including an affine mode and a bi-directional optical flow (BDOF) mode), the precision level of the displacement (e.g., at least one of a horizontal displacement or a vertical displacement) is uniform (e.g., the same).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of U.S. application Ser. No. 16 / 874,057, filed on May 14, 2020, which claims the benefit of U.S. Provisional Application No. 62 / 849,352, filed on May 17, 2019, the entire contents of each of which are hereby incorporated by reference. Technical Field

[0002] The present disclosure relates to video encoding and video decoding. Background Art

[0003] Digital video capabilities can be incorporated into a variety of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular phones or satellite wireless telephones, so-called "smartphones", video conferencing devices, video streaming devices, and the like. Digital video devices implement video decoding techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), and extensions of such standards. Video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information by implementing such video decoding techniques.

[0004] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy in a video sequence. For block-based video decoding, a video slice (e.g., a video picture or a portion of a video picture) can be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in a slice decoded intra-frame (I) of a picture are encoded using spatial prediction relative to reference samples in neighboring blocks in the same picture. Video blocks in a slice decoded inter-frame (P or B) of a picture can use spatial prediction relative to reference samples in neighboring blocks in the same picture or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0005] Generally, this disclosure describes techniques for gradient-based prediction refinement. A video decoder (e.g., a video encoder or a video decoder) (e.g., based on one or more motion vectors for a current block) determines one or more prediction blocks for performing inter prediction on the current block. In gradient-based prediction refinement, the video decoder modifies one or more samples of the prediction block based on various factors such as displacement in the horizontal direction, horizontal gradient, displacement in the vertical direction, and vertical gradient.

[0006] For example, a motion vector identifies a prediction block. The displacement in the horizontal direction (also referred to as horizontal displacement) refers to the change (e.g., increment) of the x coordinate of the motion vector, and the displacement in the vertical direction (also referred to as vertical displacement) refers to the change (e.g., increment) of the y coordinate. The horizontal gradient refers to the result of applying a filter to a first set of samples in the prediction block, and the vertical gradient refers to the result of applying a filter to a second set of samples in the prediction block.

[0007] The example techniques described in this disclosure provide gradient-based prediction refinement, where for different prediction modes, the precision level of the displacement (e.g., at least one of the horizontal displacement or the vertical displacement) is uniform (e.g., the same). For example, for a first prediction mode (e.g., affine mode), the motion vector may be at a first precision level, and for a second prediction mode (e.g., bi-directional optical flow (BDOF)), the motion vector may be at a second precision level. Thus, the vertical displacement and the horizontal displacement for the motion vector used for the affine mode and the motion vector used for BDOF may be different. In this disclosure, the video decoder may be configured to round (e.g., round up or round down) the vertical displacement and the horizontal displacement for the motion vector such that regardless of the prediction mode, the precision level of the displacement is the same (e.g., the vertical displacement and the horizontal displacement for the affine mode and BDOF have the same precision level).

[0008] By rounding the precision level of the displacement, the example techniques can improve the overall operation of the video decoder. For example, gradient-based prediction refinement involves multiplication operations and shift operations. If the precision level of the displacement is different for different modes, different logic circuits may be required to support different precision levels (e.g., a logic circuit configured for one precision level may not be suitable for other precision levels). Since the precision level of the displacement is the same for different modes, the same logic circuit can be reused for blocks, resulting in a smaller overall logic circuit and reduced power consumption because there is no need to power unused logic circuits.

[0009] In some examples, the techniques for determining displacements can be based on information already available at the video decoder. For example, the way the video decoder determines a horizontal displacement or a vertical displacement can be based on information that the video decoder can use for inter - frame prediction of the current block according to an inter - frame prediction mode. Additionally, there may be certain inter - frame prediction modes that are disabled for certain block types (e.g., based on size). In some examples, these inter - frame prediction modes that are disabled for certain block types can be enabled for those block types, but the predicted blocks for such blocks can be modified using the example techniques described in this disclosure.

[0010] In one example, this disclosure describes a method for decoding video data, the method including determining a predicted block for inter - frame prediction of a current block, determining a horizontal displacement and a vertical displacement for gradient - based prediction refinement of one or more samples of the predicted block, rounding the horizontal displacement and the vertical displacement to the same precision level for different inter - frame prediction modes including an affine mode and a BDOF mode, determining one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement, modifying one or more samples of the predicted block based on the determined one or more refinement offsets to generate a modified predicted block, and reconstructing the current block based on the modified predicted block.

[0011] In one example, this disclosure describes a method for encoding video data, the method including determining a predicted block for inter - frame prediction of a current block, determining a horizontal displacement and a vertical displacement for gradient - based prediction refinement of one or more samples of the predicted block; rounding the horizontal displacement and the vertical displacement to the same precision level for different inter - frame prediction modes including an affine mode and a BDOF mode, determining one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement, modifying one or more samples of the predicted block based on the determined one or more refinement offsets to generate a modified predicted block, determining a residual value indicating a difference between the current block and the modified predicted block, and signaling information indicating the residual value.

[0012] In one example, the present disclosure describes an apparatus for decoding video data. The apparatus includes a memory configured to store one or more samples of a prediction block and a processing circuit. The processing circuit is configured to: determine a prediction block for performing inter prediction on a current block; determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block; round the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes including an affine mode and a BDOF mode; determine one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement; modify one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block; and decode the current block based on the modified prediction block.

[0013] In one example, the present disclosure describes a computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform operations including: determining a prediction block for performing inter prediction on a current block; determining a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block; rounding the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes including an affine mode and a BDOF mode; determining one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement; modifying one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block; and decoding the current block based on the modified prediction block.

[0014] In one example, the present disclosure describes an apparatus for decoding video data. The apparatus includes: a unit configured to determine a prediction block for performing inter prediction on a current block; a unit configured to determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block; a unit configured to round the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes including an affine mode and a BDOF mode; a unit configured to determine one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement; a unit configured to modify one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block; and a unit configured to decode the current block based on the modified prediction block.

[0015] Details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a block diagram illustrating an example video encoding and decoding system that may execute the techniques of the present disclosure.

[0017] Figure 2A and Figure 2B is a conceptual diagram showing an example quadtree binary tree (QTBT) structure and a corresponding coding tree unit (CTU).

[0018] Figure 3 is a block diagram showing an example video encoder that can perform the techniques of the present disclosure.

[0019] Figure 4 is a block diagram showing an example video decoder that can perform the techniques of the present disclosure.

[0020] Figure 5 is a conceptual diagram showing an extended coding unit (CU) region used in bidirectional optical flow (BDOF).

[0021] Figure 6 is a conceptual diagram showing an example of sub-block motion vector (MV) selection.

[0022] Figure 7 is a flowchart showing an example method of coding video data. Detailed Description

[0023] The present disclosure relates to gradient-based prediction refinement. In gradient-based prediction refinement, a video coder (e.g., a video encoder or a video decoder) determines a prediction block for a current block as part of inter prediction based on a motion vector, and modifies (e.g., refines) the samples of the prediction block to generate modified prediction samples (e.g., refined prediction samples). The video encoder signals a residual value indicating the difference between the modified prediction samples and the current block. The video decoder performs the same operations as the video encoder performs to modify the samples of the prediction block to generate modified prediction samples. The video decoder adds the residual value to the modified prediction samples to reconstruct the current block.

[0024] One example way to modify the samples of the prediction block is to have the video coder determine one or more refinement offsets and add the samples of the prediction block to the refinement offsets. One example method of generating the refinement offsets is based on gradients and motion vector displacements. The gradients can be determined according to a gradient filter applied to the samples of the prediction block.

[0025] Examples of motion vector displacements include a horizontal displacement of the motion vector and a vertical displacement of the motion vector. The horizontal displacement can be a value added to or subtracted from the x coordinate of the motion vector, and the vertical displacement can be a value added to or subtracted from the y coordinate of the motion vector. For example, the horizontal displacement can be referred to as Δv x , where v xis the x coordinate of the motion vector, and the vertical displacement can be referred to as Δv y , where v y is the y coordinate of the motion vector.

[0026] The precision level of the motion vector of the current block can be different for different inter prediction modes. For example, the coordinates of the motion vector (e.g., the x coordinate or the y coordinate) include an integer part and may include a fractional part. The fractional part is referred to as the sub-pixel part of the motion vector because the integer part of the motion vector identifies the actual pixel in the reference picture including the prediction block, and the sub-pixel part of the motion vector adjusts the motion vector to identify the position between pixels in the reference picture.

[0027] The precision level of the motion vector is based on the sub-pixel part of the motion vector and indicates the granularity of the movement of the motion vector from the actual pixel in the reference picture. As an example, if the sub-pixel part of the x coordinate is 0.5, the motion vector is located in the middle of two horizontal pixels in the reference picture. If the sub-pixel part of the x coordinate is 0.25, the motion vector is at a quarter of the distance between two horizontal pixels, and so on. In these examples, the precision level of the motion vector can be equal to the sub-pixel part (e.g., the precision level is 0.5, 0.25, etc.).

[0028] In some examples, the precision levels of the horizontal displacement and the vertical displacement can be based on the precision level of the motion vector or the way the motion vector is generated. For example, in some examples, such as the merge mode (which is a form of inter prediction mode), the sub-pixel parts of the x coordinate and the y coordinate of the motion vector can be the horizontal displacement and the vertical displacement respectively. As another example, such as the affine mode (which is a form of inter prediction), the motion vector can be based on a corner motion vector, and the horizontal displacement and the vertical displacement can be determined based on the corner motion vector.

[0029] The precision levels of the horizontal displacement and the vertical displacement can be different for different inter prediction modes. For example, for some inter prediction modes, the horizontal displacement and the vertical displacement may be more precise (e.g., for the first prediction mode, the precision level is 1 / 128) compared to other inter prediction modes (e.g., for the second prediction mode, the precision level is 1 / 16).

[0030] In an implementation, a video decoder may need to include different logic circuits for handling different levels of precision. Performing gradient-based prediction refinement involves multiplication, shift operations, addition, and other arithmetic operations. A logic circuit configured for one level of precision for horizontal or vertical displacement may not be able to handle higher levels of precision for horizontal and vertical displacement. Thus, some video decoders include a set of logic circuits for performing gradient-based prediction refinement for one inter prediction mode (where horizontal and vertical displacements have a first level of precision) and a different set of logic circuits for performing gradient-based prediction refinement for another inter prediction mode (where horizontal and vertical displacements have a second level of precision).

[0031] However, having different logic circuits for performing gradient-based prediction refinement for different inter prediction modes results in an increase in the size of the video decoder and additional logic circuits that consume additional power. For example, if the current block is inter predicted in the first mode, the first set of logic circuits for gradient-based prediction refinement is used. However, the second set of logic circuits for gradient-based prediction refinement for a different inter prediction mode is still receiving power.

[0032] This disclosure describes examples of techniques for rounding the levels of precision for horizontal and vertical displacements to the same level of precision for different inter prediction modes. For example, a video decoder may round a first displacement (e.g., a first horizontal displacement or a first vertical displacement) having a first level of precision for a first block inter predicted in a first inter prediction mode to a set level of precision, and may round a second displacement (e.g., a second horizontal displacement or a second vertical displacement) having a second level of precision for a second block inter predicted in a second inter prediction mode to the same set level of precision. In other words, the video decoder may round at least one of the horizontal and vertical displacements to the same level of precision for different inter prediction modes. As an example, the first inter prediction mode may be an affine mode, and the second inter prediction mode may be a bi-directional optical flow (BDOF).

[0033] In this way, the same logic circuits can be used for gradient-based prediction refinement for different inter prediction modes, rather than having different logic circuits for different inter prediction modes. For example, the logic circuits of a video decoder may be configured to perform gradient-based prediction refinement for horizontal and vertical displacements having a set level of precision. The video decoder may round the horizontal and vertical displacements such that the precision levels of the rounded horizontal and rounded vertical displacements are equal to the set level of precision, thereby allowing the same logic circuits to perform gradient-based prediction refinement for different inter prediction modes.

[0034] Figure 1 FIG. 1 is a block diagram illustrating an example video encoding and decoding system 100 that can perform the techniques of the present disclosure. The techniques of the present disclosure generally relate to decoding (encoding and / or decoding) video data. Generally, video data includes any data for processing video. Thus, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (such as signaling data).

[0035] As Figure 1 shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. Specifically, source device 102 provides video data to destination device 116 via a computer-readable medium 110. Source device 102 and destination device 116 can include any of a variety of devices, including desktop computers, notebooks (e.g., laptops) computers, tablet computers, set-top boxes, telephone handsets (such as smart phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receiver devices, set-top boxes, etc. In some cases, source device 102 and destination device 116 can be equipped for wireless communication and can thus be referred to as wireless communication devices.

[0036] In Figure 1 the example of FIG. 1, source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. In accordance with the present disclosure, video encoder 200 of source device 102 and video decoder 300 of destination device 116 can be configured to apply techniques for gradient-based prediction refinement. Thus, source device 102 represents an example of a video encoding device, and destination device 116 represents an example of a video decoding device. In other examples, source and destination devices can include other components or arrangements. For example, source device 102 can receive video data from an external video source (such as an external camera). Similarly, destination device 116 can interface with an external display device instead of including an integrated display device.

[0037] As in Figure 1The illustrated system 100 is merely an example. In general, any digital video encoding and / or decoding device may perform techniques for gradient-based prediction refinement. Source device 102 and destination device 116 are merely examples of such decoding devices, where source device 102 generates encoded video data for transmission to destination device 116. This disclosure refers to a "decoding" device as a device that performs decoding (encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of decoding devices (specifically, a video encoder and a video decoder), respectively. In some examples, devices 102, 116 may operate in a substantially symmetric manner such that each of devices 102, 116 includes video encoding and decoding components. Thus, system 100 may support unidirectional or bi-directional video transmission between video devices 102, 116, e.g., for video streaming, video playback, video broadcast, or video telephony.

[0038] In general, video source 104 represents a source of video data (i.e., raw, unencoded video data), and provides a continuous sequence of pictures (also referred to as "frames") of the video data to video encoder 200, which encodes the data for the pictures. The video source 104 of source device 102 may include a video capture device (e.g., a camera), a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As another alternative, video source 104 may generate computer graphics-based data as source video or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 encodes the captured, pre-captured, or computer-generated video data. Video encoder 200 may reorder the pictures from the received order (sometimes referred to as "display order") into a decoding order for decoding. Video encoder 200 may generate a bitstream including the encoded video data. Source device 102 may then output the encoded video data via output interface 108 onto a computer-readable medium 110 for reception and / or retrieval by, e.g., input interface 122 of destination device 116.

[0039] The memory 106 of the source device 102 and the memory 120 of the destination device 116 represent general memories. In some examples, the memories 106, 120 may store raw video data, e.g., raw video from the video source 104 and raw decoded video data from the video decoder 300. Additionally or alternatively, the memories 106, 120 may store software instructions executable, e.g., by the video encoder 200 and the video decoder 300, respectively. Although shown separately from the video encoder 200 and the video decoder 300 in this example, it should be understood that the video encoder 200 and the video decoder 300 may also include internal memories for functionally similar or equivalent purposes. Further, the memories 106, 120 may store encoded video data (e.g., output from the video encoder 200 and input to the video decoder 300). In some examples, a portion of the memories 106, 120 may be allocated as one or more video buffers for storing, e.g., raw, decoded, and / or encoded video data.

[0040] The computer-readable medium 110 may represent any type of medium or device capable of transmitting the encoded video data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium for enabling the source device 102 to directly send the encoded video data to the destination device 116 in real time, e.g., via a radio frequency network or a computer-based network. In accordance with a communication standard (such as a wireless communication protocol), the output interface 108 may modulate the transmission signal including the encoded video data, and the input interface 122 may demodulate the received transmission signal. The communication medium may include any wireless communication medium or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (such as a local area network, a wide area network, or a global network such as the Internet). The communication medium may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from the source device 102 to the destination device 116.

[0041] In some examples, the source device 102 may output the encoded data from the output interface 108 to the storage device 112. Similarly, the destination device 116 may access the encoded data from the storage device 112 via the input interface 122. The storage device 112 may include any distributed or locally accessible data storage medium (such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data) among various distributed or locally accessible data storage media.

[0042] In some examples, the source device 102 may output the encoded video data to a file server 114 or to another intermediate storage device that may store the encoded video generated by the source device 102. The destination device 116 may access the stored video data from the file server 114 via streaming or downloading. The file server 114 may be any type of server device capable of storing the encoded video data and sending the encoded video data to the destination device 116. The file server 114 may represent, for example, a web server (for a website), a File Transfer Protocol (FTP) server, a content delivery network device, or a Network Attached Storage (NAS) device. The destination device 116 may access the encoded video data from the file server 114 via any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The file server 114 and the input interface 122 may be configured to operate according to a streaming transport protocol, a download transport protocol, or a combination thereof.

[0043] The output interface 108 and the input interface 122 may represent a wireless transmitter / receiver, a modem, a wired network component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and the input interface 122 include wireless components, the output interface 108 and the input interface 122 may be configured to transmit data, such as encoded video data, according to a cellular communication standard, such as 4G, 4G-LTE (Long Term Evolution), enhanced LTE, 5G, etc. In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 may be configured to transmit data, such as encoded video data, according to other wireless standards, such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee TM )), Bluetooth TM standards, etc. In some examples, the source device 102 and / or the destination device 116 may include respective System-on-a-Chip (SoC) devices. For example, the source device 102 may include an SoC device for performing the functionality attributed to the video encoder 200 and / or the output interface 108, and the destination device 116 may include an SoC device for performing the functionality attributed to the video decoder 300 and / or the input interface 122.

[0044] The techniques of the present disclosure can be applied to video coding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium), or other applications.

[0045] The input interface 122 of the destination device 116 receives the encoded bitstream from a computer-readable medium 110 (e.g., a storage device 112, a file server 114, etc.). The encoded bitstream can include signaling information defined by the video encoder 200, which is also used by the video decoder 300, such as syntax elements having values that describe features and / or processing of video blocks or other decoded units (e.g., slices, pictures, groups of pictures, sequences, etc.). The display device 118 can represent any of a variety of display devices (such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device).

[0046] Although not shown in Figure 1 In some aspects, the video encoder 200 and the video decoder 300 can each be integrated with an audio encoder and / or an audio decoder, and can include appropriate multiplexing-demultiplexing units or other hardware and / or software to process a multiplexed stream including both audio and video in a common data stream. If applicable, the multiplexing-demultiplexing unit can conform to the ITU H.223 multiplexer protocol or other protocols such as the User Datagram Protocol (UDP).

[0047] The video encoder 200 and the video decoder 300 can each be implemented as any encoder circuit or decoder circuit of a variety of suitable encoder circuits or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially in software, the device can store instructions for the software in a suitable, non-transitory computer-readable medium, and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in their respective devices. Devices including the video encoder 200 and / or the video decoder 300 can include integrated circuits, microprocessors, and / or wireless communication devices, such as cellular phones.

[0048] The video encoder 200 and the video decoder 300 may operate according to a video coding standard also known as High Efficiency Video Coding (HEVC), such as ITU-T H.265, or an extension thereof, such as multi-view and / or scalable video coding extensions. Alternatively, the video encoder 200 and the video decoder 300 may operate according to other proprietary or industry standards also known as Versatile Video Coding (VVC), such as ITU-T H.266. A recent draft of the VVC standard is described in “Versatile Video Coding (Draft 4)” by Bross et al., Joint Video Exploration Team (JVET) of ITU-T SG16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 13th meeting: Marrakech, MA, January 9 - 18, 2019, JVET-M1001-v5 (hereinafter referred to as “VVC Draft 4”). A more recent draft of the VVC standard is described in “Versatile Video Coding (Draft 8)” by Bross et al., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 17th meeting: Brussels, Belgium, January 7 - 17, 2020, JVET-Q2001-vD (hereinafter referred to as “VVC Draft 8”). However, the techniques of the present disclosure are not limited to any particular coding standard.

[0049] Generally, the video encoder 200 and the video decoder 300 may perform block-based coding of pictures. The term “block” generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used during the encoding and / or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, the video encoder 200 and the video decoder 300 may code video data represented in the YUV (e.g., Y, Cb, Cr) format. That is, the video encoder 200 and the video decoder 300 may code the luminance component and the chrominance component, rather than coding the red, green, and blue (RGB) data of the samples for a picture, where the chrominance component may include both a red chrominance component and a blue chrominance component. In some examples, the video encoder 200 converts the received data in the RGB format to the YUV representation before encoding, and the video decoder 300 converts the YUV representation to the RGB format. Alternatively, a preprocessing unit and a postprocessing unit (not shown) may perform these conversions.

[0050] The present disclosure can generally relate to the decoding (e.g., encoding and decoding) of pictures, including the process of encoding or decoding data of a picture. Similarly, the present disclosure can relate to the decoding of blocks of a picture, including the process of encoding or decoding data for the block (e.g., prediction and / or residual decoding). An encoded video bitstream generally includes a series of values for representing decoding decisions (e.g., decoding modes) and syntax elements that partition a picture into blocks. Thus, a reference to decoding a picture or a block should generally be understood as the decoding values of the syntax elements that form the picture or the block.

[0051] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video decoder (such as video encoder 200) divides a coding tree unit (CTU) into CUs according to a quadtree structure. That is to say, the video decoder divides a CTU and a CU into four equal and non-overlapping squares, and each node of the quadtree has zero or four child nodes. A node without child nodes can be referred to as a "leaf node", and the CU of such a leaf node can include one or more PUs and / or one or more TUs. The video decoder can further divide PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the division of TUs. In HEVC, a PU represents inter-prediction data, and a TU represents residual values. A CU that is intra-predicted includes intra-prediction information (such as an intra-mode indication).

[0052] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, a video decoder (such as video encoder 200) divides a picture into a plurality of coding tree units (CTUs). Video encoder 200 can divide a CTU according to a tree structure (such as a quadtree binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple partition types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level divided according to quadtree partitioning and a second level divided according to binary tree partitioning. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0053] In the MTT partitioning structure, a block can be divided using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning. Ternary tree partitioning is a partitioning in which a block is divided into three sub-blocks. In some examples, ternary tree partitioning divides a block into three sub-blocks without dividing the original block through the center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.

[0054] In some examples, video encoder 200 and video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance component and chrominance components, while in other examples, video encoder 200 and video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for the two chrominance components (or two QTBT / MTT structures for each chrominance component).

[0055] Video encoder 200 and video decoder 300 may be configured to use per HEVC quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures. For purposes of explanation, the techniques of the present disclosure are described with respect to QTBT partitioning. However, it should be understood that the techniques of the present disclosure may also be applied to video decoders configured to use quadtree partitioning or other types of partitioning.

[0056] The present disclosure may interchangeably use "NxN" and "N by N" to refer to the sample dimensions of a block (such as a CU or other video block) in the vertical and horizontal dimensions. For example, 16x16 samples or 16 by 16 samples. Generally, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and will have 16 samples in the horizontal direction (x = 16). Similarly, an NxN CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. The samples in a CU may be arranged in rows and columns. Additionally, a CU need not have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU may contain NxM samples, where M need not be equal to N.

[0057] Video encoder 200 encodes video data for a CU that represents prediction and / or residual information and other information. The prediction information indicates how to predict the CU in order to form a prediction block for the CU. The residual information generally represents the sample-by-sample difference between the samples of the CU before encoding and the prediction block.

[0058] To predict a CU, video encoder 200 can generally form a prediction block for the CU through inter prediction or intra prediction. Inter prediction generally refers to predicting a CU from data of a previously decoded picture, while intra prediction generally refers to predicting a CU from previously decoded data of the same picture. To perform inter prediction, video encoder 200 can use one or more motion vectors to generate a prediction block. Video encoder 200 can generally perform a motion search to identify a reference block that closely matches the CU in terms of, for example, the difference between the CU and the reference block. Video encoder 200 can calculate a difference metric using sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether the reference block closely matches the current CU. In some examples, video encoder 200 can use uni-directional prediction or bi-directional prediction to predict the current CU.

[0059] Some examples of VVC also provide an affine motion compensation mode that can be considered an inter prediction mode. In the affine motion compensation mode, video encoder 200 can determine two or more motion vectors representing non-translational motion such as zooming in or out, rotation, perspective motion, or other irregular motion types.

[0060] To perform intra prediction, video encoder 200 can select an intra prediction mode to generate a prediction block. Some examples of VVC provide sixty-seven intra prediction modes, including various directional modes as well as a planar mode and a DC mode. Generally, video encoder 200 selects an intra prediction mode that describes neighboring samples of the current block (e.g., the block of the CU) and predicts samples of the current block from the neighboring samples. Assuming that video encoder 200 decodes CTUs and CUs in a raster scan order (from left to right, top to bottom), such samples can generally be above, top-left, or left of the current block in the same picture as the current block.

[0061] Video encoder 200 encodes data representing the prediction mode for the current block. For example, for an inter prediction mode, video encoder 200 can encode data representing which one of the various available inter prediction modes is used and the motion information for the corresponding mode. For uni-directional inter prediction or bi-directional inter prediction, for example, video encoder 200 can use advanced motion vector prediction (AMVP) or the merge mode to encode the motion vectors. Video encoder 200 can use a similar mode to encode the motion vectors for the affine motion compensation mode.

[0062] After prediction (such as intra prediction or inter prediction of a block), the video encoder 200 may compute a residual value for the block. The residual value (such as a residual block) represents the sample-by-sample difference between the block and a predicted block for the block formed using the corresponding prediction mode. The video encoder 200 may apply one or more transforms to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 may apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, the video encoder 200 may apply a secondary transform after the first transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), etc. The video encoder 200 produces transform coefficients after applying one or more transforms.

[0063] As described above, after any transform used to produce transform coefficients, the video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to the process of quantizing the transform coefficients to potentially reduce the amount of data used to represent the coefficients and thus provide further compression. By performing the quantization process, the video encoder 200 may reduce the bit depth associated with some or all of the coefficients. For example, the video encoder 200 may round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 may perform a bitwise right shift of the value to be quantized.

[0064] After quantization, the video encoder 200 may scan the transform coefficients to produce a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan may be designed to place higher energy (and thus lower frequency) coefficients at the front of the vector and lower energy (and thus higher frequency) transform coefficients at the back of the vector. In some examples, the video encoder 200 may use a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 may entropy encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). The video encoder 200 may also entropy encode the values of syntax elements that describe metadata associated with the encoded video data for use by the video decoder 300 when decoding the video data.

[0065] To perform CABAC, video encoder 200 may assign contexts within context models to symbols to be sent. The context may relate to, for example, whether neighboring values of the symbol are zero values. Probability determination may be based on the context assigned to the symbol.

[0066] Video encoder 200 may further generate syntax data (such as block-based syntax data, picture-based syntax data, and sequence-based syntax data) for video decoder 300 in, for example, a picture header, a block header, a slice header, or other syntax data (such as a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS)). Video decoder 300 may similarly decode such syntax data to determine how to decode the corresponding video data.

[0067] In this way, video encoder 200 may generate a bitstream including encoded video data (e.g., syntax elements describing the partitioning of a picture into blocks (e.g., CUs) and syntax elements for prediction and / or residual information for the blocks). Finally, video decoder 300 may receive the bitstream and decode the encoded video data.

[0068] Generally, video decoder 300 performs a process inverse to the process performed by video encoder 200 to decode the encoded video data of the bitstream. For example, video decoder 300 may use CABAC to decode the values of the syntax elements for the bitstream in a manner substantially similar (although inverse) to the CABAC encoding process of video encoder 200. The syntax elements may define the partitioning information of a picture as CTUs and partition each CTU according to a corresponding partitioning structure (such as a QTBT structure) to define the CUs of the CTU. The syntax elements may further define prediction and residual information for blocks (e.g., CUs) of the video data.

[0069] The residual information may be represented by, for example, quantized transform coefficients. Video decoder 300 may inverse-quantize and inverse-transform the quantized transform coefficients of a block to reproduce the residual block for the block. Video decoder 300 uses the signaled prediction mode (intra prediction or inter prediction) and associated prediction information (e.g., motion information for inter prediction) to form a prediction block for the block. Video decoder 300 may then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. Video decoder 300 may perform additional processing, such as performing a deblocking process to reduce visual artifacts along block boundaries.

[0070] The present disclosure can generally refer to "signaling" certain information, such as syntax elements. The term "signaling" can generally refer to the conveyance of values, syntax elements, and / or other data used to decode encoded video data. That is, the video encoder 200 can signal the value of a syntax element in a bitstream. Generally, signaling refers to generating values in a bitstream. As described above, the source device 102 can transmit the bitstream to the destination device 116 substantially in real time or non-real time (such as may occur when storing the syntax elements in the storage device 112 for later retrieval by the destination device 116).

[0071] In accordance with the techniques of the present disclosure, the video encoder 200 and the video decoder 300 can be configured to perform gradient-based prediction refinement. As described above, as part of performing inter prediction on a current block, the video encoder 200 and the video decoder 300 can determine one or more prediction blocks for the current block (e.g., based on one or more motion vectors). In gradient-based prediction refinement, the video encoder 200 and the video decoder 300 modify one or more samples of the prediction block (e.g., including all samples).

[0072] For example, in gradient-based prediction refinement, an inter prediction sample at position (i,j) (e.g., a sample of a prediction block) is refined by an offset ΔI(i,j) that is derived from a displacement in the horizontal direction, a horizontal gradient, a displacement in the vertical direction, and a vertical gradient at position (i,j). In one example, the prediction refinement is described as: ΔI(i,j) = g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j), where g x (i,j) is the horizontal gradient, g y (i,j) is the vertical gradient, Δv x (i,j) is the displacement in the horizontal direction, and Δv y (i,j) is the displacement in the vertical direction.

[0073] The gradient of an image is a measure of the directional change in intensity or color in the image. For example, the gradient value is based on the rate of change of color or intensity in the direction of the maximum change in color or intensity with neighboring samples. As an example, if the rate of change is relatively high, the gradient value is greater than if the rate of change is relatively low.

[0074] Additionally, the predicted block for the current block may be in a reference picture different from the current picture including the current block. The video encoder 200 and the video decoder 300 may determine an offset (e.g., ΔI(i,j)) based on sample values in the reference picture (e.g., the gradient is determined based on sample values in the reference picture). In some examples, the values used to determine the gradient may be values within the predicted block itself, or values generated based on the values of the predicted block (e.g., interpolated, rounded, etc. values generated from the values within the predicted block). Additionally, in some examples, the values used to determine the gradient may be outside the predicted block and within the reference picture or are generated from samples outside the predicted block and within the reference picture (e.g., interpolated, rounded, etc.).

[0075] However, in some examples, the video encoder 200 and the video decoder 300 may determine the offset based on sample values in the current picture. In some examples, such as intra-block copy, the current picture and the reference picture are the same picture.

[0076] The displacement (e.g., vertical displacement and / or horizontal displacement) may be determined based on the inter-frame prediction mode. In some examples, the displacement is determined based on motion parameters. As described in more detail, for the decoder-side motion refinement mode, the displacement may be based on samples in the reference picture. For other inter-frame prediction modes, the displacement may not be based on samples in the reference picture, but the example techniques are not limited thereto, and samples in the reference picture may be used to determine the displacement. There may be various ways to determine the vertical displacement and / or the horizontal displacement, and the techniques are not limited to a specific way of determining the vertical displacement and / or the horizontal displacement.

[0077] An example way to perform gradient calculation is described below. For example, for a gradient filter, in one example, the Sobel filter may be used for gradient calculation. The gradient is calculated as follows: g x (i,j) = I(i+1,j-1) - I(i-1,j-1) + 2*I(i+1,j) - 2*I(i-1,j) + I(i+1,j+1) - I(i-1,j+1) and g y (i,j) = I(i-1,j+1) - I(i-1,j-1) + 2*I(i,j+1) - 2*I(i,j-1) + I(i+1,j+1) - I(i+1,j-1).

[0078] In some examples, a [1,0,-1] filter is applied. The gradient may be calculated as follows: g x (i,j) = I(i+1,j) - I(i-1,j) and g y(i,j) = I(i,j + 1) - I(i,j - 1). In some examples, other gradient filters (e.g., Canny filter) may be applied. The example techniques described in this disclosure are not limited to any particular gradient filter.

[0079] For gradient normalization, the computed gradient may be normalized before being used for refined offset derivation (e.g., before computing ΔI), or normalization may be done after refined offset derivation. A rounding process may be applied during normalization. For example, if a [1,0,-1] filter is applied, normalization is performed by adding one to the input value and then shifting it right by one. If the input is scaled by a power of two, normalization is performed by adding 1 << N and then shifting right by (N + 1).

[0080] For the gradient at the boundary, the gradient at the boundary of the prediction block may be calculated by expanding the prediction block by S / 2 at each boundary, where S is the filtering step size used for gradient calculation. In one example, the expanded prediction samples are generated by using the same motion vector as the prediction block used for inter-frame prediction (motion compensation). In some examples, the expanded prediction samples are generated by using the same motion vector but using a shorter filter for the interpolation process in motion compensation. In some examples, the expanded prediction samples are generated by using the rounded motion vector for integer motion compensation. In some examples, the expanded prediction samples are generated by padding, where padding is performed by copying the boundary samples. In some examples, if the prediction block is generated by sub-block based motion compensation, the expanded prediction samples are generated by using the motion vector of the nearest sub-block. In some examples, if the prediction block is generated by sub-block based motion compensation, the expanded prediction samples are generated by using a representative motion vector. In one example, the representative motion vector may be the motion vector at the center of the prediction block. In one example, the representative motion vector may be derived by averaging the motion vectors of the boundary sub-blocks.

[0081] Sub-block based gradient derivation may be applied to facilitate parallel processing in hardware or a pipeline-friendly design. The width and height of the sub-blocks (denoted as sbW and sbH) may be determined as follows: sbW = min(blkW, SB_WIDTH) and sbH = min(blkH, SB_HEIGHT). In this equation, blkW and blkH are the width and height of the prediction block respectively. SB_WIDTH (SB width) and SB_HEIGHT (SB height) are two pre-determined variables. In one example, both SB_WIDTH and SB_HEIGHT are equal to 16.

[0082] For horizontal and vertical displacements, in some examples, the horizontal displacement Δv x (i,j) and the vertical displacement Δv y (i,j) can be determined according to an inter-frame prediction mode. However, the example techniques are not limited to determining the horizontal and vertical displacements based on an inter-frame prediction mode.

[0083] For inter-frame modes of small block sizes (e.g., small-sized blocks predicted inter-frame), to reduce the worst-case memory bandwidth, the inter-frame prediction mode for small blocks can be disabled or constrained. For example, disabling inter-frame prediction for 4x4 blocks or smaller blocks can disable bidirectional prediction for 4x8, 8x4, 4x16, and 16x4. Due to the interpolation process for these small blocks, the memory bandwidth may increase. Integer motion compensation without interpolation can still be applied to those small blocks without increasing the worst-case memory bandwidth.

[0084] In one or more example techniques, inter-frame prediction can be enabled for some or all of those small blocks, but integer motion compensation and gradient-based prediction refinement are utilized. First, the motion vector is rounded to an integer motion vector for motion compensation. Then, the remainder of the rounding (i.e., the sub-pixel part of the motion vector) is used as Δv x (i,j) and Δv y (i,j) for gradient-based prediction refinement. For example, if the motion vector for a small block is (2.25, 5.75), the integer motion vector used for motion compensation will be (2, 6), and the horizontal displacement (e.g., Δv x (i,j)) will be 0.25, and the vertical displacement (e.g., Δv y (i,j)) will be 0.75. In this example, the precision level of the horizontal and vertical displacements is 0.25 (or 1 / 4). For example, the horizontal and vertical displacements can be incremented in steps of 0.25.

[0085] In some examples, for small block size inter-frame modes, gradient-based prediction refinement can be available, but only when inter-frame prediction is performed on small-sized blocks in the merge mode. Examples of the merge mode are described below. In some examples, for small-sized inter-frame modes, gradient-based prediction refinement can be disabled for blocks with an integer motion mode. In the integer motion mode, one or more motion vectors (e.g., the motion vectors signaled) are integers. In some examples, even for larger-sized blocks, if inter-frame prediction is performed on the block in the integer motion mode, gradient-based prediction refinement can be disabled for such a block.

[0086] For the normal merge mode, which is an example of an inter-frame prediction mode where motion information is derived from spatially or temporally neighboring decoded blocks, Δvx (i,j) and Δv y (i,j) can be the remainder of a motion vector rounding process (e.g., similar to the above example of a motion vector (2.25, 5.75)). In one example, a temporal motion vector predictor is derived by scaling a motion vector in a temporal motion buffer according to a picture order count difference between a current picture and a reference picture. A rounding process can be performed to round the scaled motion vector to a certain precision. The remainder can be used as Δv x (i,j) and Δv y (i,j). The precision of the remainder (i.e., the precision level of the horizontal and vertical displacements) can be predefined and can be higher than the precision of motion vector prediction. For example, if the motion vector precision is 1 / 16, the remainder precision is 1 / (16 * MaxBlkSize) (1 / (16 * maximum block size)), where MaxBlkSize is the maximum block size. In other words, the precision level for the horizontal and vertical displacements (e.g., Δv x and Δv y ) is 1 / (16 * MaxBlkSize).

[0087] For the merge with motion vector difference (MMVD) mode, which is an example of an inter prediction mode, the motion vector difference is signaled together with a merge index to represent motion information. In some techniques, the motion vector difference (e.g., the difference between an actual motion vector and a motion vector prediction value) has the same precision as the motion vector. In one or more examples described in the present disclosure, it can be allowed for the motion vector difference to have a higher precision. First, the signaled motion vector difference is rounded to the motion vector precision, and the motion vector indicated by the merge index is added to generate a final motion vector for motion compensation. In one or more examples, the remainder after rounding (e.g., the difference between the rounded value of the motion vector difference and the original value of the motion vector difference) can be used as the horizontal and vertical displacements (e.g., used as Δv x (i,j) and Δv y (i,j)) for gradient-based prediction refinement. In some examples, Δv x (i,j) and Δv y (i,j) can be signaled as candidates for the motion vector difference.

[0088] For the motion vector refinement mode at the decoder side, motion compensation using the original motion vector is performed to generate an original bi-predictive block, and the difference between the prediction in list 0 and the prediction in list 1 is calculated, which is denoted as DistOrig. List 0 refers to the first reference picture list (RefPicList0) that includes reference picture lists that can potentially be used for inter-prediction. List 1 refers to the second reference picture list (RefPicList1) that includes reference picture lists that can potentially be used for inter-prediction. Then, the motion vectors at list 0 and list 1 are rounded to the closest integer positions. That is, the motion vector that references a picture in list 0 is rounded to the closest integer position, and the motion vector that references a picture in list 1 is rounded to the closest integer position. A search algorithm is used to search within the integer displacement range to perform motion compensation using the new integer motion vectors and find the displacement pair that has the minimum distortion DistNew between the block of the picture identified in the list 0 prediction and the block of the picture identified in list 1. If DistNew is less than DistOrig, the new integer motion vectors are fed into the bidirectional optical flow (BDOF) to derive Δv x (i,j) and Δv y (i,j) are used for prediction refinement at both the list 0 prediction and the list 1 prediction. Otherwise, BDOF is performed on the original list 0 prediction and list 1 prediction for prediction refinement.

[0089] For the affine mode, a motion field can be derived for each pixel (e.g., the motion vector can be determined on a per-pixel basis). However, a 4x4-based motion field is used for affine motion compensation to reduce complexity and memory bandwidth. For example, as an example, the motion vector is determined for a sub-block instead of on a per-pixel basis, where one sub-block is 4x4. Other sub-block sizes can also be used, such as 4x2, 2x4, or 2x2. In one or more examples, gradient-based prediction refinement can be used to improve affine motion compensation. The gradient of the block can be calculated as described above. Given an affine motion model: where a, b, c, d, e, and f are values determined by the video encoder 200 and the video decoder 300 based on the control point motion vectors and the length and width of the block, as several examples. In some examples, the values for a, b, c, d, e, and f can be signaled.

[0090] Some example methods for determining a, b, c, d, e, and f are described below. In a video coder (e.g., the video encoder 200 or the video decoder 300), in the affine mode, a picture is divided into sub-blocks for block-based coding. The affine motion model for a block can also be determined by three motion vectors (MVs) at three different positions that are not in the same row ( and ) These three positions are typically referred to as control points, and these three motion vectors are referred to as control point motion vectors (CPMVs). In the case where the three control points are at the three corners of the block, the affine motion can be described as:

[0091]

[0092] where blkW and blkH are the width and height of the block.

[0093] For the affine mode, the video encoder 200 and the video decoder 300 can use representative coordinates of the sub - blocks (e.g., the center position of the sub - blocks) to determine the motion vector for each sub - block. In one example, the block is divided into non - overlapping sub - blocks. If the block width is blkW, the block height is blkH, the sub - block width is sbW, and the sub - block height is sbH, then there are blkH / sbH rows of sub - blocks and blkW / sbW sub - blocks in each row. For the six - parameter affine motion model, the motion vector of the sub - block (referred to as sub - block MV) at the i - th row (0 <= i < blkW / sbW) and the j - th column (0 <= j < blkH / sbH) is derived as:

[0094]

[0095] According to the above equations, the variables a, b, c, d, e, and f can be defined as follows:

[0096]

[0097]

[0098]

[0099]

[0100] e = v 0x

[0101] f = v 0y

[0102] For the affine mode, which is an example of an inter - frame prediction mode, the video encoder 200 and the video decoder 300 can determine the displacement (e.g., horizontal displacement or vertical displacement) by at least one of the following methods. The following are examples and should not be considered restrictive. There may be other ways in which the video encoder 200 and the video decoder 300 can determine the displacement (e.g., horizontal displacement or vertical displacement) for the affine mode.

[0103] For 4x4 block-based affine motion compensation, for 2x2-based displacement derivation, the displacement within each 2x2 block is the same. Within each 4x4 block, the Δv(i,j) for the four 2x2 blocks within the 4x4 is calculated as follows:

[0104] Upper left 2x2:

[0105] Upper right 2x2:

[0106] Lower left 2x2:

[0107] Lower right 2x2:

[0108] For 1x1 displacement derivation, the displacement is derived for each sample. The coordinates of the upper left sample in the 4x4 can be (0, 0), in which case the Δv(i,j) is derived as:

[0109] In some examples, dividing by 2 (which is implemented as a right shift operation) can be moved to the refinement offset calculation. For example, the video encoder 200 and the video decoder 300 can perform the divide-by-2 operation as part of determining ΔI (e.g., the refinement offset), rather than performing the divide-by-2 operation when deriving the horizontal and vertical displacements (e.g., Δv x and Δv y ).

[0110] For 4x2 block-based affine motion compensation, the motion field for motion vector storage is still 4x4; however, the affine motion compensation is 4x2. The motion vector (MV) for the 4x4 block can be (v x ,v y ), in which case the MV for motion compensation of the left 4x2 is (v x -a,v y -c), and the MV for motion compensation of the right 4x2 is (v x +a,v y +c).

[0111] For 2x2-based displacement derivation, in the 2x2-based displacement derivation, the displacement within each 2x2 block is the same. Within each 4x2 block, the Δv(i,j) for the two 2x2 blocks within the 4x4 is calculated as follows:

[0112] Upper 2x2:

[0113] Lower 2x2:

[0114] For 1x1 displacement derivation, the displacement is derived for each sample. Suppose the coordinates of the top-left sample in 4x2 are (0, 0), and Δv(i,j) is derived as:

[0115] Dividing by 2 (which can be implemented as a right shift operation) can be moved to the refined offset calculation. For example, video encoder 200 and video decoder 300 can perform the division-by-2 operation as part of determining ΔI (e.g., the refined offset), rather than performing the division-by-2 operation when deriving the horizontal and vertical displacements (e.g., Δv x and Δv y ).

[0116] For 2x4 sub-block based affine motion compensation, the motion field for motion vector storage is still 4x4; however, the affine motion compensation is 2x4. The MV for a 4x4 sub-block can be (v x , v y ), in which case the MV for motion compensation of the left 4x2 is (v x - b, v y - d), and the MV for motion compensation of the right 4x2 is (v x + b, v y + d).

[0117] For 2x2 displacement derivation, the displacement within each 2x2 sub-block is the same. Within each 2x4 sub-block, the calculation of Δv(i,j) for the 2 2x2 sub-blocks within 4x4 is as follows:

[0118] Left 2x2:

[0119] Right 2x2:

[0120] For 1x1 displacement derivation, in the 1x1 based displacement derivation, the displacement is derived for each sample. The coordinates of the top-left sample in 2x4 can be (0, 0), in which case Δv(i,j) is derived as:

[0121] Dividing by 2 (which can be implemented as a right shift operation) can be moved to the refined offset calculation. For example, video encoder 200 and video decoder 300 can perform the division-by-2 operation as part of determining ΔI (e.g., the refined offset), rather than performing the division-by-2 operation when deriving the horizontal and vertical displacements (e.g., Δv x and Δv y ).

[0122] The following describes the prediction refinement for the affine mode. After performing sub-block based affine motion compensation, the prediction signal can be refined by adding an offset derived from the per-pixel motion and gradient of the prediction signal. The offset at position (m,n) can be calculated as:

[0123] ΔI(m,n) = g x (m,n) * Δv x (m,n) + g y (m,n) * Δv y (m,n)

[0124] where g x (m,n) is the horizontal gradient of the prediction signal and g y (m,n) is the vertical gradient. Δv x (m,n) and Δv y (m,n) are the differences between the x and y components of the motion vector calculated at the pixel position (m,n) and the sub-block MV. Let the coordinates of the top-left sample of the sub-block be (0,0), and the center of the sub-block is Given the affine motion parameters a, b, c, and d, Δv x (m,n) and Δv y (m,n) can be derived as:

[0125]

[0126]

[0127] In the control-point based affine motion model, the affine motion parameters a, b, c, and d are calculated from the CPMV as:

[0128]

[0129]

[0130]

[0131]

[0132] The following introduces the Bidirectional Optical Flow (BDOF). The Bidirectional Optical Flow (BDOF) tool is included in VTM4. BDOF was previously called BIO. BDOF can be used to refine the bidirectional prediction signal of a Coding Unit (CU) at the 4×4 sub-block level. The BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4x4 sub-block, the motion refinement (v x ,v y) is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples (e.g., the predicted samples from the reference pictures in the first reference picture list L0 and the predicted samples from the reference pictures in the second reference picture list L1). Motion refinement is then used to adjust the sample values of the bi-directional prediction in the 4x4 sub-blocks. The following steps are applied during the BDOF process.

[0133] First, the horizontal and vertical gradients of the two prediction signals ( and ) are calculated by directly computing the difference between two neighboring samples, i.e.,

[0134]

[0135]

[0136] where I (k) (i,j) is the sample value at the coordinate (i,j) of the prediction signal in list k (k = 0,1).

[0137] Then, the auto-correlations and cross-correlations of the gradients S 1 , S 2 , S 3 , S 5 and S 6 are calculated as:

[0138] S 1 = ∑ (i,j)∈Ω ψ x (i,j)·ψ x (i,j), S 3 = ∑ (i,j)∈Ω θ(i,j)·ψ x (i,j)

[0139]

[0140] S 5 = ∑ (i,j)∈Ω ψ y (i,j)·ψ y (i,j) S 6 = ∑ (i,j)∈Ω θ(i,j)·ψ y (i,j)

[0141] where,

[0142]

[0143]

[0144] θ(i,j) = (I(1) (i,j) >> n b )-(I (0) (i,j) >> n b )

[0145] where Ω is a 6×6 window surrounding the 4×4 sub-block.

[0146] Motion refinement (v x ,v y ) is then derived using the cross-correlation term and the auto-correlation term by the following formula:

[0147]

[0148]

[0149] where, is the rounding function.

[0150] Based on the motion refinement and the gradient, the following adjustment is calculated for each sample in the 4×4 sub-block:

[0151]

[0152] Finally, the BDOF samples of the CU are calculated by adjusting the dual-prediction samples as shown below:

[0153] pred BDOF (x,y) = (I (0) (x,y) + I (1) (x,y) + b(x,y) + o offset ) >> shift

[0154] o offset = 1 << (shift - 1), i.e., the rounding offset

[0155] These values are selected such that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process remains within 32 bits.

[0156] To derive the gradient value, some prediction samples I (k) (i,j) in the list k (k = 0, 1) outside the current CU boundary need to be generated. As in Figure 5As depicted in [description], the BDOF in VTM4 uses an extended row / column around the boundary of the CU. To control the computational complexity of generating prediction samples outside the boundary, the prediction samples in the extended region (white positions) are generated by directly obtaining the reference samples at the nearby integer positions without interpolation (using the floor() operation on the coordinates), and the normal 8-tap motion compensation interpolation filter is used to generate the prediction samples within the CU (gray positions). These extended sample values are only used in the gradient calculation. For the remaining steps in the BDOF process, if any samples and gradient values outside the CU boundary are needed, such samples are filled (i.e., repeated) from their closest neighbors.

[0157] The precision of the displacement and gradient is described below. In some examples, the same precision for the horizontal displacement and the vertical displacement can be used in all modes. The precision can be predefined or signaled in the high-level syntax. Thus, if the horizontal displacement and the vertical displacement are derived from different modes with different precisions, the horizontal displacement and the vertical displacement are rounded to the predefined precision. Examples of the predefined precision are: 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, 1 / 128, etc.

[0158] As described above, the precision (also referred to as the precision level) can indicate how precise the horizontal displacement and the vertical displacement (e.g., Δv x and Δv y ) are, where the horizontal displacement and the vertical displacement can be determined using one or more of the examples described above or using some other techniques. Generally, the precision level is defined as a decimal (e.g., 0.25, 0.125, 0.0625, 0.03125, 0.015625, 0.0078125, etc.) or a fraction (e.g., 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, 1 / 128, etc.). For example, for the 1 / 4 precision level, the horizontal or vertical displacement can be represented with an increment of 0.25 (e.g., 0.25, 0.5, or 0.75). For the 1 / 8 precision level, the horizontal displacement and the vertical displacement can be represented with an increment of 0.125 (e.g., 0.125, 0.25, 0.325, 0.5, 0.625, 0.75, or 0.825). It can be seen that the lower the value of the precision level (e.g., 1 / 8 is less than 1 / 4), the coarser the increment, and the more precise values can be presented (e.g., for the 1 / 4 precision level, the displacement is rounded to the closest quarter, but for the 1 / 8 precision level, the displacement is rounded to the closest eighth).

[0159] Because horizontal and vertical displacements can have different precision levels for different inter-frame prediction modes, video encoder 200 and video decoder 300 can be configured to include different logic circuits to perform gradient-based prediction refinement for different inter-frame prediction modes. As described above, to perform gradient-based prediction refinement, video encoder 200 and video decoder 300 can perform the following operations: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x and g y are the first gradient based on a first set of samples in the samples of the prediction block and the second gradient based on a second set of samples in the samples of the prediction block, respectively, and Δv x and Δv y are the horizontal displacement and the vertical displacement, respectively. It can be seen that for gradient-based prediction refinement, video encoder 200 and video decoder 300 may need to perform multiplication and addition operations, and utilize memory to store temporary results used in the calculations.

[0160] However, the ability of logic circuits (e.g., multiplier circuits, adder circuits, memory registers) to perform mathematical operations can be limited to the precision level that the logic circuits are configured for. For example, a logic circuit configured for a first precision level may not be able to perform the operations required for gradient prediction refinement where the horizontal or vertical displacement is at a more precise second precision level.

[0161] Therefore, some techniques utilize different sets of logic circuits configured for different precision levels to perform gradient-based prediction refinement for different inter-frame prediction modes. For example, a first set of logic circuits can be configured to perform gradient-based prediction refinement for an inter-frame prediction mode where the horizontal and / or vertical displacement is 0.25, while a second set of logic circuits can be configured to perform gradient-based prediction refinement for an inter-frame prediction mode where the horizontal and / or vertical displacement is 0.125. Having these different circuit sets increases the overall size of video encoder 200 and video decoder 300, and potentially wastes power.

[0162] In some examples described in the present disclosure, the same gradient calculation process can be used for all inter-frame prediction modes. In other words, the same logic circuit can be used to perform gradient-based prediction refinement for different inter-frame prediction modes. For example, for prediction refinement in all inter-frame prediction modes, the precision of the gradient can be kept the same. In some examples, for the precision of displacement and gradient, the example techniques can ensure that the same (or unified) prediction refinement process can be applied to different modes, and the same prediction refinement module can be applied to different modes.

[0163] As an example, video encoder 200 and video decoder 300 can be configured to round at least one of the horizontal displacement and the vertical displacement to the same precision level for different inter-frame prediction modes (e.g., the same for affine mode and BDOF). For example, if the precision level to which the horizontal displacement and the vertical displacement are rounded is 0.015625 (1 / 64), then if the precision level of the horizontal displacement and / or the vertical displacement is 1 / 4 for one inter-frame prediction mode, the precision level of the horizontal displacement and / or the vertical displacement is rounded to 1 / 64. If the precision level of the horizontal displacement and / or the vertical displacement is 1 / 128, the precision level of the horizontal displacement and / or the vertical displacement is rounded to 1 / 64.

[0164] In this way, the logic circuit for gradient-based prediction refinement can be reused for different inter-frame prediction modes. For example, in the above example, video encoder 200 and video decoder 300 can include a logic circuit for a precision level of 0.125, and this logic circuit can be reused for different inter-frame prediction modes because the precision level of the horizontal displacement and / or the vertical displacement is rounded to 0.125.

[0165] In some examples, when rounding is not performed according to the techniques described in the present disclosure, if the logic circuit is designed to have a relatively high level of precision, the logic circuit for multiplication and accumulation type operations can be reused (e.g., a logic circuit designed for a specific precision level for multiplication can handle multiplication operations for values of a lower precision level). However, for shift operations, a logic circuit designed for a specific precision may not be able to handle shift operations for values of a lower precision level. With the example techniques described in the present disclosure, with the described rounding techniques, it may be possible to reuse the logic circuit including shift operations for different inter-frame prediction modes.

[0166] In one example, the prediction refinement offset is derived as:

[0167] ΔI(i,j) = (g x (i,j) * Δv x (i,j) + g y(i,j)*Δv y (i,j)+offset)>>shift

[0168] In the above equation, the offset is equal to 1<<(shift-1), and the offset is determined by the predefined precision of the displacement and the gradient, and is fixed for different modes. In some examples, the offset is equal to 0.

[0169] In some examples, the mode may include one or more of the modes described above with respect to the horizontal displacement and the vertical displacement, such as the block size inter prediction mode, the normal merge mode, the merge with motion vector difference, the decoder-side motion vector refinement mode, and the affine mode. The mode may also include the bidirectional optical flow (BDOF) described above.

[0170] There may be a separate refinement for each prediction direction. For example, in the case of bidirectional prediction, the prediction refinement may be performed separately for each prediction direction. The result of the refinement may be clipped to a certain range to ensure the same bit width as the prediction without refinement. For example, the refinement result is clipped to a 16-bit range. As described above, the example techniques may also be applied to BDOF, where it is assumed that the displacements in two different directions are in the same motion trajectory.

[0171] The following describes the N-bit (e.g., 16-bit) multiplication constraint. To reduce the complexity of the gradient-based prediction refinement, the multiplication can be kept within N bits (e.g., 16 bits). In this example, the gradient and the displacement should be able to be represented by no more than 16 bits. If not, then in this example, the gradient or the displacement is quantized within 16 bits. For example, a right shift may be applied to maintain a 16-bit representation.

[0172] The following describes the clipping of the refinement offset ΔI(i,j) and the refinement result. The refinement offset ΔI(i,j) is clipped to a certain range. In one example, the range is determined by the range of the original prediction signal. The range of ΔI(i,j) may be the same as the range of the original prediction signal, or the range may be a scaled range. The scaling may be 1 / 2, 1 / 4, 1 / 8, etc. The refinement result is clipped to have the same range as the original prediction signal (e.g., the range of the samples in the prediction block). The equation for performing the clipping is:

[0173] pbSamples[x][y]=Clip3(0,(2 BitDepth )-1,(predSamplesL0[x+1][y+1]+offset4+predSamplesL1[x+1][y+1]+bdofOffset)>>shift4)

[0174] Clipping where predSamplesL0 and predSamplesL1 are prediction samples in each single prediction direction. bdofOffset is a refinement offset derived through BDOF. Offset4 = 1<<(shift4-1) and Clip3(min, max, x) clipping are functions for clipping the value of x within a range from minimum to maximum (including boundary values).

[0175] In this way, the video encoder 200 and the video decoder 300 can be configured to determine a prediction block for performing inter-frame prediction on a current block. For example, the video encoder 200 and the video decoder 300 can determine a motion vector or a block vector pointing to the prediction block (e.g., for the intra-block copy mode).

[0176] The video encoder 200 and the video decoder 300 can determine at least one of a horizontal displacement or a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block. An example of the horizontal displacement is Δv x , and an example of the vertical displacement is Δv y . In some examples, the video encoder 200 and the video decoder 300 can determine at least one of a horizontal displacement or a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block based on an inter-frame prediction mode (e.g., using the above example techniques for the affine mode to determine Δv x and Δv y , or using the above example techniques for the merge mode to determine Δv x and Δv y , as two examples).

[0177] According to one or more examples, the video encoder 200 and the video decoder 300 can round at least one of the horizontal displacement and the vertical displacement to the same precision level for different inter-frame prediction modes. Examples of different inter-frame prediction modes include the affine mode and BDOF. For example, the precision level for a first horizontal displacement or a first vertical displacement for performing gradient-based prediction refinement on a first block for inter-frame prediction in a first inter-frame prediction mode can be at a first precision level, and the precision level for a second horizontal displacement or a second vertical displacement for performing gradient-based prediction refinement on a second block for inter-frame prediction in a second inter-frame prediction mode can be at a second precision level. The video encoder 200 and the video decoder 300 can be configured to round the first precision level for the first horizontal displacement or the first vertical displacement to the precision level, and round the second precision level for the first horizontal displacement or the first vertical displacement to the same precision level.

[0178] In some examples, the precision level can be predefined (e.g., pre-stored on video encoder 200 and video decoder 300), or can be signaled (e.g., defined by video encoder 200 and signaled to video decoder 300). In some examples, the precision level can be 1 / 64.

[0179] Video encoder 200 and video decoder 300 can be configured to determine one or more refinement offsets based on at least one of the rounded horizontal displacement or vertical displacement. For example, video encoder 200 and video decoder 300 can use at least one of the respective rounded horizontal displacement or vertical displacement to determine ΔI(i,j) for each sample of a prediction block. That is, video encoder 200 and video decoder 300 can determine a refinement offset for each sample of a prediction block. In some examples, video encoder 200 and video decoder 300 can utilize the rounded horizontal displacement and the rounded vertical displacement to determine the refinement offset (e.g., ΔI).

[0180] As described, to perform gradient-based prediction refinement, video encoder 200 and video decoder 300 can determine a first gradient based on a first sample set of one or more samples of a prediction block (e.g., determine g x (i,j), where the first sample set is the samples used to determine g x (i,j)), and determine a second gradient based on a second sample set of one or more samples of a prediction block (e.g., determine g y (i,j), where the second sample set is the samples used to determine g y (i,j)). Video encoder 200 and video decoder 300 can determine the refinement offset based on the rounded horizontal displacement and the rounded vertical displacement and the first gradient and the second gradient.

[0181] Video encoder 200 and video decoder 300 can modify one or more samples of a prediction block based on the determined one or more refinement offsets to generate a modified prediction block (e.g., form one or more modified samples of the modified prediction block). For example, video encoder 200 and video decoder 300 can add or subtract ΔI(i,j) from I(i,j), where I(i,j) refers to the sample in the prediction block at position (i,j). In some examples, video encoder 200 and video decoder 300 can clip one or more refinement offsets (e.g., clip ΔI(i,j)). Video encoder 200 and video decoder 300 can modify one or more samples of a prediction block based on the clipped one or more refinement offsets.

[0182] For encoding, the video encoder 200 may determine (e.g., based on the modified samples of the modified prediction block) a residual value indicating the difference between the current block and the modified prediction block (e.g., of the residual block), and signal information indicating the residual value. For decoding, the video decoder 300 may receive the information indicating the residual value, and reconstruct the current block based on the modified prediction block (e.g., the modified samples of the modified prediction block) and the residual value (e.g., by adding the residual value to the modified samples).

[0183] Figure 2A and Figure 2B FIG. 6 is a conceptual diagram showing an exemplary quadtree binary tree (QTBT) structure 130 and a corresponding coding tree unit (CTU) 132. Solid lines represent quadtree partitions and dashed lines indicate binary tree partitions. In each partition (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which partition type (i.e., horizontal or vertical) is used, where 0 indicates a horizontal partition and 1 indicates a vertical partition in this example. For quadtree partitions, since a quadtree node divides a block horizontally and vertically into 4 sub-blocks of equal size, there is no need to indicate the partition type. Thus, the video encoder 200 may encode syntax elements (e.g., partition information) for the region tree level (i.e., solid lines) of the QTBT structure 130 and syntax elements (e.g., partition information) for the prediction tree level (i.e., dashed lines) of the QTBT structure 130, and the video decoder 300 may decode syntax elements (e.g., partition information) for the region tree level (i.e., solid lines) of the QTBT structure 130 and syntax elements (e.g., partition information) for the prediction tree level (i.e., dashed lines) of the QTBT structure 130. The video encoder 200 may encode video data (such as prediction and transform data) for the CUs represented by the terminal leaf nodes of the QTBT structure 130, and the video decoder 300 may decode video data (such as prediction and transform data) for the CUs represented by the terminal leaf nodes of the QTBT structure 130.

[0184] Generally Figure 2B the CTU 132 may be associated with parameters that define the size of the blocks corresponding to the nodes of the QTBT structure 130 at the first and second levels. These parameters may include the CTU size (representing the size of the CTU 132 in samples), the minimum quadtree size (MinQTSize, representing the minimum allowable quadtree leaf node size), the maximum binary tree size (MaxBTSize, representing the maximum allowable binary tree root node size), the maximum binary tree depth (MaxBTDepth, representing the maximum allowable binary tree depth), and the minimum binary tree size (MinBTSize, representing the minimum allowable binary tree leaf node size).

[0185] The root node of the QTBT structure corresponding to the CTU can have four child nodes at the first level of the QTBT structure, and each of the four child nodes can be divided according to quadtree partitioning. That is, the nodes at the first level are leaf nodes (without child nodes) or have four child nodes. An example of the QTBT structure 130 represents such a node as including a parent node and child nodes with solid lines for branching. If the nodes at the first level are not greater than the maximum allowable binary tree root node size (MaxBTSize), they can be further divided by their respective binary trees. The binary tree splitting of a node can be iterated until the nodes obtained from the splitting reach the minimum allowable binary tree leaf node size (MinBTSize) or the maximum allowable binary tree depth (MaxBTDepth). An example of the QTBT structure 130 represents such a node as having dotted lines for branching. The binary tree leaf nodes are called coding units (CUs), which are used for prediction (e.g., intra-frame or inter-frame prediction) and transformation without any further division. As discussed above, the CU can also be referred to as a "video block" or a "block".

[0186] In an example of the QTBT partitioning structure, the CTU size is set to 128x128 (luminance samples and two corresponding 64x64 chrominance samples), MinQTSize is set to 16x16, MaxBTSize is set to 64x64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. First, quadtree partitioning is applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If the quadtree leaf node is 128x128, it will not be further split by the binary tree because the size exceeds MaxBTSize (i.e., 64x64 in this example). Otherwise, the quadtree leaf node will be further divided by the binary tree. Thus, the quadtree leaf node is still the root node for the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (4 in this example), further splitting is not allowed. When a binary tree node has a width equal to MinBTSize (4 in this example), it means that further horizontal splitting is not allowed. Similarly, a binary tree node with a height equal to MinBTSize means that further vertical splitting is not allowed for that binary tree node. As described above, the leaf nodes of the binary tree are called CUs and are further processed according to prediction and transformation without further division.

[0187] Figure 3 FIG. is a block diagram showing an example video encoder 200 that can perform the techniques of the present disclosure.Figure 3 is provided for purposes of explanation and should not be construed as a limitation of the techniques as broadly illustrated and described in the present disclosure. For purposes of explanation, the present disclosure describes the video encoder 200 in the context of video coding standards such as the HEVC video coding standard and the H.266 video decoding standard under development. However, the techniques of the present disclosure are not limited to these video coding standards and are generally applicable to video encoding and decoding.

[0188] In Figure 3 an example of, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any one or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy coding unit 220 may be implemented in one or more processors or in processing circuitry. Additionally, the video encoder 200 may include additional or alternative processors or processing circuitry for performing these functions and other functions.

[0189] The video data memory 230 may store video data to be encoded by components of the video encoder 200. The video encoder 200 may receive the video data stored in the video data memory 230 from, for example, a video source 104 ( Figure 1 ). The DPB 218 may act as a reference picture memory that stores reference video data for use by the video encoder 200 in predicting subsequent video data. The video data memory 230 and the DPB 218 may be formed of any one of various storage devices such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The video data memory 230 and the DPB 218 may be provided by the same storage device or separate storage devices. In various examples, the video data memory 230 may be on-chip (as shown) or off-chip relative to the other components of the video encoder 200.

[0190] In the present disclosure, references to the video data memory 230 should not be construed as being limited to a memory internal to the video encoder 200 (unless specifically so described) or a memory external to the video encoder 200 (unless specifically so described). Rather, references to the video data memory 230 should be understood as a reference memory that stores video data received by the video encoder 200 for encoding (e.g., video data for a current block to be encoded). Figure 1 The memory 106 may also provide temporary storage for outputs from the various units of the video encoder 200.

[0191] Figure 3 The various units are shown to assist in understanding the operations performed by the video encoder 200. The units may be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Fixed-function circuitry refers to circuitry that provides a specific function and is pre-set in terms of the operations that can be performed. Programmable circuitry refers to circuitry that can be programmed to perform various tasks and provides flexible functionality in terms of the operations that can be performed. For example, programmable circuitry may execute software or firmware that causes the programmable circuitry to operate in a manner defined by instructions of the software or firmware. Fixed-function circuitry may execute software instructions (e.g., to receive parameters or output parameters), but the type of operations performed by the fixed-function circuitry is generally immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0192] The video encoder 200 may include an arithmetic logic unit (ALU), a basic function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core formed from programmable circuitry. In examples where software executed by programmable circuitry is used to perform the operations of the video encoder 200, the memory 106 ( Figure 1 ) may store the object code of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.

[0193] The video data memory 230 is configured to store received video data. The video encoder 200 may retrieve pictures of video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 may be the original video data to be encoded.

[0194] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, an intra prediction unit 226, and a gradient-based prediction refinement (GBPR) unit 227. The mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes. As an example, the mode selection unit 202 may include a palette unit, a block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0195] Although the GBPR unit 227 is shown as separate from the motion estimation unit 222 and the motion compensation unit 224, in some examples, the GBPR unit 227 may be part of the motion estimation unit 222 and / or the motion compensation unit 224. The GBPR unit 227 is shown separate from the motion estimation unit 222 and the motion compensation unit 224 for ease of understanding and should not be considered restrictive.

[0196] The mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and the resulting rate-distortion values for such combinations. The encoding parameters may include the CTU-to-CU partitioning, the prediction mode for the CU, the transform type for the residual values of the CU, the quantization parameter for the residual values of the CU, etc. The mode selection unit 202 may ultimately select the combination of encoding parameters that has a better rate-distortion value than other tested combinations.

[0197] The video encoder 200 may divide a picture retrieved from the video data memory 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 may divide the CTUs of the picture according to a tree structure, such as the QTBT structure or the quadtree structure of HEVC described above. As described above, the video encoder 200 may form one or more CUs by dividing the CTUs according to the tree structure. Such CUs may generally also be referred to as "video blocks" or "blocks".

[0198] Typically, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, the intra prediction unit 226, and the GBPR unit 227) to generate a prediction block for a current block (e.g., a current CU, or in HEVC, the overlapping portion of a PU and a TU). For inter prediction of a current block, the motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in the DPB 218). Specifically, the motion estimation unit 222 may calculate, for example, values representing how similar a potential reference block is to the current block according to the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), etc. The motion estimation unit 222 typically may perform these calculations using the per-sample differences between the current block and the reference block under consideration. The motion estimation unit 222 may identify the reference block having the lowest value derived from these calculations, which indicates the reference block that most closely matches the current block.

[0199] The motion estimation unit 222 may form one or more motion vectors (MVs) defining the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion estimation unit 222 may then provide the motion vectors to the motion compensation unit 224. For example, for uni-directional inter prediction, the motion estimation unit 222 may provide a single motion vector, while for bi-directional inter prediction, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then use the motion vectors to generate a prediction block. For example, the motion compensation unit 224 may use the motion vectors to retrieve the data of the reference block. As another example, if the motion vectors have fractional sample precision, the motion compensation unit 224 may interpolate the values for the prediction block according to one or more interpolation filters. Additionally, for bi-directional inter prediction, the motion compensation unit 224 may retrieve the data for the two reference blocks identified by the respective motion vectors and combine the retrieved data, for example, by per-sample averaging or weighted averaging.

[0200] As another example, for intra prediction or intra prediction coding, the intra prediction unit 226 may generate a prediction block from samples adjacent to the current block. For example, for the directional mode, the intra prediction unit 226 may typically mathematically combine the values of the adjacent samples and fill in the calculated values in a defined direction across the current block to produce a prediction block. As another example, for the DC mode, the intra prediction unit 226 may calculate the average of the samples adjacent to the current block and generate a prediction block to include this derived average for each sample of the prediction block.

[0201] The GBPR unit 227 can be configured to perform the example techniques for gradient-based prediction refinement described in this disclosure. For example, the GBPR unit 227, together with the motion compensation unit 224, can determine a prediction block for inter-frame prediction of a current block (e.g., based on the motion vectors determined by the motion estimation unit 222). The GBPR unit 227 can determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block (e.g., Δv x and Δv y ). As an example, the GBPR unit 227 can determine an inter-frame prediction mode for inter-frame prediction of the current block based on a determination made by the mode selection unit 202. In some examples, the GBPR unit 227 can determine the horizontal displacement and the vertical displacement based on the determined inter-frame prediction mode.

[0202] The GBPR unit 227 can round the horizontal displacement and the vertical displacement to the same precision level for different inter-frame prediction modes. For example, the current block can be a first current block, the prediction block can be a first prediction block, the horizontal displacement and the vertical displacement can be a first horizontal displacement and a first vertical displacement, and the rounded horizontal displacement and the vertical displacement can be a first rounded horizontal displacement and a first rounded vertical displacement. In some examples, the GBPR unit 227 can determine a second prediction block for inter-frame prediction of a second current block and determine a second horizontal displacement and a second vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block. The GBPR unit 227 can round the second horizontal displacement and the second vertical displacement to the same precision level to which the first horizontal displacement and the first vertical displacement are rounded to generate a second rounded horizontal displacement and a second rounded vertical displacement.

[0203] In some cases, the inter-frame prediction mode for inter-frame prediction of the first current block and the inter-frame prediction mode for the second current block can be different. For example, the first mode among different inter-frame prediction modes is an affine mode, and the second mode among different inter-frame prediction modes is a bi-directional optical flow (BDOF) mode.

[0204] The precision level to which the horizontal displacement and the vertical displacement are rounded can be predefined and stored for use by the GBPR unit 227, or the GBPR unit 227 can determine the precision level, and the video encoder 200 can signal the precision level. For example, the precision level is 1 / 64.

[0205] The GBPR unit 227 may determine one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement. For example, the GBPR unit 227 may determine a first gradient of a first sample set based on one or more samples of a prediction block (e.g., determining g x (i,j)) using the samples of the prediction block described above, and determine a second gradient of a second sample set based on one or more samples of the prediction block (e.g., determining g y (i,j)) using the samples of the prediction block described above. The GBPR unit 227 may determine one or more refinement offsets based on the rounded horizontal displacement, the rounded vertical displacement, the first gradient, and the second gradient. In some examples, if the value of one or more refinement offsets is too high (e.g., greater than a threshold), the GBPR unit 227 may clip the one or more refinement offsets.

[0206] The GBPR unit 227 may modify one or more samples of the prediction block based on the determined one or more refinement offsets or the clipped one or more refinement offsets to generate a modified prediction block (e.g., form one or more modified samples of the modified prediction block). For example, the GBPR unit 227 may determine: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample in one or more samples located at (i,j), Δv x (i,j) is the rounded horizontal displacement for the sample in one or more samples located at (i,j), g y (i,j) is the second gradient for the sample in one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for the sample in one or more samples located at (i,j). In some examples, for each sample in the sample (i,j) of the prediction block, Δv x and Δv y may be the same.

[0207] The resulting modified sample can form a prediction block (e.g., a modified prediction block) in gradient-based prediction refinement. That is, the modified prediction block serves as the prediction block in gradient-based prediction refinement. The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives the original unencoded version of the current block from the video data memory 230 and the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block for the current block. In some examples, the residual generation unit 204 may also determine the differences between the sample values in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, the residual generation unit 204 may be formed using one or more subtractor circuits that perform binary subtraction.

[0208] In an example where the mode selection unit 202 divides a CU into PUs, each PU may be associated with a luminance prediction unit and a corresponding chrominance prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As indicated above, the size of a CU may refer to the size of the luminance decoding block of the CU, and the size of a PU may refer to the size of the luminance prediction unit of the PU. Assuming that the size of a particular CU is 2Nx2N, the video encoder 200 may support PU sizes of 2Nx2N or NxN for intra prediction and symmetric PU sizes such as 2Nx2N, 2NxN, Nx2N, NxN, etc. for inter prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitioning of PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter prediction.

[0209] In an example where the mode selection unit 202 does not further divide a CU into PUs, each CU may be associated with a luminance decoding block and a corresponding chrominance decoding block. As above, the size of a CU may refer to the size of the luminance decoding block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0210] For other video decoding techniques (to name a few examples, such as in-block copy mode decoding, affine mode decoding, and linear model (LM) mode decoding), the mode selection unit 202 generates a prediction block for the current block being encoded via respective units associated with the decoding technique. In some examples (such as palette mode decoding), the mode selection unit 202 may not generate a prediction block but instead generate syntax elements indicating the manner in which the block is to be reconstructed based on a selected palette. In such a mode, the mode selection unit 202 may provide these syntax elements to the entropy coding unit 220 for encoding.

[0211] As described above, the residual generation unit 204 receives video data for a current block and a corresponding predicted block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the per-sample difference between the predicted block and the current block.

[0212] The transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, the transform processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transform processing unit 206 may perform multiple transforms (e.g., a primary transform and a secondary transform) such as a rotation transform on the residual block. In some examples, the transform processing unit 206 does not apply a transform to the residual block.

[0213] The quantization unit 208 may quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. The quantization unit 208 may quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) may adjust the degree of quantization applied to the coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and thus, the quantized transform coefficients may have lower precision than the original transform coefficients produced by the transform processing unit 206.

[0214] The inverse quantization unit 210 and the inverse transform processing unit 212 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficient block to reconstruct the residual block from the transform coefficient block. The reconstruction unit 214 may generate a reconstructed block corresponding to the current block (although potentially with some degree of distortion) based on the reconstructed residual block and the predicted block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add the samples of the reconstructed residual block to the corresponding samples of the predicted block generated by the mode selection unit 202 to produce the reconstructed block.

[0215] The filter unit 216 may perform one or more filtering operations on the reconstructed block. For example, the filter unit 216 may perform a deblocking operation to reduce block effect artifacts along the edges of the CU. In some examples, the operation of the filter unit 216 may be skipped.

[0216] Video encoder 200 stores the reconstructed blocks in DPB 218. For example, in an example where the operation of filter unit 216 is not required, reconstruction unit 214 may store the reconstructed blocks into DPB 218. In an example where the operation of filter unit 216 is required, filter unit 216 may store the filtered reconstructed blocks into DPB 218. Motion estimation unit 222 and motion compensation unit 224 may retrieve reference pictures formed by the reconstructed (and possibly filtered) blocks from DPB 218 to perform inter prediction on the blocks of a subsequently encoded picture. Additionally, intra prediction unit 226 may use the reconstructed blocks in DPB 218 of the current picture to perform intra prediction on other blocks in the current picture.

[0217] Generally, entropy coding unit 220 may perform entropy coding on syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 may perform entropy coding on the quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 may perform entropy coding on prediction syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from mode selection unit 202. Entropy coding unit 220 may perform one or more entropy coding operations on syntax elements as another example of video data to generate entropy-coded data. For example, entropy coding unit 220 may perform context-adaptive variable-length coding (CAVLC) operations, CABAC operations, variable-to-variable (V2V) length coding operations, syntax-based context-adaptive binary arithmetic coding (SBAC) operations, probability interval partitioning entropy (PIPE) coding operations, exponential-Golomb coding operations, or another type of entropy coding operation on the data. In some examples, entropy coding unit 220 may operate in a bypass mode where the syntax elements are not entropy-coded.

[0218] Video encoder 200 may output a bitstream that includes the entropy-coded syntax elements required to reconstruct the blocks of a slice or picture. Specifically, entropy coding unit 220 may output the bitstream.

[0219] The operations described above are described with respect to blocks. Such descriptions should be understood as operations for luminance decoding blocks and / or chrominance decoding blocks. As described above, in some examples, the luminance decoding block and the chrominance decoding block are the luminance component and the chrominance component of a CU. In some examples, the luminance decoding block and the chrominance decoding block are the luminance component and the chrominance component of a PU.

[0220] In some examples, for chrominance decoding blocks, operations performed relative to luminance decoding blocks need not be repeated. As an example, operations for identifying motion vectors (MVs) and reference pictures for luminance decoding blocks need not be repeated to identify MVs and reference pictures for chrominance blocks. Instead, the MVs for luminance decoding blocks can be scaled to determine the MVs for chrominance blocks, and the reference pictures can be the same. As another example, for luminance decoding blocks and chrominance decoding blocks, the intra prediction process can be the same.

[0221] Figure 4 FIG. 4 is a block diagram illustrating an example video decoder 300 that may perform the techniques of the present disclosure. Figure 4 FIG. 4 is provided for purposes of explanation and should not be considered a limitation of the techniques broadly illustrated and described in the present disclosure. For purposes of explanation, the present disclosure describes the video decoder 300 described in accordance with the techniques of VVC and HEVC. However, the techniques of the present disclosure may be performed by video decoding devices configured for other video coding standards.

[0222] In Figure 4 an example, video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any one or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 may be implemented in one or more processors or in processing circuitry. Additionally, video decoder 300 may include additional or alternative processors or processing circuitry for performing these functions and other functions.

[0223] Prediction processing unit 304 includes a motion compensation unit 316, an intra prediction unit 318, and a gradient-based prediction refinement (GBPR) unit 319. Prediction processing unit 304 may include additional units for performing prediction according to other prediction modes. As an example, prediction processing unit 304 may include a palette unit, a block copy unit (which may form part of motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, video decoder 300 may include more, fewer, or different functional components.

[0224] Although the GBPR unit 319 is shown as being separate from the motion compensation unit 316, in some examples, the GBPR unit 319 can be part of the motion compensation unit 316. The GBPR unit 319 is shown separate from the motion compensation unit 316 for ease of understanding and should not be considered restrictive.

[0225] The CPB memory 320 can store video data (such as an encoded video bitstream) to be decoded by components of the video decoder 300. The video data stored in the CPB memory 320 can be obtained, for example, from the computer-readable medium 110 ( Figure 1 ). The CPB memory 320 can include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Additionally, the CPB memory 320 can store video data other than the syntax elements of decoded pictures, such as temporary data representing the outputs of various units from the video decoder 300. The DPB 314 generally stores decoded pictures that the video decoder 300 can output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 can be formed by any one of various storage devices (such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices). The CPB memory 320 and the DPB 314 can be provided by the same storage device or separate storage devices. In various examples, the CPB memory 320 can be on-chip or off-chip relative to other components of the video decoder 300.

[0226] Additionally or alternatively, in some examples, the video decoder 300 can retrieve decoded video data from the memory 120 ( Figure 1 ). That is, the memory 120 can store data together with the CPB memory 320 as discussed above. Similarly, when some or all of the functions of the video decoder 300 are implemented in software executed by the processing circuitry of the video decoder 300, the memory 120 can store instructions to be executed by the video decoder 300.

[0227] In Figure 4 the various units shown are shown to aid in understanding the operations performed by the video encoder 200. The units can be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Similar to Figure 3, A fixed - function circuit refers to a circuit that provides a specific function and is pre - set in terms of the operations it can perform. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality in terms of the operations it can perform. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed - function circuit can execute software instructions (e.g., to receive parameters or output parameters), but the type of operations performed by the fixed - function circuit is generally immutable. In some examples, one or more units in a cell can be different circuit blocks (fixed - function or programmable), and in some examples, one or more units can be integrated circuits.

[0228] The video decoder 300 can include an ALU, an EFU, digital circuits, analog circuits, and / or a programmable core formed from a programmable circuit. In an example where the operation of the video decoder 300 is performed by software executed on the programmable circuit, on - chip or off - chip memory can store the instructions (e.g., object code) of the software that the video decoder 300 receives and executes.

[0229] The entropy decoding unit 302 can receive the encoded video data from the CPB and perform entropy decoding on the video data to reproduce the syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 can generate the decoded video data based on the syntax elements extracted from the bitstream.

[0230] Generally, the video decoder 300 reconstructs pictures on a block - by - block basis. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the “current block”).

[0231] The entropy decoding unit 302 can perform entropy decoding on the syntax elements that define the quantized transform coefficients of the quantized transform coefficient block and the transform information (such as the quantization parameter (QP) and / or the transform mode indication). The inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the degree of quantization and, accordingly, the degree of inverse quantization to be applied by the inverse quantization unit 306. The inverse quantization unit 306 can, for example, perform a bit - by - bit left - shift operation to inverse - quantize the quantized transform coefficients. The inverse quantization unit 306 can thereby form a transform coefficient block including the transform coefficients.

[0232] After the inverse quantization unit 306 forms a transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the coefficient block.

[0233] In addition, the prediction processing unit 304 generates a prediction block based on the prediction information syntax element entropy decoded by the entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is inter-predicted, the motion compensation unit 316 can generate a prediction block. In this case, the prediction information syntax element can indicate the reference picture in the DPB 314 from which the reference block is retrieved, and a motion vector that identifies the position of the reference block in the reference picture relative to the current block in the current picture. The motion compensation unit 316 can generally perform the inter-prediction process in a manner substantially similar to the manner described with respect to the motion compensation unit 224 ( Figure 3 ).

[0234] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, the intra-prediction unit 318 can generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, the intra-prediction unit 318 can generally perform the intra-prediction process in a manner substantially similar to the manner described with respect to the intra-prediction unit 226 ( Figure 3 ). The intra-prediction unit 318 can retrieve data of neighboring samples of the current block from the DPB 314.

[0235] As another example, if the prediction information syntax element indicates that gradient-based prediction refinement is enabled, the GBPR unit 319 can modify the samples of the prediction block to generate a modified prediction block for reconstructing the current block (e.g., generate modified samples that form the modified prediction block).

[0236] The GBPR unit 319 can be configured to perform the example techniques for gradient-based prediction refinement described in the present disclosure. For example, the GBPR unit 319 together with the motion compensation unit 316 can (e.g., based on the motion vector determined by the prediction processing unit 304) determine a prediction block for inter-predicting the current block. The GBPR unit 319 can determine a horizontal displacement and a vertical displacement of gradient-based prediction refinement for one or more samples of the prediction block (e.g., Δv x and Δv y)。As an example, the GBPR unit 319 may determine an inter prediction mode for inter predicting a current block based on a prediction information syntax element. In some examples, the GBPR unit 319 may determine a horizontal displacement and a vertical displacement based on the determined inter prediction mode.

[0237] The GBPR unit 319 may round the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes. For example, the current block may be a first current block, the prediction block may be a first prediction block, the horizontal displacement and the vertical displacement may be a first horizontal displacement and a first vertical displacement, and the rounded horizontal displacement and the vertical displacement may be a first rounded horizontal displacement and a first rounded vertical displacement. In some examples, the GBPR unit 319 may determine a second prediction block for inter predicting a second current block, and determine a second horizontal displacement and a second vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block. The GBPR unit 319 may round the second horizontal displacement and the second vertical displacement to the same precision level to which the first horizontal displacement and the first vertical displacement are rounded to generate a second rounded horizontal displacement and a second rounded vertical displacement.

[0238] In some cases, the inter prediction mode for inter predicting the first current block and the inter prediction mode for the second current block may be different. For example, the first mode in the different inter prediction modes is an affine mode, and the second mode in the different inter prediction modes is a bi-directional optical flow (BDOF) mode.

[0239] The precision level to which the horizontal displacement and the vertical displacement are rounded may be predefined and stored for use by the GBPR unit 319, or the GBPR unit 319 may receive information indicating the precision level in the information signaled (e.g., the precision level is signaled). As an example, the precision level is 1 / 64.

[0240] The GBPR unit 319 may determine one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement. For example, the GBPR unit 319 may determine a first gradient of a first sample set of one or more samples of the prediction block (e.g., determining g x (i,j) using the samples of the prediction block described above), and determine a second gradient of a second sample set of one or more samples of the prediction block (e.g., determining g y(i,j)). The GBPR unit 319 may determine one or more refined offsets based on the rounded horizontal displacement and the rounded vertical displacement, as well as the first gradient and the second gradient. In some examples, if the value of one or more refined offsets is too high (e.g., greater than a threshold), the GBPR unit 319 may clip the one or more refined offsets.

[0241] The GBPR unit 319 may modify one or more samples of the prediction block based on the determined one or more refined offsets or the clipped one or more refined offsets to generate a modified prediction block (e.g., form one or more modified samples of the modified prediction block). For example, the GBPR unit 319 may determine: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for a sample among one or more samples located at (i,j), Δv x (i,j) is the rounded horizontal displacement for a sample among one or more samples located at (i,j), g y (i,j) is the second gradient for a sample among one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for a sample among one or more samples located at (i,j). In some examples, for each sample in the sample (i,j) of the prediction block, Δv x and Δv y may be the same.

[0242] The resulting modified samples may form a prediction block in gradient-based prediction refinement. That is to say, the modified prediction block may be used as a prediction block in gradient-based prediction refinement. The reconstruction unit 310 may use the prediction block and the residual block to reconstruct the current block. For example, the reconstruction unit 310 may add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.

[0243] The filter unit 312 may perform one or more filtering operations on the reconstructed block. For example, the filter unit 312 may perform a deblocking operation to reduce block effect artifacts along the edges of the reconstructed block. The operation of the filter unit 312 is not necessarily performed in all examples.

[0244] Video decoder 300 may store the reconstructed blocks in DPB 314. As discussed above, DPB 314 may provide reference information to prediction processing unit 304, such as samples of the current picture for intra prediction and previously decoded pictures for subsequent motion compensation. Additionally, video decoder 300 may output the decoded pictures from the DPB for subsequent rendering on a display device (such as Figure 1 the display device 118 in

[0245] According to a first technique of the present disclosure, a video coder (e.g., video encoder 200 and / or video decoder 300) may derive the differences in the x and y components between a motion vector calculated at a pixel position (m,n) and a sub-block MV (i.e., Δv x (m,n) and Δv y (m,n)) based on the sub-block MV. For example, if the affine motion parameters a, b, c, d, e, and f in the derivation of Δv x (m,n) and Δv y (m,n) are calculated from CPMVS (control point motion vectors), then the CPMVS for each block may need to be stored in a motion buffer. This storage of CPMVS for each block can significantly increase the buffer size because CPMV has 3 MVs for each prediction direction instead of 1 MV in the normal inter-frame mode. Thus, the present disclosure describes that the video coder performs the derivation of Δv x (m,n) and Δv y (m,n) based on the sub-block MV.

[0246] For a 6-parameter affine model, 3 different sub-block MVs that are not all in the same sub-block row or column can be selected. In a 4-parameter affine model, 2 different sub-block MVs are selected. In some examples, the selected sub-block MVs can be used as similar to the CPMV described above, where and are in the same sub-block row, and and are in the same sub-block column. Then, parameter a is calculated as Parameter b is calculated as Parameter c is calculated as And parameter d is calculated as In the case of a 4-parameter affine mode, parameter a is calculated as Parameter c is calculated as Parameter b is set equal to -c, and parameter d is set equal to a. W is the distance between and , and H is the distance between and The distance between. However, in some examples, regardless of whether a 6-parameter affine model or a 4-parameter affine model is used, 3 sub-block MVs are selected.

[0247] The video decoder selects sub-block MVs such that W is equal to blkW / 2 and H is equal to blkH / 2. In one example, as Figure 6 shown in is the sub-block MV of the upper-left sub-block at position (0,0), and is the sub-block MV of the upper-middle sub-block at position (blkW / 2,0), and is the sub-block MV of the left-middle sub-block at position (0,blkH / 2). In another example, is the sub-block MV of the upper-middle sub-block at position (blkW / 2 - sbW,0), and is the sub-block MV of the upper-right sub-block at position (blkW - sbW,0), and is the sub-block MV of the middle sub-block at position (blkW / 2 - sbW,blkH / 2).

[0248] According to the second technique of the present disclosure, the video decoder may perform clipping on Δv x (m,n) and Δv y (m,n). The gradient-based refinement offset calculation may assume that Δv x (m,n) and Δv y (m,n) are small. In this technique, the video decoder may clip Δv x (m,n) and Δv y (m,n) such that the absolute value is less than or equal to a predefined threshold ΔTH.

[0249] As an example, a predefined threshold may be set such that in the offset calculation, Δv x (m,n) / Δv y (m,n) multiplied by the gradient g x (m,n) / g y (m,n) does not cause a buffer overflow. For example, if the budget for the multiplication result is 16 bits, the maximum absolute value is 1<<15 (1 bit for the sign), and ΔTH * g x (m,n) or ΔTH * g y (m,n) should not exceed 1<<15. Assuming the gradient is represented by k bits, ΔTH is set equal to 1<<(15 - k).

[0250] As another example, a predefined threshold (e.g., ΔTH) can be set equal to the same value as the value in the bidirectional optical flow (BDOF), i.e., th′ BIO . th′ BIO can represent half a pixel. If for Δv x (m,n) and Δv y (m,n) the basic unit is 1 / q pixel, then th′ BIO is q / 2.

[0251] As another example, the predefined threshold (e.g., ΔTH) can be set equal to the minimum value between th′ BIO and 1<<(15-k).

[0252] According to a third technique of the present disclosure, the video decoder can set the precision of Δv x (m,n) and Δv y (m,n) to the same precision as in the BDOF. In one example, the precision of Δv x (m,n) and Δv y (m,n) is determined by shfit1 in Section 1.3. Thus, one unit of Δv x (m,n) and Δv y (m,n) is 1 / (1<<shift1) pixel. In one example, shift is set equal to 6. In another example, shift1 is set equal to max(2, 14-bitDepth), where bitDepth is the internal bit depth of the video signal used for encoding / decoding.

[0253] According to a fourth technique of the present disclosure, the video decoder can use the same process as in the BDOF to perform gradient calculation for prediction refinement in the affine mode. Accordingly, the same module of the video decoder can be used for both the BDOF and the gradient calculation for prediction refinement in the affine mode. However, the video decoder can use different filling methods for the prediction samples in the extended region.

[0254] As an example, the video decoder can generate prediction samples in the extended region (white positions) by directly obtaining reference samples at nearby integer positions without interpolation (using the floor() operation on the coordinates).

[0255] As another example, the video decoder can generate prediction samples in the extended region (white positions) by directly obtaining reference samples at the closest integer positions without interpolation (using the round() operation on the coordinates).

[0256] As another example, if any sample values outside the sub-block boundary are needed, the video decoder can fill the required samples from its closest neighbor (i.e., repeat). This can also be applied to the gradient calculation in BDOF.

[0257] According to a fifth technique of the present disclosure, the video decoder can perform clipping on the refinement result. In inter prediction, the motion compensation prediction signal of a block is typically clipped to the same range as the original signal of the block. However, in bidirectional motion compensation, the motion compensation prediction signals for each direction are maintained at an intermediate precision and range to improve accuracy. After the weighted average process of bidirectional motion compensation, the result is rounded and clipped to the same range and precision as the original signal of the block. In this fifth technique, in the case of bidirectional prediction, the video decoder can clip the result of prediction refinement to have the same intermediate precision and range as in normal motion compensation. For example, if the number of bits for intermediate precision is 14, then the video decoder can clip the prediction refinement result to the range from -(1<<14) to (1<<14).

[0258] Figure 7 is a flowchart showing an example method for decoding video data. The current block may include a current CU. Figure 7 Examples are described with respect to a processing circuit. Examples of the processing circuit include fixed function and / or programmable circuits for a video encoder 200 (such as a GBPR unit 227) and a video decoder 300 (such as a GBPR unit 319).

[0259] In one or more examples, the memory may be configured to store samples of a predicted block. For example, DPB 218 or DPB 314 may be configured to store samples of a predicted block for inter prediction. Intra-block copy may be considered an example inter prediction mode, in which case the block vector for intra-block copy is an example of a motion vector.

[0260] The processing circuit may determine a predicted block (350) stored in the memory for inter predicting the current block. The processing circuit may determine a horizontal displacement and a vertical displacement of gradient-based prediction refinement for one or more samples of the predicted block (e.g., Δv x and Δv y )(352). As an example, the processing circuit may determine an inter prediction mode for inter predicting the current block. In some examples, the processing circuit may determine the horizontal displacement and the vertical displacement based on the determined inter prediction mode.

[0261] The processing circuit can round the horizontal displacement and the vertical displacement to the same precision level (354) for different inter-frame prediction modes. For example, the current block can be a first current block, the prediction block can be a first prediction block, the horizontal displacement and the vertical displacement can be a first horizontal displacement and a first vertical displacement, and the rounded horizontal displacement and the rounded vertical displacement can be a first rounded horizontal displacement and a first rounded vertical displacement. In some examples, the processing circuit can determine a second prediction block for inter-frame prediction of a second current block, and determine a second horizontal displacement and a second vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block. The processing circuit can round the second horizontal displacement and the second vertical displacement to the same precision level to which the first horizontal displacement and the first vertical displacement are rounded, to generate a second rounded horizontal displacement and a second rounded vertical displacement.

[0262] In some cases, the inter-frame prediction mode for inter-frame prediction of the first current block and the inter-frame prediction mode for the second current block can be different. For example, the first mode among different inter-frame prediction modes is an affine mode, and the second mode among different inter-frame prediction modes is a bi-directional optical flow (BDOF) mode.

[0263] The precision level to which the horizontal displacement and the vertical displacement are rounded can be predefined or signaled. As an example, the precision level is 1 / 64.

[0264] The processing circuit can determine one or more refinement offsets (356) based on the rounded horizontal displacement and the rounded vertical displacement. For example, the processing circuit can determine a first gradient of a first sample set based on one or more samples of the prediction block (e.g., determining g x (i,j) using the samples of the prediction block described above), and determine a second gradient of a second sample set based on one or more samples of the prediction block (e.g., determining g y (i,j) using the samples of the prediction block described above). The processing circuit can determine one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement and the first gradient and the second gradient. In some examples, if the value of one or more refinement offsets is too high (e.g., greater than a threshold), the processing circuit can clip the one or more refinement offsets.

[0265] The processing circuit can modify one or more samples of the prediction block based on the determined one or more refinement offsets or the clipped one or more refinement offsets to generate a modified prediction block (e.g., forming one or more modified samples of the modified prediction block) (358). For example, the processing circuit can determine: g x (i,j)*Δv x (i,j)+gy (i,j) * Δv y (i,j), where g x (i,j) is the first gradient of the samples in one or more samples located at (i,j), Δv x (i,j) is the rounded horizontal displacement of the samples in one or more samples located at (i,j), g y (i,j) is the second gradient of the samples in one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement of the samples in one or more samples located at (i,j). In some examples, for each sample in sample (i,j) of the prediction block, Δv x and Δv y can be the same.

[0266] The processing circuit can decode (e.g., encode or decode) the current block based on the modified prediction block (e.g., one or more modified samples of the modified prediction block) (360). For example, for video decoding, the processing circuit (e.g., video decoder 300) can reconstruct the current block based on the modified prediction block (e.g., by adding one or more modified samples to the received residual values). For video encoding, the processing circuit (e.g., video encoder 200) can determine the residual value (e.g., of the residual block) between the current block and the modified prediction block (e.g., one or more modified samples of the modified prediction block), and signal information indicating the residual value.

[0267] The following describes a non - restrictive illustrative list of examples of the present disclosure.

[0268] Example 1. A method for decoding video data, the method comprising: performing sub - block - based affine motion compensation to obtain a prediction signal for a current block of video data; and refining the prediction signal by adding at least an offset to pixel positions within the prediction signal, wherein the value of the offset is derived based on values of a plurality of sub - block motion vectors (MVs) of sub - blocks of the current block, and the sub - blocks of the current block are not all in the same sub - block row or column of the current block.

[0269] Example 2. The method according to Example 1, further comprising: determining the value of the offset at position (m,n) in the prediction signal based on Δv x (m,n) and Δv y (m,n).

[0270] Example 3. The method according to Example 2, wherein determining the value of the offset at the position (m, n) of the prediction signal includes determining the value of the offset according to the following equation:

[0271] ΔI(m,n) = g x (m,n) * Δv x (m,n) + g y (m,n) * Δv y (m,n)

[0272] where g x (m,n) is the horizontal gradient of the prediction signal, and g y (m,n) is the vertical gradient of the prediction signal.

[0273] Example 4. The method according to any one of Examples 2 or 3, further comprising: determining Δv x (m,n) and Δv y (m,n) values based on the values of the multiple sub-block MVs.

[0274] Example 5. The method according to Example 4, wherein determining the values of Δv x (m,n) and Δv y (m,n) includes: determining Δv x (m,n) and Δv y (m,n) values based on the a parameter, b parameter, c parameter, and d parameter.

[0275] Example 6. The method according to Example 5, wherein deriving Δv x (m,n) and Δv y (m,n) includes: deriving Δv x (m,n) and Δv y (m,n) according to the following equation:

[0276]

[0277]

[0278] where sbW represents the width of the sub-block in the sub-blocks of the current block, sbH represents the height of the sub-block in the sub-blocks of the current block, and (m, n) represents the pixel position within the current block.

[0279] Example 7. The method according to any one of Examples 5 or 6, further comprising: deriving the a parameter, b parameter, c parameter, and d parameter according to the following equation, wherein the affine motion model is represented by 6 parameters:

[0280]

[0281]

[0282]

[0283]

[0284] Wherein, W represents the distance between the first MV of the affine motion model and the second MV of the affine motion model, and H represents the distance between the first MV of the affine motion model and the third MV of the affine motion model.

[0285] Example 8. The method according to any one of Examples 5-7, further comprising: deriving the a parameter, b parameter, c parameter, and d parameter according to the following equations, wherein the affine motion model is represented by 4 parameters:

[0286]

[0287]

[0288] b = -c

[0289] d = a

[0290] Wherein, W represents the distance between the first MV of the affine motion model and the second MV of the affine motion model.

[0291] Example 9. The method according to Example 7 or Example 8, wherein the first MV of the affine motion model is the second MV of the affine motion model is and the third MV of the affine motion model is

[0292] Example 10. The method according to any one of Examples 2-9, further comprising: clipping Δv x (m,n) and Δv y (m,n) to have an absolute value less than or equal to a predefined threshold.

[0293] Example 11. The method according to any one of Examples 2-10, further comprising: performing bidirectional optical flow (BDOF) refinement on the prediction signal for the current block.

[0294] Example 12. The method according to Example 11, further comprising: storing Δv x (m,n) and Δv y(m, n).

[0295] Example 13. The method according to any of Examples 11 or 12, wherein performing BDOF refinement includes performing gradient calculation.

[0296] Example 14. The method according to Example 13, wherein performing the gradient calculation for performing BDOF refinement uses the same process as calculating the horizontal gradient and / or the vertical gradient of the prediction signal.

[0297] Example 15. The method according to any of Examples 1 - 14, further comprising: clipping the refined prediction signal to have an intermediate precision same as that in non-affine motion compensation.

[0298] Example 16. The method according to Example 15, wherein when the number of bits for the intermediate precision is n, clipping the refined prediction signal includes: clipping the refined prediction signal to a range from –(1<<n) to (1<<n).

[0299] Example 17. The method according to any of Examples 1 - 16, wherein decoding includes decoding.

[0300] Example 18. The method according to any of Examples 1 - 17, wherein decoding includes encoding.

[0301] Example 19. An apparatus for decoding video data, the apparatus comprising one or more units for performing the method according to any of Examples 1 - 18.

[0302] Example 20. The apparatus according to Example 19, wherein the one or more units include one or more processors implemented in a circuit.

[0303] Example 21. The apparatus according to any of Examples 19 and 20, further comprising a memory for storing the video data.

[0304] Example 22. The apparatus according to any of Examples 19 - 21, further comprising a display configured to display the decoded video data.

[0305] Example 23. The apparatus according to any of Examples 19 - 22, wherein the apparatus includes one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0306] Example 24. The apparatus according to any of Examples 19 - 23, wherein the apparatus includes a video decoder.

[0307] Example 25. The apparatus according to any one of Examples 19-24, wherein the apparatus includes a video encoder.

[0308] Example 26. A computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to perform the method according to any one of Examples 1-18.

[0309] It should be recognized that, depending on the example, certain actions or events of any of the techniques described herein may be performed in a different order, may be added, combined, or entirely omitted (e.g., not all described actions or events are necessary for the practice of the technique). Additionally, in some examples, the actions or events may be performed, for example, by multithreaded processing, interrupt processing, or concurrently by multiple processors, rather than sequentially.

[0310] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium or a communication medium including any medium that facilitates transfer of a computer program from one place to another according to a communication protocol. In this manner, the computer-readable medium generally may correspond to: (1) a tangible computer-readable storage medium that is non-transitory; or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0311] By way of example, and not limitation, such a computer-readable storage medium can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if the instructions are transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transitory, tangible storage media. As used herein, disk and optical disks include compact disk (CD), laser disk, optical disk, digital versatile disk (DVD), floppy disk, and Blu-ray disk, where disks typically reproduce data magnetically, while optical disks utilize lasers to optically reproduce data. Combinations of the above should also be included within the scope of computer-readable media.

[0312] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, as used herein, the terms "processor" and "processing circuitry" refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functions described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Further, the techniques can be fully implemented in one or more circuits or logic elements.

[0313] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or a group of ICs (e.g., a chip set). Various components, modules, or units are described in the present disclosure to emphasize aspects of the functionality of the devices configured to perform the disclosed techniques, but are not necessarily required to be implemented by different hardware units. Rather, as described above, the various units can be combined in codec hardware units, or provided by a set of interoperating hardware units including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0314] Various examples have been described. These examples and other examples are within the scope of the appended claims.

Claims

1. A method for decoding video data, the method comprises: determining a prediction block for performing inter-frame prediction on a current block; determining a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block; rounding the horizontal displacement and the vertical displacement to the same precision level for different inter-frame prediction modes including an affine mode and a bi-directional optical flow (BDOF) mode; determining a first gradient of a first sample set based on the one or more samples of the prediction block; determining a second gradient of a second sample set based on the one or more samples of the prediction block; Determine one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement and the first gradient and the second gradient, wherein determining the one or more refinement offsets includes determining: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample among the one or more samples located at (i,j), Δv x (i,j) is the rounded horizontal displacement for the sample among the one or more samples located at (i,j), g y (i,j) is the second gradient for the sample among the one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for the sample among the one or more samples located at (i,j); performing gradient-based prediction refinement by modifying at least the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block; and reconstructing the current block based on the modified prediction block.

2. The method according to claim 1, further comprises: clipping the one or more refinement offsets, wherein modifying the one or more samples of the prediction block comprises: modifying the one or more samples of the prediction block based on the clipped one or more refinement offsets.

3. The method according to claim 1, further comprises: determining an inter-frame prediction mode for performing inter-frame prediction on the current block, wherein determining the horizontal displacement and the vertical displacement comprises: determining the horizontal displacement and the vertical displacement based on the determined inter-frame prediction mode.

4. The method according to claim 1, wherein, the precision level is 1 / 64.

5. The method according to claim 1, wherein, the prediction block is a first prediction block, the current block is a first current block, the horizontal displacement and the vertical displacement are a first horizontal displacement and a first vertical displacement, the one or more refinement offsets are a first one or more refinement offsets, the rounded horizontal displacement and the rounded vertical displacement are a first rounded horizontal displacement and a first rounded vertical displacement, and the modified prediction block is a first modified prediction block, the method further comprises: determining a second prediction block for performing inter-frame prediction on a second current block; determining a second horizontal displacement and a second vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block; rounding the second horizontal displacement and the second vertical displacement to the same precision level to which the first horizontal displacement and the first vertical displacement are rounded to generate a second rounded horizontal displacement and a second rounded vertical displacement; determining a second one or more refinement offsets based on the second rounded horizontal displacement and the second rounded vertical displacement; modifying the one or more samples of the second prediction block based on the determined second one or more refinement offsets to generate a second modified prediction block; and reconstructing the second current block based on the second modified prediction block.

6. A method for encoding video data, the method comprises: determining a prediction block for performing inter-frame prediction on a current block; Determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block; Round the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes including an affine mode and a bi-directional optical flow (BDOF) mode; Determine a first gradient based on a first sample set of the one or more samples of the prediction block; Determine a second gradient based on a second sample set of the one or more samples of the prediction block; Determine one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement, and the first gradient and the second gradient, wherein determining the one or more refinement offsets includes determining: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample among the one or more samples located at (i,j), Δv x (i,j) is the rounded horizontal displacement for the sample among the one or more samples located at (i,j), g y (i,j) is the second gradient for the sample among the one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for the sample among the one or more samples located at (i,j); Perform gradient-based prediction refinement by modifying at least the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block; Determine a residual value indicating a difference between the current block and the modified prediction block; and Signal information indicating the residual value.

7. The method according to claim 6, further comprising: Clamp the one or more refinement offsets, wherein modifying the one or more samples of the prediction block comprises: modifying the one or more samples of the prediction block based on the clamped one or more refinement offsets.

8. The method according to claim 6, further comprising: Determine an inter prediction mode for inter predicting the current block, wherein determining the horizontal displacement and the vertical displacement comprises: determining the horizontal displacement and the vertical displacement based on the determined inter prediction mode.

9. The method according to claim 6, wherein, the precision level is 1 / 64.

10. The method according to claim 6, wherein, the prediction block is a first prediction block, the current block is a first current block, the horizontal displacement and the vertical displacement are a first horizontal displacement and a first vertical displacement, the one or more refinement offsets are a first one or more refinement offsets, the rounded horizontal displacement and the rounded vertical displacement are a first rounded horizontal displacement and a first rounded vertical displacement, the modified prediction block is a first modified prediction block, and the residual value includes a first residual value, the method further comprises: Determine a second prediction block for inter predicting a second current block; Determine a second horizontal displacement and a second vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block; Round the second horizontal displacement and the second vertical displacement to the same precision level to which the first horizontal displacement and the first vertical displacement are rounded to generate a second rounded horizontal displacement and a second rounded vertical displacement; Determine a second one or more refinement offsets based on the second rounded horizontal displacement and the second rounded vertical displacement; Modify the one or more samples of the second prediction block based on the determined second one or more refinement offsets to generate a second modified prediction block; Determine a second residual value indicating a difference between the second current block and the second modified prediction block; and Signal information indicating the second residual value.

11. An apparatus for decoding video data, the apparatus comprising: A memory configured to store one or more samples of a prediction block; and A processing circuit configured to: Determine the prediction block for performing inter prediction on a current block; Determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of the one or more samples of the prediction block; Round the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes including an affine mode and a bi-directional optical flow (BDOF) mode; Determine a first gradient of a first sample set based on the one or more samples of the prediction block; Determine a second gradient of a second sample set based on the one or more samples of the prediction block; Perform gradient-based prediction refinement, wherein, to perform gradient-based prediction refinement, the processing circuitry is configured to: determine at least one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement and the first gradient and the second gradient, wherein, to determine the one or more refinement offsets, the processing circuitry is configured to determine: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample among the one or more samples located at (i,j), Δv x (i,j) is the rounded horizontal displacement for the sample among the one or more samples located at (i,j), g y (i,j) is the second gradient for the sample among the one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for the sample among the one or more samples located at (i,j); Modify the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block; and Decode the current block based on the modified prediction block.

12. The apparatus according to claim 11, wherein, To decode the current block, the processing circuit is configured to reconstruct the current block based on the modified prediction block.

13. The apparatus according to claim 11, wherein, To decode the current block, the processing circuit is configured to: Determine a residual value indicating a difference between the current block and the modified prediction block; and Signal information indicating the residual value.

14. The apparatus according to claim 11, wherein, The processing circuit is configured to: Limit the one or more refinement offsets, wherein, to modify the one or more samples of the prediction block, the processing circuit is configured to modify the one or more samples of the prediction block based on the limited one or more refinement offsets.

15. The apparatus according to claim 11, wherein, The processing circuit is configured to: Determine an inter prediction mode for performing inter prediction on the current block, wherein, to determine the horizontal displacement and the vertical displacement, the processing circuit is configured to determine the horizontal displacement and the vertical displacement based on the determined inter prediction mode.

16. The apparatus according to claim 11, wherein, The precision level is 1 / 64.

17. The apparatus according to claim 11, wherein, The prediction block is a first prediction block, the current block is a first current block, the horizontal displacement and the vertical displacement are a first horizontal displacement and a first vertical displacement, the one or more refinement offsets are a first one or more refinement offsets, the rounded horizontal displacement and the rounded vertical displacement are a first rounded horizontal displacement and a first rounded vertical displacement, and the modified prediction block is a first modified prediction block, and wherein, the processing circuit is configured to: Determine a second prediction block for performing inter prediction on a second current block; Determine a second horizontal displacement and a second vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block; Round the second horizontal displacement and the second vertical displacement to the same precision level to which the first horizontal displacement and the first vertical displacement are rounded, to generate a second rounded horizontal displacement and a second rounded vertical displacement; Determine second one or more refinement offsets based on the second rounded horizontal displacement and the second rounded vertical displacement; Modify the one or more samples of the second prediction block based on the determined second one or more refinement offsets to generate a second modified prediction block; and Decode the second current block based on the second modified prediction block.

18. The apparatus according to claim 11, further comprising a display configured to display decoded video data.

19. The apparatus according to claim 11, further comprising a camera configured to capture the video data to be encoded.

20. The apparatus according to claim 11, wherein, the apparatus comprises one or more of a camera, a computer, a wireless communication device, a broadcast receiver device, or a set-top box.

21. A non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform the following operations: Determine a prediction block for performing inter prediction on a current block; Determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block; Round the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes including an affine mode and a bi-directional optical flow (BDOF) mode; Determine a first gradient of a first sample set based on the one or more samples of the prediction block; Determine a second gradient of a second sample set based on the one or more samples of the prediction block; Determine one or more refinement offsets based on the rounded horizontal displacement and the rounded vertical displacement and the first gradient and the second gradient, wherein, The instructions that cause the one or more processors to determine the one or more refinement offsets include instructions that cause the one or more processors to perform the following operations: Determine: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample among the one or more samples located at (i,j), Δv x (i,j) is the rounded horizontal displacement for the sample among the one or more samples located at (i,j), g y (i,j) is the second gradient for the sample among the one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for the sample among the one or more samples located at (i,j); Perform gradient-based prediction refinement, wherein instructions that cause the one or more processors to perform gradient-based prediction refinement include instructions that cause the one or more processors to perform the following operations: at least modify the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block; and Decode the current block based on the modified prediction block.

22. An apparatus for decoding video data, the apparatus comprises: a unit for determining a prediction block for performing inter prediction on a current block; a unit for determining a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block; a unit for rounding the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes including an affine mode and a bi-directional optical flow (BDOF) mode; a unit for determining a first gradient of a first sample set based on the one or more samples of the prediction block; a unit for determining a second gradient of a second sample set based on the one or more samples of the prediction block; A unit for determining one or more refined offsets based on the rounded horizontal displacement and the rounded vertical displacement, and the first gradient and the second gradient, wherein the unit for determining the one or more refined offsets includes units for the following operations: determining: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample among the one or more samples located at (i,j), Δv x (i,j) is the rounded horizontal displacement for the sample among the one or more samples located at (i,j), g y (i,j) is the second gradient for the sample among the one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for the sample among the one or more samples located at (i,j); A unit for performing gradient-based prediction refinement, wherein the unit for performing gradient-based prediction refinement includes a unit for modifying one or more samples of the prediction block based on one or more determined refinement offsets to generate a modified prediction block; and A unit for decoding the current block based on the modified prediction block.