Gradient-Based Prediction Refinement for Video Coding

By rounding the displacement accuracy levels in video encoding and decoding technologies, it is unified in different inter-frame prediction modes, the problem of inconsistent displacement accuracy levels in the existing technology is solved, and the simplification of logic circuits and improvement of equipment performance is achieved.

CN113940079BActive Publication Date: 2025-06-17QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080023997.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-31
Filing Date
2020-04-01
Publication Date
2025-06-17
Estimated Expiration
2040-04-01

AI Technical Summary

Technical Problem

The existing video encoding and decoding technologies have problems with inconsistent displacement accuracy levels in inter-frame prediction, which leads to the need for different logic circuits to support different accuracy levels, increasing the size and power consumption of the device.

Method used

The same logic circuit is used for gradient-based prediction refinement by rounding the horizontal and vertical displacements so that they have the same level of accuracy in different inter prediction modes.

Benefits of technology

The logic circuit design of video encoder and decoder is simplified, reducing device size and power consumption, while improving overall operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113940079B_ABST
    Figure CN113940079B_ABST
Patent Text Reader

Abstract

The present disclosure describes gradient-based prediction refinement. A video decoder (e.g., a video encoder or a video decoder) determines one or more prediction blocks for inter prediction of a current block (e.g., based on one or more motion vectors for the current block). In gradient-based prediction refinement, the video decoder modifies one or more samples of the prediction block based on various factors such as displacement in the horizontal direction, horizontal gradient, displacement in the vertical direction, and vertical gradient. The present disclosure provides gradient-based prediction refinement in which the precision level of the displacement (e.g., at least one of horizontal displacement or vertical displacement) is uniform (e.g., the same) for different prediction modes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the benefit of priority of U.S. Application No. 16 / 836,013, filed Mar. 31, 2020, which claims the benefit of U.S. Provisional Application 62 / 827,677, filed Apr. 1, 2019, and U.S. Provisional Application 62 / 837,405, filed Apr. 23, 2019, the entire contents of each of which are hereby incorporated by reference. Technical Field

[0002] This disclosure relates to video encoding and video decoding. Background Art

[0003] Digital video capabilities can be incorporated into a wide variety of devices, including digital televisions, digital live systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones (so-called “smart phones”), video teleconferencing devices, video streaming devices, etc. Digital video devices implement video decoding techniques (such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 (Part 10, Advanced Video Coding (AVC)), ITU-T H.265 / High Efficiency Video Coding (HEVC)) and extensions of such standards). By implementing such video decoding techniques, video devices can send, receive, encode, decode, and / or store digital video information more efficiently.

[0004] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in a video sequence. For block-based video decoding, a video slice (e.g., a video picture or a portion of a video picture) can be divided into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction relative to reference samples in adjacent blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture can use spatial prediction relative to reference samples in adjacent blocks in the same picture, or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0005] Generally speaking, the present disclosure describes techniques for gradient-based prediction refinement. A video decoder (e.g., a video encoder or a video decoder) determines one or more prediction blocks for performing inter prediction on a current block (e.g., based on one or more motion vectors for the current block). In gradient-based prediction refinement, the video decoder modifies one or more samples of the prediction block based on various factors such as displacement in the horizontal direction, horizontal gradient, displacement in the vertical direction, and vertical gradient.

[0006] For example, a motion vector identifies a prediction block. The displacement in the horizontal direction (also referred to as horizontal displacement) refers to the change (e.g., difference) in the x coordinate of the motion vector, and the displacement in the vertical direction (also referred to as vertical displacement) refers to the change (e.g., difference) in the y coordinate. The horizontal gradient refers to the result of applying a filter to a first set of samples in the prediction block, and the vertical gradient refers to the result of applying a filter to a second set of samples in the prediction block.

[0007] The example techniques described in the present disclosure provide gradient-based prediction refinement, where the precision level of the displacement (e.g., at least one of the horizontal displacement or the vertical displacement) is uniform (e.g., the same) for different prediction modes. For example, for a first prediction mode (e.g., affine mode), the motion vector may be at a first precision level, and for a second prediction mode (e.g., bidirectional optical flow (BDOF)), the motion vector may be at a second precision level. Thus, the vertical and horizontal displacements of the motion vector for the affine mode and the motion vector for the BDOF may be different. In the present disclosure, the video decoder may be configured to round (e.g., round up or round down) the vertical and horizontal displacements for the motion vector such that the precision level of the displacement is the same regardless of the prediction mode (e.g., the vertical and horizontal displacements for the affine mode and the BDOF have the same precision level).

[0008] By rounding the precision level of the displacement, the example techniques can improve the overall operation of the video decoder. For example, gradient-based prediction refinement involves multiplication and shift operations. If the precision level of the displacement is different for different modes, different logic circuits may be required to support different precision levels (e.g., a logic circuit configured for one precision level may not be applicable to other precision levels). Since the precision level for the displacement is the same for different modes, the same logic circuit can be reused for blocks, resulting in a smaller overall logic circuit and reduced power consumption as there is no need to power unused logic circuits.

[0009] In some examples, the techniques for determining displacement can be based on information already available at the video decoder. For example, the way the video decoder determines a horizontal displacement or a vertical displacement can be based on information available to the video decoder for performing inter - prediction on a current block according to an inter - prediction mode. Additionally, there may be certain inter - prediction modes that are disabled for certain block types (e.g., based on size). In some examples, these inter - prediction modes that are disabled for certain block types may be enabled for those block types, but the example techniques described in this disclosure can be used to modify the predicted block for such blocks.

[0010] In one example, this disclosure describes a method for decoding video data, the method comprising: determining a predicted block for performing inter - prediction on a current block; determining a horizontal displacement and a vertical displacement for gradient - based prediction refinement of one or more samples of the predicted block; rounding the horizontal displacement and the vertical displacement to the same precision level for different inter - prediction modes; determining one or more refinement offsets based on the rounded horizontal displacement and vertical displacement; modifying the one or more samples of the predicted block based on the determined one or more refinement offsets to generate a modified predicted block; and reconstructing the current block based on the modified predicted block.

[0011] In one example, this disclosure describes a method for encoding video data, the method comprising: determining a predicted block for performing inter - prediction on a current block; determining a horizontal displacement and a vertical displacement for gradient - based prediction refinement of one or more samples of the predicted block; rounding the horizontal displacement and the vertical displacement to the same precision level for different inter - prediction modes; determining one or more refinement offsets based on the rounded horizontal displacement and vertical displacement; modifying the one or more samples of the predicted block based on the determined one or more refinement offsets to generate a modified predicted block; determining a residual value for indicating a difference between the current block and the modified predicted block; and signaling information for indicating the residual value.

[0012] In one example, the present disclosure describes an apparatus for decoding video data, the apparatus comprising: a memory configured to store one or more samples of a prediction block; and a processing circuit. The processing circuit is configured to: determine a prediction block for performing inter prediction on a current block; determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of the one or more samples of the prediction block; round the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes; determine one or more refinement offsets based on the rounded horizontal displacement and vertical displacement; modify the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block; and decode the current block based on the modified prediction block.

[0013] In one example, the present disclosure describes a computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to perform the following operations: determine a prediction block for performing inter prediction on a current block; determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of the one or more samples of the prediction block; round the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes; determine one or more refinement offsets based on the rounded horizontal displacement and vertical displacement; modify the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block; and decode the current block based on the modified prediction block.

[0014] In one example, the present disclosure describes an apparatus for decoding video data, the apparatus comprising: a unit configured to determine a prediction block for performing inter prediction on a current block; a unit configured to determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of the one or more samples of the prediction block; a unit configured to round the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes; a unit configured to determine one or more refinement offsets based on the rounded horizontal displacement and vertical displacement; a unit configured to modify the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block; and a unit configured to decode the current block based on the modified prediction block.

[0015] Details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1is a block diagram showing an example video encoding and decoding system that can execute the technology of the present disclosure.

[0017] Figure 2A and Figure 2B is a conceptual diagram showing an example quadtree binary tree (QTBT) structure and a corresponding coding tree unit (CTU).

[0018] Figure 3 is a block diagram showing an example video encoder that can execute the technology of the present disclosure.

[0019] Figure 4 is a block diagram showing an example video decoder that can execute the technology of the present disclosure.

[0020] Figure 5 is a flowchart showing an example method for decoding video data. Detailed Description

[0021] The present disclosure relates to gradient-based prediction refinement. In gradient-based prediction refinement, a video decoder (e.g., a video encoder or a video decoder) determines a prediction block for a current block based on a motion vector as part of inter prediction, and modifies (e.g., refines) the samples of the prediction block to generate modified prediction samples (e.g., refined prediction samples). The video encoder signals a residual value for indicating the difference between the modified prediction samples and the current block. The video decoder performs the same operations as those performed by the video encoder to modify the samples of the prediction block to generate modified prediction samples. The video decoder adds the residual value to the modified prediction samples to reconstruct the current block.

[0022] An example method for modifying the samples of a prediction block is to make the video decoder determine one or more refinement offsets and add the samples of the prediction block to the refinement offsets. An example method for generating the refinement offsets is based on gradients and motion vector displacements. The gradient can be determined according to a gradient filter applied to the samples of the prediction block.

[0023] Examples of motion vector displacements include a horizontal displacement of the motion vector and a vertical displacement of the motion vector. The horizontal displacement can be a value added to or subtracted from the x coordinate of the motion vector, and the vertical displacement can be a value added to or subtracted from the y coordinate of the motion vector. For example, the horizontal displacement can be referred to as Δvx x where vx x is the x coordinate of the motion vector, and the vertical displacement can be referred to as Δvy y where vy y is the y coordinate of the motion vector.

[0024] For different inter - frame prediction modes, the precision level of the motion vector of the current block may be different. For example, the coordinates of the motion vector (e.g., the x - coordinate or the y - coordinate) include an integer part and may include a fractional part. The fractional part is referred to as the sub - pel part of the motion vector because the integer part of the motion vector identifies the actual pixel in the reference picture that includes the predicted block, and the sub - pel part of the motion vector adjusts the motion vector to identify a position between pixels in the reference picture.

[0025] The precision level of the motion vector is based on the sub - pel part of the motion vector and indicates the granularity of the movement of the motion vector from the actual pixel in the reference picture. As an example, if the sub - pel part of the x - coordinate is 0.5, the motion vector is located in the middle of two horizontal pixels in the reference picture. If the sub - pel part of the x - coordinate is 0.25, the motion vector is located at a quarter of the distance between two horizontal pixels, and so on. In these examples, the precision level of the motion vector can be equal to the sub - pel part (e.g., precision levels of 0.5, 0.25, etc.).

[0026] In some examples, the precision levels of the horizontal displacement and the vertical displacement can be based on the precision level of the motion vector or the way the motion vector is generated. For example, in some examples (such as the merge mode in the form of an inter - frame prediction mode), the sub - pel parts of the x - coordinate and the y - coordinate of the motion vector can be the horizontal displacement and the vertical displacement, respectively. As another example (such as for the affine mode in the form of inter - frame prediction), the motion vector can be based on corner motion vectors, and the horizontal displacement and the vertical displacement can be determined based on the corner motion vectors.

[0027] For different inter - frame prediction modes, the precision levels of the horizontal displacement and the vertical displacement may be different. For example, for some inter - frame prediction modes, the horizontal displacement and the vertical displacement may be more precise (e.g., for the first prediction mode, the precision level is 1 / 128) compared to other inter - frame prediction modes (e.g., for the second prediction mode, the precision level is 1 / 16).

[0028] In implementation, a video decoder may need to include different logic circuits to handle different precision levels. Performing gradient - based prediction refinement involves multiplication, shift operations, addition, and other arithmetic operations. The logic circuits configured for one precision level of the horizontal displacement or the vertical displacement may not be able to handle the horizontal displacement and the vertical displacement of a higher precision level. Therefore, some video decoders include: a set of logic circuits for performing gradient - based prediction refinement for an inter - frame prediction mode in which the horizontal displacement and the vertical displacement have a first precision level, and a different set of logic circuits for performing gradient - based prediction refinement for another inter - frame prediction mode in which the horizontal displacement and the vertical displacement have a second precision level.

[0029] However, having different logic circuits for performing gradient-based prediction refinement for different inter-frame prediction modes results in additional logic circuits, which increases the size of the video decoder and consumes additional power. For example, if an inter-frame prediction is performed on a current block in a first mode, a first set of logic circuits for gradient-based prediction refinement is used. However, a second set of logic circuits for gradient-based prediction refinement for different inter-frame prediction modes is still receiving power.

[0030] This disclosure describes examples of techniques for rounding the precision levels of horizontal and vertical displacements to the same precision level for different inter-frame prediction modes. For example, a video decoder may round a first displacement (e.g., a first horizontal displacement or a first vertical displacement) having a first precision level for a first block inter-frame predicted in a first inter-frame prediction mode to a set precision level, and may round a second displacement (e.g., a second horizontal displacement or a second vertical displacement) having a second precision level for a second block inter-frame predicted in a second inter-frame prediction mode to the same set precision level. In other words, for different inter-frame prediction modes, the video decoder may round at least one of the horizontal and vertical displacements to the same precision level. As an example, the first inter-frame prediction mode may be an affine mode, and the second inter-frame prediction mode may be bidirectional optical flow (BDOF).

[0031] In this way, the same logic circuit can be used for gradient-based prediction refinement for different inter-frame prediction modes, rather than having different logic circuits for different inter-frame prediction modes. For example, the logic circuit of a video decoder may be configured to perform gradient-based prediction refinement on horizontal and vertical displacements having a set precision level. The video decoder may round the horizontal and vertical displacements such that the precision levels of the rounded horizontal and vertical displacements are equal to the set precision level, thereby allowing the same logic circuit to perform gradient-based prediction refinement for different inter-frame prediction modes.

[0032] Figure 1 is a block diagram of an example video encoding and decoding system 100 that may perform the techniques of this disclosure. Generally, the techniques of this disclosure relate to decoding (encoding and / or decoding) video data. Typically, video data includes any data for processing video. Thus, video data may include raw unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (e.g., signaling data).

[0033] As Figure 1As shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. Specifically, the source device 102 provides the video data to the destination device 116 via a computer-readable medium 110. The source device 102 and the destination device 116 can include any of a variety of devices, including desktop computers, notebook computers (i.e., laptop computers), tablet computers, set-top boxes, cellular phones such as smart phones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receiver devices, set-top boxes, etc. In some cases, the source device 102 and the destination device 116 can be equipped for wireless communication and can thus be referred to as wireless communication devices.

[0034] In Figure 1 the example of, the source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. The destination device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. According to the present disclosure, the video encoder 200 of the source device 102 and the video decoder 300 of the destination device 116 can be configured to apply techniques for gradient-based prediction refinement. Thus, the source device 102 represents an example of a video encoding device, while the destination device 116 represents an example of a video decoding device. In other examples, the source device and the destination device can include other components or arrangements. For example, the source device 102 can receive video data from an external video source such as an external camera. Similarly, the destination device 116 can interface with an external display device instead of including an integrated display device.

[0035] As Figure 1 shown, the system 100 is merely an example. Generally, any digital video encoding and / or decoding device can perform techniques for gradient-based prediction refinement. The source device 102 and the destination device 116 are merely examples of such decoding devices, where the source device 102 generates decoded video data for transmission to the destination device 116. The present disclosure refers to a "decoding" device as a device that performs decoding (e.g., encoding and / or decoding) of data. Thus, the video encoder 200 and the video decoder 300 represent examples of decoding devices (specifically, a video encoder and a video decoder), respectively. In some examples, the devices 102 and 116 can operate in a substantially symmetric manner such that each of the devices 102, 116 includes video encoding and decoding components. Thus, the system 100 can support one-way or two-way video transmission between the video devices 102, 116, e.g., for video streaming, video playback, video broadcast, or video telephony.

[0036] Typically, video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a sequential series of pictures (also referred to as “frames”) of the video data to video encoder 200, which encodes the data for the pictures. The video source 104 of source device 102 may include a video capture device, such as a camera, a video archive unit containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As another alternative, the video source 104 may generate computer graphics-based data as the source video, or generate a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 may encode the captured, pre-captured, or computer-generated video data. The video encoder 200 may reorder the pictures from the received order (sometimes referred to as “display order”) into a decoding order for decoding. The video encoder 200 may generate a bitstream including the encoded video data. Then, the source device 102 may output the encoded video data to a computer-readable medium 110 via output interface 108 for reception and / or retrieval by, for example, an input interface 122 of destination device 116.

[0037] The memories 106 of source device 102 and 120 of destination device 116 represent general memories. In some examples, the memories 106, 120 may store raw video data, e.g., raw video from video source 104 and raw decoded video data from video decoder 300. Additionally or alternatively, the memories 106, 120 may store software instructions that may be executed, for example, by video encoder 200 and video decoder 300, respectively. Although shown as separate from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memories for functionally similar or equivalent purposes. Further, the memories 106, 120 may store, for example, encoded video data output from video encoder 200 and input to video decoder 300. In some examples, portions of the memories 106, 120 may be allocated as one or more video buffers, e.g., to store raw decoded and / or encoded video data.

[0038] The computer-readable medium 110 can represent any type of medium or device capable of conveying the encoded video data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium such that the source device 102 can directly send the encoded video data to the destination device 116 in real time, for example, via a radio frequency network or a computer-based network. The output interface 108 can modulate the transmission signal including the encoded video data according to a communication standard such as a wireless communication protocol, and the input interface 122 can demodulate the received transmission signal according to a communication standard such as a wireless communication protocol. The communication medium can include any wireless or wired communication medium, for example, the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other device that can be useful for facilitating communication from the source device 102 to the destination device 116.

[0039] In some examples, the source device 102 can output the encoded data from the output interface 108 to the storage device 112. Similarly, the destination device 116 can access the encoded data from the storage device 112 via the input interface 122. The storage device 112 can include any one of a variety of distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data.

[0040] In some examples, the source device 102 may output the encoded video data to a file server 114 or to another intermediate storage device that may store the encoded video generated by the source device 102. The destination device 116 may access the stored video data from the file server 114 via streaming or downloading. The file server 114 may be any type of server device capable of storing the encoded video data and sending the encoded video data to the destination device 116. The file server 114 may represent a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a content delivery network device, or a Network Attached Storage (NAS) device. The destination device 116 may access the encoded video data from the file server 114 via any standard data connection, including an Internet connection. This may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server 114. The file server 114 and the input interface 122 may be configured to operate according to: a streaming protocol, a download transfer protocol, or a combination thereof.

[0041] The output interface 108 and the input interface 122 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any one of the various IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and the input interface 122 include wireless components, the output interface 108 and the input interface 122 may be configured to transmit data (such as encoded video data) according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), enhanced LTE, 5G, etc. In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 may be configured to transmit data (such as encoded video data) according to other wireless standards such as the IEEE 802.11 specifications, the IEEE 802.15 specifications (e.g., ZigBee TM )、Bluetooth TM standards, etc.). In some examples, the source device 102 and / or the destination device 116 may include respective System-on-Chip (SoC) devices. For example, the source device 102 may include an SoC device for performing the functions ascribed to the video encoder 200 and / or the output interface 108, and the destination device 116 may include an SoC device for performing the functions ascribed to the video decoder 300 and / or the input interface 122.

[0042] The techniques of the present disclosure may be applied to video coding to support any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions (such as HTTP-based Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0043] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a storage device 112, a file server 114, etc.). The computer-readable medium 110 of the encoded video bitstream may include signaling information defined by the video encoder 200 and also used by the video decoder 300, such as syntax elements having values for describing the characteristics and / or processing of video blocks or other coding units (e.g., slices, pictures, groups of pictures, sequences, etc.). The display device 118 displays the decoded pictures of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0044] Although not shown in Figure 1 In some examples, the video encoder 200 and the video decoder 300 may each be integrated with an audio encoder and / or an audio decoder and may include appropriate MUX-DEMUX units or other hardware and / or software to process a multiplexed stream including both audio and video in a common data stream. If applicable, the MUX-DEMUX unit may follow the ITU H.223 multiplexer protocol or other protocols (such as the User Datagram Protocol (UDP)).

[0045] Video encoder 200 and video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the techniques of the present disclosure. Each of video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, and any of the encoders or decoders may be integrated as part of a combined encoder / decoder (CODEC) in a corresponding device. Devices including video encoder 200 and / or video decoder 300 may include integrated circuits, microprocessors, and / or wireless communication devices (such as cellular phones).

[0046] Video encoder 200 and video decoder 300 may operate according to a video coding standard, such as ITU-T H.265 (also known as the High Efficiency Video Coding (HEVC) standard) or an extension thereof (such as multi-view and / or scalable video coding extensions). Alternatively, video encoder 200 and video decoder 300 may operate according to other proprietary or industry standards (such as the ITU-T H.266 standard, also known as Versatile Video Coding (VVC)). The latest draft of the VVC standard is described in the following document: Bross et al., "Versatile Video Coding (Draft 4)", Joint Video Team (JVT) of ITU-T SG 16 WP3 and ISO / IEC JTC 1 / SC 29 / WG 11, 13th meeting: Marrakech, Morocco, January 9 - 18, 2019, JVET-M1001-v5 (hereinafter referred to as "VVC Draft 4"). An updated draft of the VVC standard is described in the following document: Bross et al., "Versatile Video Coding (Draft 8)", Joint Video Team (JVT) of ITU-T SG 16 WP 3 and ISO / IEC JTC1 / SC 29 / WG 11, 17th meeting: Brussels, Belgium, January 7 - 17, 2020, JVET-Q2001-vD (hereinafter referred to as "VVC Draft 8"). However, the techniques of the present disclosure are not limited to any particular coding standard.

[0047] Generally, video encoder 200 and video decoder 300 may perform block-based coding of pictures. The term "block" generally refers to a structure that includes data to be processed (e.g., data to be encoded, decoded, or otherwise used during an encoding and / or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, video encoder 200 and video decoder 300 may code video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, instead of coding the red, green, and blue (RGB) data of the samples for a picture, video encoder 200 and video decoder 300 may code the luminance and chrominance components, where the chrominance components may include both a red hue and a blue hue chrominance component. In some examples, video encoder 200 converts the received RGB-formatted data to a YUV representation before encoding, and video decoder 300 converts the YUV representation to an RGB format. Alternatively, a preprocessing unit and a postprocessing unit (not shown) may perform these conversions.

[0048] Generally speaking, the present disclosure may relate to coding (e.g., encoding and decoding) of pictures to include a process of encoding or decoding data of a picture. Similarly, the present disclosure may relate to coding of blocks of a picture to include a process of encoding or decoding data for a block (e.g., prediction and / or residual coding). An encoded video bitstream generally includes a series of values for representing coding decisions (e.g., coding modes) and syntax elements that partition a picture into blocks. Thus, a reference to coding a picture or a block should generally be understood as coding the values of the syntax elements used to form the picture or the block.

[0049] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder (such as video encoder 200) divides a coding tree unit (CTU) into CUs according to a quadtree structure. That is, the video coder divides the CTU and CUs into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. A node without child nodes may be referred to as a "leaf node", and the CU of such a leaf node may include one or more PUs and / or one or more TUs. The video coder may further divide PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the division of TUs. In HEVC, a PU represents inter-prediction data, while a TU represents residual values. An intra-predicted CU includes intra-prediction information, such as an intra-mode indicator.

[0050] As another example, video encoder 200 and video decoder 300 may be configured to operate according to VVC. According to VVC, a video coder (such as video encoder 200) divides a picture into multiple coding tree units (CTUs). Video encoder 200 may divide a CTU according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure removes the concept of multiple partitioning types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level divided according to quadtree partitioning and a second level divided according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0051] In the MTT partitioning structure, quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning may be used to partition a block. Ternary tree partitioning is a partitioning in which a block is split into three sub-blocks. In some examples, ternary tree partitioning divides a block into three sub-blocks without dividing the original block through the center. The partitioning types in MTT (e.g., QT, BT, and TT) may be symmetric or asymmetric.

[0052] In some examples, video encoder 200 and video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance component and the chrominance components, while in other examples, video encoder 200 and video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for the two chrominance components (or two QTBT / MTT structures for the respective chrominance components).

[0053] Video encoder 200 and video decoder 300 may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures per HEVC. For purposes of explanation, a description of the techniques of the present disclosure is given with respect to QTBT partitioning. However, it should be understood that the techniques of the present disclosure may also be applied to video coders configured to use quadtree partitioning or other types of partitioning as well.

[0054] The present disclosure may interchangeably use "NxN" and "N by N" to refer to the sample dimensions of a block (such as a CU or other video block) in terms of vertical and horizontal dimensions. For example, 16x16 samples or 16 by 16 samples. Generally, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an NxN CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non - negative integer value. The samples in a CU can be arranged in rows and columns. Additionally, a CU does not necessarily need to have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU can include NxM samples, where M does not necessarily equal N.

[0055] Video encoder 200 encodes video data for the prediction and / or residual information and other information used for a CU. The prediction information indicates how the CU is to be predicted to form a prediction block for the CU. The residual information generally represents the sample - by - sample difference between the samples of the CU before encoding and the prediction block.

[0056] To predict a CU, video encoder 200 can generally form a prediction block for the CU through inter - frame prediction or intra - frame prediction. Inter - frame prediction generally refers to predicting the CU based on the data of previously decoded pictures, while intra - frame prediction generally refers to predicting the CU based on the previously decoded data of the same picture. To perform inter - frame prediction, video encoder 200 can use one or more motion vectors to generate the prediction block. Video encoder 200 can generally perform a motion search to identify, for example, a reference block that closely matches the CU in terms of the difference between the CU and the reference block. Video encoder 200 can use the following to calculate a difference metric to determine whether the reference block closely matches the current CU: sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations. In some examples, video encoder 200 can use uni - directional prediction or bi - directional prediction to predict the current CU.

[0057] Some examples of VVC also provide an affine motion compensation mode, which can be considered an inter - frame prediction mode. In the affine motion compensation mode, video encoder 200 can determine two or more motion vectors for representing non - translational motion (such as zooming in or out, rotation, perspective motion, or other irregular motion types).

[0058] To perform intra prediction, video encoder 200 may select an intra prediction mode to generate a prediction block. Some examples of VVC provide sixty-seven intra prediction modes, including various directional modes, as well as planar mode and DC mode. Generally, video encoder 200 selects an intra prediction mode that describes the neighboring samples of the current block from which the samples of the current block (e.g., the block of a CU) are to be predicted. Assuming that video encoder 200 decodes CTUs and CUs in raster scan order (from left to right, top to bottom), such samples can generally be above, top-left, or to the left of the current block in the same picture as the current block.

[0059] Video encoder 200 encodes data representing the prediction mode for the current block. For example, for an inter prediction mode, video encoder 200 may encode data representing which one of the various available inter prediction modes is used, as well as the motion information for the corresponding mode. For unidirectional or bidirectional inter prediction, for example, video encoder 200 may use advanced motion vector prediction (AMVP) or merge mode to encode the motion vectors. Video encoder 200 may use a similar mode to encode the motion vectors for the affine motion compensation mode.

[0060] After prediction such as intra prediction or inter prediction of a block, video encoder 200 may calculate a residual value for the block. The residual value (such as a residual block) represents the sample-by-sample difference between the block and the prediction block for the block, which is formed using the corresponding prediction mode. Video encoder 200 may apply one or more transforms to the residual block to produce transformed data in the transform domain rather than in the sample domain. For example, video encoder 200 may apply a discrete cosine transform (DCT), integer transform, wavelet transform, or conceptually similar transform to the residual video data. Additionally, video encoder 200 may apply a secondary transform after the first transform, such as a mode-dependent non-separable secondary transform (MDNSST), signal-dependent transform, Karhunen-Loeve transform (KLT), etc. Video encoder 200 produces transform coefficients after applying one or more transforms.

[0061] As described above, after any transform to produce transform coefficients, video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to the process in which transform coefficients are quantized to potentially reduce the amount of data used to represent the coefficients, thereby providing further compression. By performing the quantization process, video encoder 200 may reduce the bit depth associated with some or all of the coefficients. For example, video encoder 200 may round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, video encoder 200 may perform a bitwise right shift on the value to be quantized.

[0062] After quantization, video encoder 200 may scan the transform coefficients to produce a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place higher energy (and thus lower frequency) coefficients at the front of the vector and lower energy (and thus higher frequency) transform coefficients at the back of the vector. In some examples, video encoder 200 may use a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy code the quantized transform coefficients of the vector. In other examples, video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 200 may entropy code the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). Video encoder 200 may also entropy code the values of syntax elements used to describe metadata associated with the encoded video data for use by video decoder 300 when decoding the video data.

[0063] To perform CABAC, video encoder 200 may assign a context within a context model to the symbol to be sent. The context may relate, for example, to whether adjacent values of the symbol are zero values. Probability determination may be based on the context assigned to the symbol.

[0064] Video encoder 200 may also generate syntax data (such as block-based syntax data, picture-based syntax data, and sequence-based syntax data) to video decoder 300, or other syntax data (such as sequence parameter set (SPS), picture parameter set (PPS), or video parameter set (VPS)) in, for example, a picture header, a block header, a slice header. Similarly, video decoder 300 may decode such syntax data to determine how to decode the corresponding video data.

[0065] In this way, video encoder 200 can generate a bitstream that includes encoded video data, e.g., syntax elements for describing the partitioning of a picture into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Eventually, video decoder 300 can receive the bitstream and decode the encoded video data.

[0066] Generally, video decoder 300 performs a process opposite to that performed by video encoder 200 to decode the encoded video data of the bitstream. For example, video decoder 300 can use CABAC to decode the values of the syntax elements for the bitstream in a manner that is substantially similar to, but opposite to, the CABAC encoding process of video encoder 200. The syntax elements can define partitioning information for partitioning a picture into CTUs and for further partitioning each CTU according to a corresponding partitioning structure (such as a QTBT structure) to define the CUs of the CTU. The syntax elements can also define prediction and residual information for blocks (e.g., CUs) of the video data.

[0067] The residual information can be represented by, e.g., quantized transform coefficients. Video decoder 300 can inverse-quantize and inverse-transform the quantized transform coefficients of a block to reproduce the residual block for the block. Video decoder 300 uses the signalized prediction mode (intra prediction or inter prediction) and associated prediction information (e.g., motion information for inter prediction) to form a prediction block for the block. Video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. Video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the boundaries of the blocks.

[0068] In general, the present disclosure may relate to “signaling” certain information, such as syntax elements. The term “signaling” can generally refer to the conveyance of values for syntax elements and / or other data used to decode the encoded video data. That is, video encoder 200 can signal the values for the syntax elements in the bitstream. Generally, signaling refers to generating values in the bitstream. As described above, source device 102 can transmit the bitstream to destination device 116 substantially in real time or not in real time (such as may occur when storing the syntax elements to storage device 112 for later retrieval by destination device 116).

[0069] According to the technology of the present disclosure, the video encoder 200 and the video decoder 300 can be configured to perform gradient-based prediction refinement. As described above, as part of performing inter prediction on a current block, the video encoder 200 and the video decoder 300 can determine one or more prediction blocks for the current block (e.g., based on one or more motion vectors). In gradient-based prediction refinement, the video encoder 200 and the video decoder 300 modify one or more samples of the prediction block (e.g., including all samples).

[0070] For example, in gradient-based prediction refinement, the inter prediction sample at position (i, j) (e.g., a sample of the prediction block) is refined by an offset ΔI(i, j), which is derived from the displacement in the horizontal direction, the horizontal gradient, the displacement in the vertical direction, and the vertical gradient at position (i, j). In one example, the prediction refinement is described as: ΔI(i, j) = g x (i, j) * Δv x (i, j) + g y (i, j) * Δv y (i, j), where g x (i, j) is the horizontal gradient, and g y (i, j) is the vertical gradient, and Δv x (i, j) is the displacement in the horizontal direction, and Δv y (i, j) is the displacement in the vertical direction.

[0071] The gradient of an image is a measure of the directional change of intensity or color in the image. For example, the gradient value is the rate of change of color or intensity in the direction of the maximum change in color or intensity based on adjacent samples. As an example, if the rate of change is relatively high, the gradient value is larger than when the rate of change is relatively low.

[0072] In addition, the prediction block for the current block can be a reference picture different from the current picture including the current block. The video encoder 200 and the video decoder 300 can determine the offset (e.g., ΔI(i, j)) based on the sample values in the reference picture (e.g., the gradient is determined based on the sample values in the reference picture). In some examples, the values used to determine the gradient can be values within the prediction block itself or values generated based on the values in the prediction block (e.g., values such as interpolation, rounding, etc. generated according to the values in the prediction block). In addition, in some examples, the values used to determine the gradient can be outside the prediction block and within the reference picture, or are generated (e.g., interpolated, rounded, etc.) according to samples outside the prediction block and within the reference picture.

[0073] However, in some examples, the video encoder 200 and the video decoder 300 may determine an offset based on sample values in the current picture. In some examples (such as intra-block copy), the current picture and the reference picture are the same picture.

[0074] The displacement (e.g., vertical and / or horizontal displacement) may be determined based on the inter-frame prediction mode. In some examples, the displacement is determined based on motion parameters. As described in more detail, for the decoder-side motion refinement mode, the displacement may be based on samples in the reference picture. For other inter-frame prediction modes, the displacement may not be based on samples in the reference picture, but the example techniques are not limited thereto, and samples in the reference picture may be used to determine the displacement. There may be various ways to determine the vertical and / or horizontal displacement, and the techniques are not limited to a specific way of determining the vertical and / or horizontal displacement.

[0075] Example methods for performing gradient calculation are described below. For example, for a gradient filter, in one example, the Sobel filter may be used for gradient calculation. The gradient is calculated as follows: g x (i,j) = I(i+1,j-1) - I(i-1,j-1) + 2*I(i+1,j) - 2*I(i-1,j) + I(i+1,j+1) - I(i-1,j+1) and g y (i,j) = I(i-1,j+1) - I(i-1,j-1) + 2*I(i,j+1) - 2*I(i,j-1) + I(i+1,j+1) - I(i+1,j-1).

[0076] In some examples, a [1, 0, -1] filter is applied. The gradient can be calculated as follows: g x (i,j) = I(i+1,j) - I(i-1,j) and g y (i,j) = I(i,j+1) - I(i,j-1). In some examples, other gradient filters, such as the Canny filter, may be applied.

[0077] For gradient normalization, the calculated gradient may be normalized before being used for refinement offset derivation (e.g., before calculating ΔI), or normalization may be performed after refinement offset derivation. A rounding process may be applied during normalization. For example, if a [1, 0, -1] filter is applied, normalization is performed by adding one to the input value and then right-shifting by one. If the input is scaled by a power of two, normalization is performed by adding 1<<N and then right-shifting by (N+1).

[0078] For the gradient at the boundary, the gradient at the boundary of the prediction block can be calculated by expanding the prediction block by S / 2 at each boundary, where S is the filtering step used for gradient calculation. In one example, the extended prediction samples are generated by using the same motion vector as the prediction block used for inter-frame prediction (motion compensation). In some examples, the extended prediction samples are generated by using the same motion vector but a shorter filter for the interpolation process in motion compensation. In some examples, the extended prediction samples are generated by using the rounded motion vector for integer motion compensation. In some examples, the extended prediction samples are generated by padding, where the padding is performed by copying the boundary samples. In some examples, if the prediction block is generated by sub-block-based motion compensation, the extended prediction samples are generated by using the motion vector of the nearest sub-block. In some examples, if the prediction block is generated by sub-block-based motion compensation, the extended prediction samples are generated by using a representative motion vector. In one example, the representative motion vector can be the motion vector at the center of the prediction block. In one example, the representative motion vector can be derived by averaging the motion vectors of the boundary sub-blocks.

[0079] Sub-block-based gradient derivation can be applied to facilitate parallel processing or pipeline-friendly design in hardware. The width and height of the sub-blocks (denoted as sbW and sbH) can be determined as follows: sbW = min(blkW, SB_WIDTH) and sbH = min(blkH, SB_HEIGHT). In this equation, blkW and blkH are the width and height of the prediction block respectively. SB_WIDTH and SB_HEIGHT are two predetermined variables. In one example, both SB_WIDTH and SB_HEIGHT are equal to 16.

[0080] In some examples, for the horizontal displacement and the vertical displacement, the horizontal displacement Δv x (i,j) and the vertical displacement Δv y (i,j) used in the refinement derivation can be determined according to the inter-frame prediction mode. However, the example techniques are not limited to determining the horizontal and vertical displacements based on the inter-frame prediction mode.

[0081] For small-block-size inter-frame modes (e.g., small-sized blocks predicted inter-frame), to reduce the memory bandwidth in the worst case, the inter-frame prediction mode for small blocks can be disabled or restricted. For example, disabling the inter-frame prediction for 4x4 or smaller blocks can disable the bi-directional prediction for 4x8, 8x4, 4x16, and 16x4. Due to the interpolation process for those small blocks, the memory bandwidth may increase. Integer motion compensation without interpolation can still be applied to those small blocks without increasing the memory bandwidth in the worst case.

[0082] In one or more example techniques, inter - frame prediction may be enabled for some or all of those tiles, but with integer motion compensation and gradient - based prediction refinement. First, the motion vectors are rounded to integer motion vectors for motion compensation. Then, the remainder of the rounding (i.e., the sub - pixel element part of the motion vector) is used as Δv x (i,j) and Δv y (i,j) for gradient - based prediction refinement. For example, if the motion vector for a tile is (2.25, 5.75), the integer motion vector for motion compensation is (2, 6), and the horizontal displacement (e.g., Δv x (i,j)) is 0.25, and the vertical displacement (e.g., Δv y (i,j)) is 0.75. In this example, the precision level of the horizontal and vertical displacements is 0.25 (or 1 / 4). For example, the horizontal and vertical displacements can be incremented in steps of 0.25.

[0083] In some examples, for the inter - frame mode of tile size, gradient - based prediction refinement may be available, but only when inter - frame prediction is performed on small - sized blocks in the merge mode. Examples of the merge mode are described below. In some examples, for small - sized inter - frame modes, gradient - based prediction refinement may be disabled for blocks with integer motion patterns. In the integer motion pattern, one or more motion vectors (e.g., the signaled motion vectors) are integers. In some examples, even for larger - sized blocks, if inter - frame prediction is performed on the block in the integer motion pattern, gradient - based prediction refinement may also be disabled for such blocks.

[0084] For the normal merge mode (which is an example of an inter - frame prediction mode) where the motion information is derived from spatially or temporally adjacent decoded blocks, Δv x (i,j) and Δv y (i,j) can be the remainder of the motion vector rounding process (e.g., similar to the above example of the motion vector (2.25, 5.75)). In one example, the temporal motion vector predictor is derived by scaling the motion vectors in the temporal motion buffer according to the picture order count that is different between the current picture and the reference picture. A rounding process may be performed to round the scaled motion vectors to a certain precision. The remainder can be used as Δv x (i,j) and Δv y (i,j). The precision of the remainder (i.e., the precision level of the horizontal and vertical displacements) can be predefined and can be higher than the precision of the motion vector prediction. For example, if the motion vector precision is 1 / 16, the remainder precision is 1 / (16 * MaxBlkSize), where MaxBlkSize is the maximum block size. In other words, for the horizontal and vertical displacements (e.g., Δvx and Δv y ) has a precision level of 1 / (16 * MaxBlkSize).

[0085] For merge using motion vector differences (MMVD) mode, which is an example of an inter prediction mode, the motion vector difference is signaled together with the merge index to represent the motion information. In some techniques, the motion vector difference (e.g., the difference between the actual motion vector and the motion vector predictor) has the same precision as the motion vector. In one or more examples described in the present disclosure, the motion vector difference may be allowed to have a higher precision. First, the signaled motion vector difference is rounded to the motion vector precision, and the motion vector indicated by the merge index is added to generate the final motion vector for motion compensation. In one or more examples, the remainder after rounding (e.g., the difference between the rounded value of the motion vector difference and the original value of the motion vector difference) can be used as the horizontal displacement and vertical displacement for gradient-based prediction refinement (e.g., used as Δv x (i, j) and Δv y (i, j)). In some examples, Δv x (i, j) and Δv y (i, j) can be signaled as candidates for the motion vector difference.

[0086] For the decoder-side motion vector refinement mode, motion compensation using the original motion vector is performed to generate the original bi-directional prediction block, and the difference between the prediction in list 0 and the prediction in list 1 is calculated, denoted as DistOrig. List 0 refers to the first reference picture list (RefPicList0), which includes the reference picture list that may potentially be used for inter prediction. List 1 refers to the second reference picture list (RefPicList1), which includes the reference picture list that may potentially be used for inter prediction. Then the motion vectors at list 0 and list 1 are rounded to the nearest integer positions. That is, the motion vector referring to the picture in list 0 is rounded to the nearest integer position, and the motion vector referring to the picture in list 1 is rounded to the nearest integer position. A search algorithm is used to search within the range of integer displacements to find the displacement pair with the minimum distortion DistNew between the block of the picture identified in the list 0 prediction and the block of the picture identified in list 1 using the new integer motion vector for motion compensation. If DistNew is less than DistOrig, the new integer motion vector is fed into the bi-directional optical flow (BDOF) to derive Δv x (i, j) and Δv y (i, j) for prediction refinement at both list 0 and list 1 predictions. Otherwise, BDOF is performed on the original list 0 and list 1 predictions for prediction refinement.

[0087] For the affine mode, the motion field can be derived for each pixel (e.g., the motion vector can be determined on a per-pixel basis). However, a 4x4 motion field is used for affine motion compensation to reduce complexity and memory bandwidth. For example, as an example, instead of determining the motion vector on a per-pixel basis, the motion vector is determined for sub-blocks, where as an example, a sub-block is 4x4. Some other sub-block sizes can also be used, such as 4x2, 2x4, or 2x2. In one or more examples, gradient-based prediction refinement can be used to improve affine motion compensation. The gradient of the block can be calculated as described above. Given the affine motion model: where a, b, c, d, e, and f are values determined by the video encoder 200 and the video decoder 300 based on the control point motion vectors and the length and width of the block, as several examples. In some examples, the values for a, b, c, d, e, and f can be signaled.

[0088] Some example methods for determining a, b, c, d, e, and f are described below. In a video decoder (e.g., the video encoder 200 or the video decoder 300), in the affine mode, the picture is divided into sub-blocks for block-based decoding. The affine motion model for the block can also be described by three motion vectors (MVs) at three different positions that are not collinear and These three positions are typically referred to as control points, and these three motion vectors are referred to as control point motion vectors (CPMVs). In the case where these three control points are at the three corners of the block, the affine motion can be described as

[0089]

[0090] where blkW and blkH are the width and height of the block.

[0091] For the affine mode, the video encoder 200 and the video decoder 300 can use the representative coordinates of the sub-blocks (e.g., the center position of the sub-block) to determine the motion vector for each sub-block. In one example, the block is divided into non-overlapping sub-blocks. If the block width is blkW, the block height is blkH, the sub-block width is sbW, and the sub-block height is sbH, then there are blkH / sbH rows of sub-blocks and blkW / sbW sub-blocks in each row. For the six-parameter affine motion model, the motion vector for the sub-block (referred to as the sub-block MV) at the i-th row (0 <= i < blkW / sbW) and the j-th (0 <= j < blkH / sbH) column is derived as follows:

[0092]

[0093] According to the above equations, variables a, b, c, d, e, and f can be defined as follows:

[0094]

[0095]

[0096]

[0097]

[0098] e = v 0x

[0099] f = v 0y

[0100] For an affine mode (which is an example of an inter prediction mode), the video encoder 200 and the video decoder 300 can determine a displacement (e.g., a horizontal displacement or a vertical displacement) by at least one of the following methods. The following are examples and should not be considered limiting. There may be other ways in which the video encoder 200 and the video decoder 300 can determine the displacement (e.g., a horizontal displacement or a vertical displacement) for the affine mode.

[0101] For 4x4 sub-block based affine motion compensation, for 2x2 based displacement derivation, the displacement is the same for each 2x2 sub-block. In each 4x4 sub-block, Δv(i,j) for the four 2x2 sub-blocks within the 4x4 is calculated as follows:

[0102] Upper left 2x2: Upper right 2x2: Lower left 2x2: Lower right 2x2:

[0103] For 1x1 displacement derivation, the displacement is derived for each sample. The coordinates of the upper left sample in the 4x4 can be (0,0), in which case Δv(i,j) is derived as follows:

[0104] In some examples, the divide-by-2 implemented as a right shift operation can be moved to the refinement offset calculation. For example, instead of performing the divide-by-2 operation when deriving the horizontal displacement and the vertical displacement (e.g., Δv x and Δv y ), the video encoder 200 and the video decoder 300 can perform the divide-by-2 operation as part of determining ΔI (e.g., the refinement offset).

[0105] For 4x2 sub-block based affine motion compensation, the motion field for motion vector storage remains 4x4; however, the affine motion compensation is 4x2. The motion vector (MV) for a 4x4 sub-block can be (v x , v y ). In this case, the MV for motion compensation of the left 4x2 is (v x - a, v y - c), and the MV for motion compensation of the right 4x2 is (v x + a, v y + c).

[0106] For 2x2 displacement derivation, in 2x2 displacement derivation, the displacement in each 2x2 sub-block is the same. In each 4x2 sub-block, calculate Δv(i,j) for the 2 2x2 sub-blocks within the 4x4 as follows:

[0107] Upper 2x2:

[0108] Lower 2x2:

[0109] For 1x1 displacement derivation, derive the displacement for each sample. Let the coordinates of the top-left sample in the 4x2 be (0,0), and derive Δv(i,j) as follows:

[0110] The divide-by-2 that can be implemented as a right shift operation can be moved to the refinement offset calculation. For example, instead of performing the divide-by-2 operation when deriving the horizontal displacement and vertical displacement (e.g., Δv x and Δv y ), the video encoder 200 and the video decoder 300 can perform the divide-by-2 operation as part of determining ΔI (e.g., the refinement offset).

[0111] For 2x4 sub-block based affine motion compensation, the motion field for motion vector storage remains 4x4; however, the affine motion compensation is 2x4. The MV for a 4x4 sub-block can be (v x , v y ). In this case, the MV for motion compensation of the left 4x2 is (v x - b, v y - d), and the MV for motion compensation of the right 4x2 is (v x + b, v y + d).

[0112] For 2x2 displacement derivation, the displacement in each 2x2 sub-block is the same. In each 2x4 sub-block, calculate Δv(i,j) for the 2 2x2 sub-blocks within the 2x4 as follows:

[0113] Left 2x2:

[0114] Right 2x2:

[0115] For 1x1 displacement derivation, in the 1x1-based displacement derivation, the displacement is derived for each sample. The coordinates of the top-right sample in 2x4 can be (0, 0), and in this case, Δv(i,j) is derived as follows:

[0116] The divide-by-2 that can be implemented as a right-shift operation can be moved to the refinement offset calculation. For example, instead of performing the divide-by-2 operation when deriving the horizontal and vertical displacements (e.g., Δv x and Δv y ), the video encoder 200 and the video decoder 300 can perform the divide-by-2 operation as part of determining ΔI (e.g., the refinement offset).

[0117] The following describes the precision of the displacement and the gradient. In some examples, the same precision for the horizontal and vertical displacements can be used in all modes. The precision can be predefined or signaled in the high-level syntax. Thus, if the horizontal and vertical displacements are derived from different modes with different precisions, the horizontal and vertical displacements are rounded to the predefined precision. Examples of predefined precision are: 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, 1 / 128, etc.

[0118] As described above, the precision (also referred to as the precision level) can indicate the horizontal and vertical displacements (e.g., Δv x and Δv y) of accuracy, where the horizontal displacement and the vertical displacement can be determined using one or more of the above examples or using some other techniques. Generally, the accuracy level is defined in decimal (e.g., 0.25, 0.125, 0.0625, 0.03125, 0.015625, 0.0078125, etc.) or fraction (e.g., 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, 1 / 128, etc.). For example, for a 1 / 4 accuracy level, the horizontal displacement or the vertical displacement can be represented using an increment of 0.25 (e.g., 0.25, 0.5, or 0.75). For a 1 / 8 accuracy level, the horizontal displacement and the vertical displacement can be represented using an increment of 0.125 (e.g., 0.125, 0.25, 0.325, 0.5, 0.625, 0.75, or 0.825). It can be seen that the lower the value of the accuracy level (e.g., 1 / 8 is less than 1 / 4), the more granular the increments are, and the more precise the values that can be given (e.g., for a 1 / 4 accuracy level, the displacement is rounded to the nearest quarter, but for a 1 / 8 accuracy level, the displacement is rounded to the nearest eighth).

[0119] Because for different inter-frame prediction modes, the horizontal displacement and the vertical displacement can have different accuracy levels, the video encoder 200 and the video decoder 300 can be configured to include different logic circuits to perform gradient-based prediction refinement for different inter-frame prediction modes. As described above, to perform gradient-based prediction refinement, the video encoder 200 and the video decoder 300 can perform the following operations: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x and g y are the first gradient based on the first sample set in the samples of the prediction block and the second gradient based on the second sample set in the samples of the prediction block, respectively, and Δv x and Δv y are the horizontal displacement and the vertical displacement, respectively. It can be seen that for gradient-based prediction refinement, the video encoder 200 and the video decoder 300 may need to perform multiplication and addition operations, as well as use a memory to store the temporary results used in the calculations.

[0120] However, the ability of the logic circuits (e.g., multiplier circuits, adder circuits, memory registers) to perform mathematical operations may be limited to the accuracy level that the logic circuits are configured for. For example, a logic circuit configured for a first accuracy level may not be able to perform the operations required for gradient prediction refinement in which the horizontal displacement or the vertical displacement is at a more precise second accuracy level.

[0121] Thus, some techniques utilize different sets of logic circuits configured for different levels of precision to perform gradient-based prediction refinement for different inter-frame prediction modes. For example, a first set of logic circuits may be configured to perform gradient-based prediction refinement for an inter-frame prediction mode in which the horizontal and / or vertical displacement is 0.25, and a second set of logic circuits may be configured to perform gradient-based prediction refinement for an inter-frame prediction mode in which the horizontal and / or vertical displacement is 0.125. Having these different sets of circuits increases the overall size of video encoder 200 and video decoder 300, and potentially wastes power.

[0122] In some examples described in this disclosure, the same gradient calculation process can be used for all inter-frame prediction modes. In other words, the same logic circuits can be used to perform gradient-based prediction refinement for different inter-frame prediction modes. For example, the precision of the gradient can be kept the same for prediction refinement in all inter-frame prediction modes. In some examples, for the precision of displacement and gradient, the example techniques can ensure that the same (or unified) prediction refinement process can be applied to different modes, and the same prediction refinement module can be applied to different modes.

[0123] As an example, video encoder 200 and video decoder 300 can be configured to round at least one of the horizontal displacement and the vertical displacement to the same level of precision for different inter-frame prediction modes. For example, if the level of precision to which the horizontal displacement and the vertical displacement are rounded is 0.015625 (1 / 64), then if for an inter-frame prediction mode, the level of precision of the horizontal and / or vertical displacement is 1 / 4, the level of precision of the horizontal and / or vertical displacement is rounded to 1 / 64. If the level of precision of the horizontal and / or vertical displacement is 1 / 128, the level of precision of the horizontal and / or vertical displacement is rounded to 1 / 64.

[0124] In this way, the logic circuits for gradient-based prediction refinement can be reused for different inter-frame prediction modes. For example, in the above example, video encoder 200 and video decoder 300 can include logic circuits for a level of precision of 0.125, and this logic circuit can be reused for different inter-frame prediction modes because the level of precision of the horizontal and / or vertical displacement is rounded to 0.125.

[0125] In some examples, when rounding is not performed according to the techniques described in this disclosure, if the logic circuit is designed to have a relatively high level of precision, the logic circuit for multiply-accumulate type operations can be reused (e.g., the logic circuit designed for a specific precision level of multiplication can handle multiplication operations for values of a lower precision level). However, for shift operations, the logic circuit designed for a specific precision may not be able to handle shift operations for values of a lower precision level. Using the example techniques described in this disclosure and the described rounding techniques, it may be possible to reuse the logic circuit, including for shift operations for different inter-frame prediction modes.

[0126] In one example, the prediction refinement offset is derived as follows:

[0127] ΔI(i,j) = (g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j) + offset) >> shift

[0128] In the above equation, the offset is equal to 1 << (shift - 1), and the shift is determined by the predefined precision of the displacement and the gradient and is fixed for different modes. In some examples, the offset is equal to 0.

[0129] In some examples, the mode can include one or more of the modes described above with respect to horizontal displacement and vertical displacement, such as the inter-frame mode of the block size, the normal merge mode, the merge using the motion vector difference, the decoder-side motion vector refinement mode, and the affine mode. The mode can also include bidirectional optical flow (BDOF).

[0130] For each prediction direction, there can be separate refinements. For example, in the case of bidirectional prediction, the prediction refinement can be performed separately for each prediction direction. The result of the refinement can be clipped to a certain range to ensure the same bit width as the prediction without refinement. For example, the refinement result is clipped to a 16-bit range. As described above, the example techniques can also be applied to BDOF, where it is assumed that the displacements in two different directions are on the same motion trajectory.

[0131] The following describes the N-bit (e.g., 16-bit) multiplication constraint. To reduce the complexity of gradient-based prediction refinement, the multiplication can be kept within N bits (e.g., 16 bits). In this example, the gradient and the displacement should be able to be represented by no more than 16 bits. If not, in this example, the gradient or the displacement is quantized to within 16 bits. For example, a right shift can be applied to maintain a 16-bit representation.

[0132] The clipping of the refinement offset ΔI(i,j) and the refinement result is described below. The refinement offset ΔI(i,j) is clipped to a certain range. In one example, the range is determined by the range of the original prediction signal. The range of ΔI(i,j) can be the same as the range of the original prediction signal, or the range can be a scaled range. The scaling can be 1 / 2, 1 / 4, 1 / 8, etc. The refinement result is clipped to have the same range as the original prediction signal (e.g., the range of samples in the prediction block). The equation for performing the clipping is: pbSamples[x][y] = Clip3(0, (2 BitDepth ) - 1, (predSamplesL0[x + 1][y + 1] + offset4 + predSamplesL1[x + 1][y + 1] + bdofOffset) >> shift4)

[0133] In this way, the video encoder 200 and the video decoder 300 can be configured to determine a prediction block for performing inter - frame prediction on a current block. For example, the video encoder 200 and the video decoder 300 can determine a motion vector or a block vector (e.g., for the intra - block copy mode) pointing to the prediction block.

[0134] The video encoder 200 and the video decoder 300 can determine at least one of a horizontal displacement or a vertical displacement for gradient - based prediction refinement of one or more samples of the prediction block. An example of the horizontal displacement is Δv x , and an example of the vertical displacement is Δv y . In some examples, the video encoder 200 and the video decoder 300 can determine at least one of a horizontal displacement or a vertical displacement based on an inter - frame prediction mode for gradient - based prediction refinement of one or more samples of the prediction block (e.g., using the example techniques for the affine mode described above to determine Δv x and Δv y , or using the example techniques for the merge mode described above to determine Δv x and Δv y , as two examples).

[0135] According to one or more examples, the video encoder 200 and the video decoder 300 may round at least one of a horizontal displacement and a vertical displacement to the same precision level for different inter prediction modes. Examples of different inter prediction modes include an affine mode and BDOF. For example, the precision level of a first horizontal displacement or vertical displacement for performing gradient-based prediction refinement on a first block inter predicted in a first inter prediction mode may be at a first precision level, and the precision level of a second horizontal displacement or vertical displacement for performing gradient-based prediction refinement on a second block inter predicted in a second inter prediction mode may be at a second precision level. The video encoder 200 and the video decoder 300 may be configured to round the first precision level for the first horizontal displacement or vertical displacement to a precision level, and round the second precision level for the first horizontal displacement or vertical displacement to the same precision level.

[0136] In some examples, the precision level may be predefined (e.g., pre-stored on the video encoder 200 and the video decoder 300) or may be signaled (e.g., defined by the video encoder 200 and signaled to the video decoder 300). In some examples, the precision level may be 1 / 64.

[0137] The video encoder 200 and the video decoder 300 may be configured to determine one or more refinement offsets based on at least one of the rounded horizontal displacement or vertical displacement. For example, the video encoder 200 and the video decoder 300 may use at least one of the respective rounded horizontal displacement or vertical displacement to determine ΔI(i,j) for each sample of a prediction block. That is, the video encoder 200 and the video decoder 300 may determine a refinement offset for each sample of a prediction block. In some examples, the video encoder 200 and the video decoder 300 may utilize the rounded horizontal displacement and vertical displacement to determine the refinement offset (e.g., ΔI).

[0138] As described above, to perform gradient-based prediction refinement, the video encoder 200 and the video decoder 300 may determine a first gradient based on a first set of samples among one or more samples of a prediction block (e.g., determine g x (i,j), where the first set of samples is the samples for determining g x (i,j)), and determine a second gradient based on a second set of samples among one or more samples of the prediction block (e.g., determine g y (i,j), where the second set of samples is the samples for determining g y (i,j)). The video encoder 200 and the video decoder 300 may determine the refinement offset based on the rounded horizontal displacement and vertical displacement and the first gradient and the second gradient.

[0139] The video encoder 200 and the video decoder 300 may modify one or more samples of a prediction block based on the determined one or more refinement offsets to generate a modified prediction block (e.g., form one or more modified samples of the modified prediction block). For example, the video encoder 200 and the video decoder 300 may add or subtract ΔI(i,j) from I(i,j), where I(i,j) refers to the sample in the prediction block located at position I(i,j). In some examples, the video encoder 200 and the video decoder 300 may clip one or more refinement offsets (e.g., clip ΔI(i,j)). The video encoder 200 and the video decoder 300 may modify one or more samples of the prediction block based on the clipped one or more refinement offsets.

[0140] For encoding, the video encoder 200 may determine a residual value (e.g., of a residual block) indicating the difference between the current block and the modified prediction block (e.g., based on the modified samples of the modified prediction block), and signal information indicating the residual value. For decoding, the video decoder 300 may receive the information indicating the residual value, and reconstruct the current block based on the modified prediction block (e.g., the modified samples of the modified prediction block) and the residual value (e.g., by adding the residual value to the modified samples).

[0141] Figure 2A and 2B FIG. 10 is a conceptual diagram showing an example quadtree binary tree (QTBT) structure 130 and a corresponding coding tree unit (CTU) 132. Solid lines represent quadtree splits, and dashed lines indicate binary tree splits. In each split (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which split type is used (i.e., horizontal or vertical), where, in this example, 0 indicates a horizontal split and 1 indicates a vertical split. For a quadtree split, since a quadtree node splits a block horizontally and vertically into 4 sub-blocks of equal size, there is no need to indicate the split type. Thus, the video encoder 200 may encode the following, and the video decoder 300 may decode the following: syntax elements (such as split information) for the region tree level (i.e., solid lines) of the QTBT structure 130, and syntax elements (such as split information) for the prediction tree level (i.e., dashed lines) of the QTBT structure 130. The video encoder 200 may encode video data (such as prediction and transform data) for a CU represented by a terminal leaf node of the QTBT structure 130, and the video decoder 300 may decode the video data.

[0142] Generally Figure 2BThe CTU 132 can be associated with parameters for defining the size of blocks corresponding to nodes at the first and second levels of the QTBT structure 130. These parameters can include the CTU size (representing the size of the CTU 132 in the sample), the minimum quadtree size (MinQTSize, which represents the minimum allowable quadtree leaf node size), the maximum binary tree size (MaxBTSize, which represents the maximum allowable binary tree root node size), the maximum binary tree depth (MaxBTDepth, which represents the maximum allowable binary tree depth), and the minimum binary tree size (MinBTSize, which represents the minimum allowable binary tree leaf node size).

[0143] The root node of the QTBT structure corresponding to the CTU can have four child nodes at the first level of the QTBT structure, and each child node can be divided according to quadtree partitioning. That is, the nodes at the first level are leaf nodes (without child nodes) or have four child nodes. An example of the QTBT structure 130 represents such a node as including a parent node and child nodes with solid branches. If the nodes at the first level are not larger than the maximum allowable binary tree root node size (MaxBTSize), they can be further divided by the corresponding binary tree. The binary tree splitting of a node can be iterated until the nodes resulting from the splitting reach the minimum allowable binary tree leaf node size (MinBTSize) or the maximum allowable binary tree depth (MaxBTDepth). An example of the QTBT structure 130 represents such a node as having dashed lines for the branches. The binary tree leaf nodes are called coding units (CUs), which are used for prediction (e.g., intra-picture or inter-picture prediction) and transformation without any further division. As discussed above, the CU can also be referred to as a "video block" or a "block".

[0144] In an example of the QTBT partitioning structure, the CTU size is set to 128x128 (luma samples and two corresponding 64x64 chroma samples), MinQTSize is set to 16x16, MaxBTSize is set to 64x64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. First, quadtree partitioning is applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If the leaf quadtree node is 128x128, since this size exceeds MaxBTSize (i.e., 64x64 in this example), the leaf quadtree node will not be further split by the binary tree. Otherwise, the leaf quadtree node will be further partitioned by the binary tree. Thus, the quadtree leaf node is also the root node for the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (4 in this example), further splitting is not allowed. When a binary tree node has a width equal to MinBTSize (4 in this example), it means that further horizontal splitting is not allowed. Similarly, a binary tree node with a height equal to MinBTSize means that further vertical splitting is not allowed for that binary tree node. As described above, the leaf nodes of the binary tree are called CUs and are further processed according to prediction and transformation without further partitioning.

[0145] Figure 3 is a block diagram showing an example video encoder 200 in which the techniques of the present disclosure may be implemented. Figure 3 is provided for purposes of explanation and should not be construed as limiting the techniques generally illustrated and described in the present disclosure. For purposes of explanation, the present disclosure describes the video encoder 200 in the context of video coding standards such as the HEVC video coding standard and the H.266 video coding standard currently under development. However, the techniques of the present disclosure are not limited to these video coding standards and are generally applicable to video encoding and decoding.

[0146] In Figure 3In the example, video encoder 200 includes video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, decoded picture buffer (DPB) 218, and entropy encoding unit 220. Any one or all of video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy encoding unit 220 may be implemented in one or more processors or in processing circuitry. Additionally, video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0147] Video data memory 230 may store video data to be encoded by components of video encoder 200. Video encoder 200 may receive the video data stored in video data memory 230 from, for example, video source 104 ( Figure 1 ). DPB 218 may act as a reference picture memory that stores reference video data for use in predicting subsequent video data by video encoder 200. Video data memory 230 and DPB 218 may be formed of any one of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. Video data memory 230 and DPB 218 may be provided by the same memory device or by separate memory devices. In various examples, video data memory 230 may be on-chip (as shown) with other components of video encoder 200 or off-chip relative to those components.

[0148] In the present disclosure, a reference to video data memory 230 should not be construed as limited to memory internal to video encoder 200 (unless so specifically described) or limited to memory external to video encoder 200 (unless so specifically described). Rather, a reference to video data memory 230 should be understood as a reference memory that stores video data received by video encoder 200 for encoding (e.g., video data for a current block to be encoded). Figure 1 Memory 106 may also provide temporary storage of outputs from the various units of video encoder 200.

[0149] illustrates Figure 3The various units to assist in understanding the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide specific functions and are preset with respect to the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions with respect to the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by fixed-function circuits are generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more units can be integrated circuits.

[0150] The video encoder 200 can include an arithmetic logic unit (ALU), a basic function unit (EFU), digital circuits, analog circuits, and / or programmable cores formed by programmable circuits. In examples where software executed by programmable circuits is used to perform the operations of the video encoder 200, the memory 106( Figure 1 ) can store the object code of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 can store such instructions.

[0151] The video data memory 230 is configured to store the received video data. The video encoder 200 can retrieve pictures of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 can be the original video data to be encoded.

[0152] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, an intra prediction unit 226, and a gradient-based prediction refinement (GBPR) unit 227. The mode selection unit 202 can include additional functional units to perform video prediction according to other prediction modes. As an example, the mode selection unit 202 can include a palette unit, a block copy unit (which can be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0153] Although the GBPR unit 227 is shown as being separate from the motion estimation unit 222 and the motion compensation unit 224, in some examples, the GBPR unit 227 can be part of the motion estimation unit 222 and / or the motion compensation unit 224. The GBPR unit 227 is shown as being separate from the motion estimation unit 222 and the motion compensation unit 224 for ease of understanding and should not be considered limiting.

[0154] The mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and the rate-distortion values obtained for such combinations. The encoding parameters can include dividing a CTU into CUs, the prediction mode for a CU, the transform type for the residual values of a CU, the quantization parameter for the residual values of a CU, etc. The mode selection unit 202 can ultimately select the combination of encoding parameters that has a better rate-distortion value than other tested combinations.

[0155] The video encoder 200 can divide a picture retrieved from the video data memory 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can divide the CTUs of the picture according to a tree structure (such as the QTBT structure or the quadtree structure of HEVC described above). As described above, the video encoder 200 can form one or more CUs by dividing the CTUs according to a tree structure. Such CUs can generally also be referred to as "video blocks" or "blocks".

[0156] Generally, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, the intra prediction unit 226, and the GBPR unit 227) to generate a prediction block for the current block (e.g., the current CU, or the overlapping part of the PU and TU in HEVC). To perform inter prediction on the current block, the motion estimation unit 222 can perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in the DPB 218). Specifically, the motion estimation unit 222 can calculate, for example, values representing how closely a potential reference block will resemble the current block based on the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), etc. The motion estimation unit 222 can generally use the sample-by-sample differences between the current block and the considered reference block to perform these calculations. The motion estimation unit 222 can identify the reference block with the lowest value obtained from these calculations, which indicates the reference block that most closely matches the current block.

[0157] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to the position of a current block in the current picture. The motion estimation unit 222 may then provide the motion vectors to the motion compensation unit 224. For example, for uni-directional inter prediction, the motion estimation unit 222 may provide a single motion vector, and for bi-directional inter prediction, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then use the motion vectors to generate a predicted block. For example, the motion compensation unit 224 may use the motion vectors to retrieve data of the reference block. As another example, if the motion vectors have fractional sample precision, the motion compensation unit 224 may interpolate values for the predicted block according to one or more interpolation filters. Further, for bi-directional inter prediction, the motion compensation unit 224 may retrieve data of two reference blocks identified by the respective motion vectors and combine the retrieved data, for example, by per-sample averaging or weighted averaging.

[0158] As another example, for intra prediction or intra prediction coding, the intra prediction unit 226 may generate a predicted block according to samples adjacent to the current block. For example, for a directional mode, the intra prediction unit 226 may generally mathematically combine values of adjacent samples and fill the calculated values across the current block in the defined direction to produce the predicted block. As another example, for the DC mode, the intra prediction unit 226 may calculate an average value of adjacent samples of the current block and generate a predicted block to include the obtained average value for each sample of the predicted block.

[0159] The GBPR unit 227 may be configured to perform example techniques for gradient-based prediction refinement described in the present disclosure. For example, the GBPR unit 227 together with the motion compensation unit 224 may determine a predicted block for inter prediction of the current block (e.g., based on the motion vectors determined by the motion estimation unit 222). The GBPR unit 227 may determine a horizontal displacement and a vertical displacement (e.g., Δv x and Δv y ) for gradient-based prediction refinement of one or more samples of the predicted block. As an example, the GBPR unit 227 may determine an inter prediction mode for inter prediction of the current block based on a determination made by the mode selection unit 202. In some examples, the GBPR unit 227 may determine the horizontal displacement and the vertical displacement based on the determined inter prediction mode.

[0160] The GBPR unit 227 may round the horizontal displacement and the vertical displacement to the same precision level for different inter-frame prediction modes. For example, the current block may be a first current block, the prediction block may be a first prediction block, the horizontal displacement and the vertical displacement may be a first horizontal displacement and a first vertical displacement, and the rounded horizontal displacement and vertical displacement may be a first rounded horizontal displacement and a first rounded vertical displacement. In some examples, the GBPR unit 227 may determine a second prediction block for inter-frame prediction of a second current block, and determine a second horizontal displacement and a second vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block. The GBPR unit 227 may round the second horizontal displacement and the second vertical displacement to the same precision level to which the first horizontal displacement and the first vertical displacement are rounded to generate a second rounded horizontal displacement and a second rounded vertical displacement.

[0161] In some cases, the inter-frame prediction mode used for inter-frame prediction of the first current block and the inter-frame prediction mode used for the second current block may be different. For example, the first mode in different inter-frame prediction modes is an affine mode, and the second mode in different inter-frame prediction modes is a bi-directional optical flow (BDOF) mode.

[0162] The precision level to which the horizontal displacement and the vertical displacement are rounded may be predefined and stored for use by the GBPR unit 227, or the GBPR unit 227 may determine the precision level, and the video encoder 200 may signal the precision level. As an example, the precision level is 1 / 64.

[0163] The GBPR unit 227 may determine one or more refinement offsets based on the rounded horizontal displacement and the vertical displacement. For example, the GBPR unit 227 may determine a first gradient based on a first set of samples among one or more samples of the prediction block (e.g., using the samples of the prediction block described above to determine g x (i,j)), and determine a second gradient based on a second set of samples among one or more samples of the prediction block (e.g., using the samples of the prediction block described above to determine g y (i,j)). The GBPR unit 227 may determine one or more refinement offsets based on the rounded horizontal displacement and the vertical displacement and the first gradient and the second gradient. In some examples, if the value of one or more refinement offsets is too high (e.g., greater than a threshold), the GBPR unit 227 may clip one or more refinement offsets.

[0164] The GBPR unit 227 may modify one or more samples of the prediction block based on the determined one or more refinement offsets or the clipped one or more refinement offsets to generate a modified prediction block (e.g., one or more modified samples for forming the modified prediction block). For example, the GBPR unit 227 may determine: gx (i, j) * Δv x (i, j) + g y (i, j) * Δv y (i, j), where g x (i, j) is the first gradient for the sample located at (i, j) in one or more samples, Δv x (i, j) is the rounded horizontal displacement for the sample located at (i, j) in one or more samples, g y (i, j) is the second gradient for the sample located at (i, j) in one or more samples, and Δv y (i, j) is the rounded vertical displacement for the sample located at (i, j) in one or more samples. In some examples, for each sample (i, j) in the sample (i, j) of the prediction block, Δv x and Δv y may be the same.

[0165] The resulting modified samples can form a prediction block (e.g., a modified prediction block) in gradient-based prediction refinement. That is, the modified prediction block is used as a prediction block in gradient-based prediction refinement. The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives the original unencoded version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the per-sample difference between the current block and the prediction block. The resulting per-sample difference defines the residual block for the current block. In some examples, the residual generation unit 204 may also determine the differences between the sample values in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, one or more subtractor circuits that perform binary subtraction can be used to form the residual generation unit 204.

[0166] In examples where the mode selection unit 202 divides a CU into PUs, each PU can be associated with a luminance prediction unit and a corresponding chrominance prediction unit. The video encoder 200 and the video decoder 300 can support PUs of various sizes. As noted above, the size of a CU can refer to the size of the luminance coding block of the CU, and the size of a PU can refer to the size of the luminance prediction unit of the PU. Assuming that the size of a particular CU is 2Nx2N, the video encoder 200 can support PU sizes of 2Nx2N or NxN for intra prediction, and PU sizes of 2Nx2N, 2NxN, Nx2N, NxN, or similar symmetric PU sizes for inter prediction. The video encoder 200 and the video decoder 300 can also support asymmetric partitioning for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter prediction.

[0167] In an example where the mode selection unit 202 does not further divide a CU into PUs, each CU may be associated with a luminance decoding block and a corresponding chrominance decoding block. As described above, the size of a CU may refer to the size of the luminance decoding block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0168] For other video decoding techniques (to name a few examples, such as intra-block copy mode decoding, affine mode decoding, and linear model (LM) mode decoding), the mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the decoding technique. In some examples (such as palette mode decoding), the mode selection unit 202 may not generate a prediction block, but instead generates a syntax element for indicating the manner in which a block is to be reconstructed based on a selected palette. In such a mode, the mode selection unit 202 may provide these syntax elements to the entropy encoding unit 220 for encoding.

[0169] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.

[0170] The transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, the transform processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transform processing unit 206 may perform multiple transforms on the residual block, for example, a primary transform and a secondary transform (such as a rotation transform). In some examples, the transform processing unit 206 does not apply a transform to the residual block.

[0171] The quantization unit 208 may quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. The quantization unit 208 may quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) may adjust the degree of quantization applied to the coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may cause information loss, and thus, the quantized transform coefficients may have lower precision than the original transform coefficients produced by the transform processing unit 206.

[0172] The inverse quantization unit 210 and the inverse transform processing unit 212 may apply inverse quantization and inverse transform to the quantized transform coefficient block respectively to reconstruct the residual block from the transform coefficient block. The reconstruction unit 214 may generate a reconstructed block corresponding to the current block (although potentially with a certain degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add the samples of the reconstructed residual block to the corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.

[0173] The filter unit 216 may perform one or more filter operations on the reconstructed block. For example, the filter unit 216 may perform a deblocking operation to reduce the block effect artifacts along the edges of the CU. In some examples, the operation of the filter unit 216 may be skipped.

[0174] The video encoder 200 stores the reconstructed block in the DPB 218. For example, in an example where the operation of the filter unit 216 is not required, the reconstruction unit 214 may store the reconstructed block into the DPB 218. In an example where the operation of the filter unit 216 is required, the filter unit 216 may store the filtered reconstructed block into the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may retrieve the reference picture formed from the reconstructed (and potentially filtered) block from the DPB 218 to perform inter prediction on the blocks of the subsequent encoded pictures. Additionally, the intra prediction unit 226 may use the reconstructed blocks of the current picture in the DPB 218 to perform intra prediction on other blocks in the current picture.

[0175] Generally, the entropy coding unit 220 may perform entropy coding on the syntax elements received from other functional components of the video encoder 200. For example, the entropy coding unit 220 may perform entropy coding on the quantized transform coefficient block from the quantization unit 208. As another example, the entropy coding unit 220 may perform entropy coding on the prediction syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from the mode selection unit 202. The entropy coding unit 220 may perform one or more entropy coding operations on the syntax elements as another example of the video data to generate the entropy coded data. For example, the entropy coding unit 220 may perform context adaptive variable length coding (CAVLC) operation, CABAC operation, variable-to-variable (V2V) length coding operation, syntax-based context adaptive binary arithmetic coding (SBAC) operation, probability interval partitioning entropy (PIPE) coding operation, exponential Golomb coding operation, or another type of entropy coding operation on the data. In some examples, the entropy coding unit 220 may operate in a bypass mode where the syntax elements are not entropy coded.

[0176] Video encoder 200 may output a bitstream, which includes entropy-coded syntax elements required to reconstruct blocks of a slice or picture. Specifically, entropy coding unit 220 may output a bitstream.

[0177] The above operations are described with respect to blocks. Such a description should be understood as operations for a luma decoding block and / or a chroma decoding block. As described above, in some examples, the luma decoding block and the chroma decoding block are the luma component and the chroma component of a CU. In some examples, the luma decoding block and the chroma decoding block are the luma component and the chroma component of a PU.

[0178] In some examples, it is not necessary to repeat the operations performed on the luma coding block for the chroma decoding block. As an example, it is not necessary to repeat the operations for identifying the motion vector (MV) and the reference picture for the luma decoding block to identify the MV and the reference picture for the chroma block. Instead, the MV for the luma decoding block may be scaled to determine the MV for the chroma block, and the reference picture may be the same. As another example, for the luma decoding block and the chroma decoding block, the intra prediction process may be the same.

[0179] Figure 4 is a block diagram illustrating an example video decoder 300 that may implement the techniques of the present disclosure. Figure 4 is provided for explanatory purposes and does not limit the techniques generally illustrated and described in the present disclosure. For explanatory purposes, the present disclosure describes video decoder 300 in accordance with the techniques of VVC and HEVC. However, the techniques of the present disclosure may be performed by a video decoding device configured for other video coding standards.

[0180] In Figure 4 example, video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 134. Any one or all of CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 134 may be implemented in one or more processors or in processing circuitry. Additionally, video decoder 300 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0181] The prediction processing unit 304 includes a motion compensation unit 316, an intra prediction unit 318, and a gradient-based prediction refinement (GBPR) unit 319. The prediction processing unit 304 may include an addition unit that performs prediction according to other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, a block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0182] Although the GBPR unit 319 is shown as separate from the motion compensation unit 316, in some examples, the GBPR unit 319 may be part of the motion compensation unit 316. The GBPR unit 319 is shown as separate from the motion compensation unit 316 for ease of understanding and should not be considered limiting.

[0183] The CPB memory 320 may store video data to be decoded by components of the video decoder 300, such as an encoded video bitstream. For example, the video data stored in the CPB memory 320 may be obtained from a computer-readable medium 110 ( Figure 1 ). The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. In addition, the CPB memory 320 may store video data other than the syntax elements of the decoded pictures, such as temporary data for representing the outputs of the respective units from the video decoder 300. The DPB 314 generally stores decoded pictures, the video decoder 300 may output decoded pictures, and / or use the decoded pictures as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 may be formed of any of various memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The CPB memory 320 and the DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300 or off-chip relative to those components.

[0184] Additionally or alternatively, in some examples, the video decoder 300 may obtain from a memory 120 ( Figure 1)Retrieve the decoded video data. That is, the memory 120 can use the CPB memory 320 to store data as discussed above. Similarly, when some or all of the functions of the video decoder 300 are implemented with software to be executed by the processing circuitry of the video decoder 300, the memory 120 can store the instructions to be executed by the video decoder 300.

[0185] illustrates Figure 4 the various units illustrated in to assist in understanding the operations performed by the video decoder 300. These units can be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Similar to Figure 3 , fixed-function circuitry refers to circuitry that provides a specific function and is preset with respect to the operations that can be performed. Programmable circuitry refers to circuitry that can be programmed to perform various tasks and provides flexible functionality in terms of the operations that can be performed. For example, programmable circuitry can execute software or firmware that causes the programmable circuitry to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuitry can execute software instructions (e.g., to receive parameters or output parameters), but the type of operations performed by the fixed-function circuitry is generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more of the units can be an integrated circuit.

[0186] The video decoder 300 can include an ALU, an EFU, digital circuitry, analog circuitry, and / or programmable cores formed by programmable circuitry. In examples where the operations of the video decoder 300 are performed by software executed on programmable circuitry, on-chip or off-chip memory can store the instructions (e.g., object code) of the software received and executed by the video decoder 300.

[0187] The entropy decoding unit 302 can receive the encoded video data from the CPB and perform entropy decoding on the video data to reproduce the syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 can generate the decoded video data based on the syntax elements extracted from the bitstream.

[0188] Generally, the video decoder 300 reconstructs pictures on a block-by-block basis. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the “current block”).

[0189] The entropy decoding unit 302 can perform entropy decoding on the syntax elements of the quantized transform coefficients that define the quantized transform coefficient block, as well as transform information such as quantization parameter (QP) and / or transform mode indication. The inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization to be applied by the inverse quantization unit 306. The inverse quantization unit 306 can, for example, perform a bitwise left shift operation to inverse-quantize the quantized transform coefficients. The inverse quantization unit 306 can thus form a transform coefficient block including the transform coefficients.

[0190] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the coefficient block.

[0191] In addition, the prediction processing unit 304 generates a prediction block based on the prediction information syntax elements entropy decoded by the entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is inter-frame predicted, the motion compensation unit 316 can generate a prediction block. In this case, the prediction information syntax element can indicate the reference picture in the DPB 314 from which the reference block is to be retrieved, and the motion vector for identifying the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can generally perform the inter-frame prediction process in a manner substantially similar to that described with respect to the motion compensation unit 224 ( Figure 3 ).

[0192] As another example, if the prediction information syntax element indicates that the current block is intra-frame predicted, the intra-frame prediction unit 318 can generate a prediction block according to the intra-frame prediction mode indicated by the prediction information syntax element. Again, the intra-frame prediction unit 318 can generally perform the intra-frame prediction process in a manner substantially similar to that described with respect to the intra-frame prediction unit 226 ( Figure 3 ). The intra-frame prediction unit 318 can retrieve the data of the neighboring samples of the current block from the DPB 314.

[0193] As another example, if the prediction information syntax element indicates that gradient-based prediction refinement is enabled, the GBPR unit 319 can modify the samples of the prediction block to generate a modified prediction block (e.g., generate modified samples for forming the modified prediction block), and the modified prediction block is used to reconstruct the current block.

[0194] The GBPR unit 319 may be configured to perform the example techniques for gradient-based prediction refinement described in this disclosure. For example, the GBPR unit 319 together with the motion compensation unit 316 may determine a prediction block for performing inter prediction on a current block (e.g., based on a motion vector determined by the prediction processing unit 304). The GBPR unit 319 may determine a horizontal displacement and a vertical displacement (e.g., Δv x and Δv y ) for gradient-based prediction refinement of one or more samples of the prediction block. As an example, the GBPR unit 319 may determine an inter prediction mode for performing inter prediction on the current block based on a prediction information syntax element. In some examples, the GBPR unit 319 may determine the horizontal displacement and the vertical displacement based on the determined inter prediction mode.

[0195] The GBPR unit 319 may round the horizontal displacement and the vertical displacement to the same precision level for different inter prediction modes. For example, the current block may be a first current block, the prediction block may be a first prediction block, the horizontal displacement and the vertical displacement may be a first horizontal displacement and a first vertical displacement, and the rounded horizontal displacement and the vertical displacement may be a first rounded horizontal displacement and a first rounded vertical displacement. In some examples, the GBPR unit 319 may determine a second prediction block for performing inter prediction on a second current block and determine a second horizontal displacement and a second vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block. The GBPR unit 319 may round the second horizontal displacement and the second vertical displacement to the same precision level to which the first horizontal displacement and the first vertical displacement are rounded to generate a second rounded horizontal displacement and a second rounded vertical displacement.

[0196] In some cases, the inter prediction mode for performing inter prediction on the first current block and the inter prediction mode for the second current block may be different. For example, the first mode among different inter prediction modes is an affine mode, and the second mode among different inter prediction modes is a bi-directional optical flow (BDOF) mode.

[0197] The precision level to which the horizontal displacement and the vertical displacement are rounded may be predefined and stored for use by the GBPR unit 319, or the GBPR unit 319 may receive information for indicating the precision level in the signaled information (e.g., the precision level may be signaled). As an example, the precision level is 1 / 64.

[0198] The GBPR unit 319 may determine one or more refinement offsets based on the rounded horizontal displacement and the vertical displacement. For example, the GBPR unit 319 may determine a first gradient based on a first set of samples among one or more samples of the prediction block (e.g., determine g using the samples of the prediction block described above)x (i,j)), and determine a second gradient based on a second set of samples in one or more samples of the prediction block (e.g., determine g using the samples of the prediction block as described above y (i,j)). The GBPR unit 319 can determine one or more refinement offsets based on the rounded horizontal displacement and vertical displacement and the first gradient and the second gradient. In some examples, if the value of one or more refinement offsets is too high (e.g., greater than a threshold), the GBPR unit 319 can clip one or more refinement offsets.

[0199] The GBPR unit 319 can modify one or more samples of the prediction block based on the determined one or more refinement offsets or the clipped one or more refinement offsets to generate a modified prediction block (e.g., one or more modified samples for forming the modified prediction block). For example, the GBPR unit 319 can determine: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample located at (i,j) in one or more samples, Δv x (i,j) is the rounded horizontal displacement for the sample located at (i,j) in one or more samples, g y (i,j) is the second gradient for the sample located at (i,j) in one or more samples, and Δv y (i,j) is the rounded vertical displacement for the sample located at (i,j) in one or more samples. In some examples, for each sample (i,j) in the samples (i,j) of the prediction block, Δv x and Δv y can be the same.

[0200] The resulting modified samples can form a modified prediction block in gradient-based prediction refinement. That is, the modified prediction block can be used as a prediction block in gradient-based prediction refinement. The reconstruction unit 310 can reconstruct the current block using the prediction block and the residual block. For example, the reconstruction unit 310 can add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.

[0201] The filter unit 312 can perform one or more filter operations on the reconstructed block. For example, the filter unit 312 can perform a deblocking operation to reduce block effect artifacts along the edges of the reconstructed block. The operations of the filter unit 312 are not necessarily performed in all examples.

[0202] Video decoder 300 may store the reconstructed blocks in DPB 314. As discussed above, DPB 314 may provide reference information, such as samples of the current picture for intra prediction and previously decoded pictures for subsequent motion compensation, to prediction processing unit 304. Additionally, video decoder 300 may output the decoded pictures from DPB for subsequent presentation on a display device, such as Figure 1 display device 118.

[0203] Figure 5 is a flowchart illustrating an example method for decoding video data. The current block may include a current CU. Examples are described Figure 5 with respect to processing circuitry. Examples of processing circuitry include fixed-function and / or programmable circuitry for video encoder 200, such as GBPR unit 227, and fixed-function and / or programmable circuitry for video decoder 300, such as GBPR unit 319.

[0204] In one or more examples, the memory may be configured to store samples of a predicted block. For example, DPB 218 or DPB 314 may be configured to store samples of a predicted block for inter prediction. Intra-block copy may be considered an example inter prediction mode, in which case the block vector for intra-block copy is an example of a motion vector.

[0205] The processing circuitry may determine a predicted block (350) stored in the memory for inter prediction of the current block. The processing circuitry may determine a horizontal displacement and a vertical displacement (e.g., Δv x and Δv y ) for gradient-based prediction refinement (352) of one or more samples of the predicted block. As an example, the processing circuitry may determine an inter prediction mode for inter prediction of the current block. In some examples, the processing circuitry may determine the horizontal displacement and the vertical displacement based on the determined inter prediction mode.

[0206] The processing circuit may round the horizontal displacement and the vertical displacement to the same precision level (354) for different inter-frame prediction modes. For example, the current block may be a first current block, the prediction block may be a first prediction block, the horizontal displacement and the vertical displacement may be a first horizontal displacement and a first vertical displacement, and the rounded horizontal displacement and vertical displacement may be a first rounded horizontal displacement and a first rounded vertical displacement. In some examples, the processing circuit may determine a second prediction block for inter-frame prediction of a second current block, and determine a second horizontal displacement and a second vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block. The processing circuit may round the second horizontal displacement and the second vertical displacement to the same precision level to which the first horizontal displacement and the first vertical displacement are rounded, to generate a second rounded horizontal displacement and a second rounded vertical displacement.

[0207] In some cases, the inter-frame prediction mode used for inter-frame prediction of a first current block and the inter-frame prediction mode for a second current block may be different. For example, the first mode in different inter-frame prediction modes is an affine mode, and the second mode in different inter-frame prediction modes is a bi-directional optical flow (BDOF) mode.

[0208] The precision level to which the horizontal displacement and the vertical displacement are rounded may be predefined or signaled. As an example, the precision level is 1 / 64.

[0209] The processing circuit may determine one or more refinement offsets (356) based on the rounded horizontal displacement and the vertical displacement. For example, the processing circuit may determine a first gradient based on a first set of samples among one or more samples of the prediction block (e.g., determining g x (i,j) using the samples of the prediction block described above), and determine a second gradient based on a second set of samples among one or more samples of the prediction block (e.g., determining g y (i,j) using the samples of the prediction block described above). The processing circuit may determine one or more refinement offsets based on the rounded horizontal displacement and the vertical displacement, and the first gradient and the second gradient. In some examples, if the value of one or more refinement offsets is too high (e.g., greater than a threshold), the processing circuit may clip one or more refinement offsets.

[0210] The processing circuit may modify one or more samples of the prediction block based on the determined one or more refinement offsets or the clipped one or more refinement offsets to generate a modified prediction block (e.g., one or more modified samples for forming the modified prediction block). For example, the processing circuit may determine: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y(i, j), where g x (i, j) is the first gradient of the sample located at (i, j) in one or more samples, Δv x (i, j) is the rounded horizontal displacement of the sample located at (i, j) in one or more samples, g y (i, j) is the second gradient of the sample located at (i, j) in one or more samples, and Δv y (i, j) is the rounded vertical displacement of the sample located at (i, j) in one or more samples. In some examples, for each sample (i, j) in the sample (i, j) of the prediction block, Δv x and Δv y can be the same.

[0211] The processing circuit can decode (e.g., encode or decode) the current block (360) based on the modified prediction block (e.g., one or more modified samples of the modified prediction block). For example, for video decoding, the processing circuit (e.g., video decoder 300) can reconstruct the current block based on the modified prediction block (e.g., by adding one or more modified samples to the received residual values). For video encoding, the processing circuit (e.g., video encoder 200) can determine the residual value (e.g., of the residual block) between the current block and the modified prediction block (e.g., one or more modified samples of the modified prediction block), and signal the information for indicating the residual value.

[0212] The following describes a non - restrictive illustrative list of examples of the present disclosure.

[0213] Example 1. A method for decoding video data, the method comprising: determining one or more prediction blocks for performing inter - frame prediction on a current block; determining an inter - frame prediction mode for performing inter - frame prediction on the current block; based on the determined inter - frame prediction mode, determining at least one of a horizontal displacement or a vertical displacement for gradient - based prediction refinement of one or more samples of one or more prediction blocks; modifying one or more samples of one or more prediction blocks based on the determined at least one of the horizontal displacement or the vertical displacement to generate one or more modified samples; and reconstructing the current block based on the one or more modified samples.

[0214] Example 2. The method according to Example 1, wherein determining the inter - frame prediction mode includes: determining that inter - frame prediction is applied to a current block having a small size, the method further comprising: rounding the motion vector for the current block to an integer motion vector, wherein determining at least one of the horizontal displacement or the vertical displacement includes: determining at least one of the horizontal displacement or the vertical displacement based on the remainder of the rounding.

[0215] Example 3. The method according to Example 1, wherein the current block includes a first block and the prediction block includes a first prediction block, the method further comprising: determining an inter prediction mode for a second block having a small size; and based on the inter prediction mode for the second block being the merge mode, using gradient-based prediction refinement to modify one or more samples of a second prediction block for the second block, wherein gradient-based prediction refinement is disabled for blocks having a small size that are not inter predicted in the merge mode.

[0216] Example 4. The method according to Example 1, wherein the current block includes a first block and the prediction block includes a first prediction block, the method further comprising: determining an inter prediction mode for a second block having a small size; and based on the inter prediction mode for the second block not being an integer motion mode, using gradient-based prediction refinement to modify one or more samples of a second prediction block for the second block, wherein gradient-based prediction refinement is disabled for blocks having an integer motion mode, wherein in the integer motion mode, one or more signaled motion vectors are integers.

[0217] Example 5. The method according to Example 1, wherein determining the inter prediction mode includes: determining that the current block is inter predicted in the merge mode, the method further comprising: rounding a motion vector derived from a spatially or temporally adjacent block, wherein determining at least one of a horizontal displacement or a vertical displacement includes: determining at least one of a horizontal displacement or a vertical displacement based on a remainder of the rounding.

[0218] Example 6. The method according to Example 1, wherein determining the inter prediction mode includes: determining that the current block is inter predicted in the merge mode using a motion vector difference, the method further comprising: rounding the motion vector difference, wherein determining at least one of a horizontal displacement or a vertical displacement includes: determining at least one of a horizontal displacement or a vertical displacement based on a remainder of the rounding.

[0219] Example 7. The method according to Example 1, wherein determining the inter prediction mode includes: determining that the current block is inter predicted in a decoder-side motion vector refinement mode, the method further comprising: using an original motion vector to determine an original dual prediction block; determining DistOrig based on a difference between the dual prediction blocks; and determining DistNew based on a search within an integer displacement range of the rounded original vector, wherein determining at least one of a horizontal displacement or a vertical displacement includes: performing bidirectional optical flow (BDOF) based on DistNew being less than DistOrig to determine at least one of a horizontal displacement or a vertical displacement.

[0220] Example 8. The method according to Example 1, wherein determining an inter-frame prediction mode includes: determining that a current block is inter-frame predicted in an affine mode, and wherein determining at least one of a horizontal displacement or a vertical displacement includes: determining at least one of a horizontal displacement or a vertical displacement based on positions of sub-blocks of the current block.

[0221] Example 9. The method according to Example 1, further comprising: rounding at least one of a horizontal displacement and a vertical displacement to a same predefined precision for different inter-frame prediction modes, wherein modifying one or more samples includes: modifying one or more samples of one or more prediction blocks based on at least one of the rounded horizontal displacement or vertical displacement to generate one or more modified samples.

[0222] Example 10. The method according to Example 1, further comprising: clipping one or more modified samples, wherein reconstructing the current block includes: reconstructing the current block based on one or more clipped and modified samples.

[0223] Example 11. A method comprising a combination of one or more features according to any one of Examples 1-10.

[0224] Example 12. A method for encoding video data, the method comprising: determining one or more prediction blocks for inter-frame prediction of a current block; determining an inter-frame prediction mode for inter-frame prediction of the current block; based on the determined inter-frame prediction mode, determining at least one of a horizontal displacement or a vertical displacement for gradient-based prediction refinement of one or more samples of one or more prediction blocks; based on the determined at least one of a horizontal displacement or a vertical displacement, modifying one or more samples of one or more prediction blocks to generate one or more modified samples; determining a residual value based on the current block and the one or more modified samples; and signaling information for indicating the residual value.

[0225] Example 13. The method according to Example 12, wherein determining the inter-frame prediction mode includes: determining that inter-frame prediction is applied to a current block having a small size, the method further comprising: rounding a motion vector for the current block to an integer motion vector, and wherein determining at least one of a horizontal displacement or a vertical displacement includes: determining at least one of a horizontal displacement or a vertical displacement based on a remainder of the rounding.

[0226] Example 14. The method according to Example 12, wherein the current block includes a first block and the prediction block includes a first prediction block, the method further comprising: determining an inter prediction mode for a second block having a small size; and based on the inter prediction mode for the second block being a merge mode, using gradient-based prediction refinement to modify one or more samples of a second prediction block for the second block, wherein gradient-based prediction refinement is disabled for blocks having a small size that are not inter predicted in the merge mode.

[0227] Example 15. The method according to Example 12, wherein the current block includes a first block and the prediction block includes a first prediction block, the method further comprising: determining an inter prediction mode for a second block having a small size; and based on the inter prediction mode for the second block not being an integer motion mode, using gradient-based prediction refinement to modify one or more samples of a second prediction block for the second block, wherein gradient-based prediction refinement is disabled for blocks having an integer motion mode, wherein, in the integer motion mode, one or more signaled motion vectors are integers.

[0228] Example 16. The method according to Example 12, wherein determining the inter prediction mode includes: determining that the current block is inter predicted in the merge mode, the method further comprising: rounding a motion vector derived from a spatially or temporally adjacent block, wherein determining at least one of a horizontal displacement or a vertical displacement includes: determining at least one of a horizontal displacement or a vertical displacement based on a remainder of the rounding.

[0229] Example 17. The method according to Example 12, wherein determining the inter prediction mode includes: determining that the current block is inter predicted in the merge mode using motion vector difference, the method further comprising: rounding the motion vector difference, wherein determining at least one of a horizontal displacement or a vertical displacement includes: determining at least one of a horizontal displacement or a vertical displacement based on a remainder of the rounding.

[0230] Example 18. The method according to Example 12, wherein determining the inter prediction mode includes: determining that the current block is inter predicted in the decoder-side motion vector refinement mode, the method further comprising: using an original motion vector to determine an original dual prediction block; determining DistOrig based on a difference between the dual prediction blocks; and determining DistNew based on a search within an integer displacement range of the rounded original vector, wherein determining at least one of a horizontal displacement or a vertical displacement includes: performing bidirectional optical flow (BDOF) based on DistNew being less than DistOrig to determine at least one of a horizontal displacement or a vertical displacement.

[0231] Example 19. The method according to Example 12, wherein determining an inter-frame prediction mode includes: determining that a current block is inter-frame predicted in an affine mode, and wherein determining at least one of a horizontal displacement or a vertical displacement includes: determining at least one of a horizontal displacement or a vertical displacement based on positions of sub-blocks of the current block.

[0232] Example 20. The method according to Example 12, further comprising: rounding at least one of a horizontal displacement and a vertical displacement to a same predefined precision for different inter-frame prediction modes, wherein modifying one or more samples includes: modifying one or more samples of one or more prediction blocks based on at least one of the rounded horizontal displacement or vertical displacement to generate one or more modified samples.

[0233] Example 21. The method according to Example 12, further comprising: clipping one or more modified samples, wherein determining a residual value includes: determining a residual value based on the current block and one or more clipped and modified samples.

[0234] Example 22. A method comprising a combination of features according to any one of Examples 12 - 21.

[0235] Example 23. A device for decoding video data, the device comprising a memory configured to store video data including prediction blocks and a video decoder including at least one fixed-function or programmable circuit, wherein the video decoder is configured to perform the method according to any one of Examples 1 - 11.

[0236] Example 24. The device according to Example 23, further comprising: a display configured to display the decoded video data.

[0237] Example 25. The device according to any one of Examples 23 and 24, wherein the device includes one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0238] Example 26. A device for encoding video data, the device comprising a memory configured to store video data including prediction blocks and a video encoder including at least one fixed-function or programmable circuit, wherein the video encoder is configured to perform the method according to any one of Examples 12 - 22.

[0239] Example 27. The device according to Example 26, further comprising: a camera configured to capture video data to be encoded.

[0240] Example 28. The device according to any one of Examples 26 and 27, wherein the device includes one or more of a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0241] Example 29. An apparatus for decoding video data, the apparatus comprising a unit for performing the method according to any one of Examples 1-11.

[0242] Example 30. An apparatus for encoding video data, the apparatus comprising a unit for performing the method according to any one of Examples 12-22.

[0243] Example 31. A computer-readable storage medium including instructions stored thereon, the instructions, when executed, causing one or more processors of an apparatus for decoding video data to perform the method according to any one of Examples 1-11.

[0244] Example 32. A computer-readable storage medium including instructions stored thereon, the instructions, when executed, causing one or more processors of an apparatus for encoding video data to perform the method according to any one of Examples 12-22.

[0245] It should be appreciated that, in accordance with examples, certain actions or events of any of the techniques described herein may be performed in a different order, may be added, combined, or entirely omitted (e.g., not all described actions or events are necessary for implementing the techniques). Additionally, in certain examples, the actions or events may be performed concurrently rather than sequentially, such as by multithreading, interrupt processing, or by multiple processors.

[0246] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium or a communication medium including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0247] By way of example and not limitation, such a computer-readable storage medium can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave) are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead are directed to non-transitory, tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0248] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, as used herein, the terms “processor” and “processing circuitry” can refer to any one of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Further, the techniques can be implemented entirely in one or more circuits or logic elements.

[0249] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a group of ICs (e.g., a chip set). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but need not be implemented by different hardware units. Rather, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) in conjunction with appropriate software and / or firmware.

[0250] Various examples have been described. These and other examples are within the scope of the appended claims.

Claims

1. A method for decoding video data, the method comprising: Determine a motion vector for inter - frame prediction of the current block; Based on the determined motion vector, determine a prediction block for inter - frame prediction of the current block; According to gradient - based prediction refinement, determine one or more refinement offsets for modifying one or more samples of the prediction block, wherein determining the one or more refinement offsets includes: Determine a horizontal displacement and a vertical displacement of the gradient - based prediction refinement for the one or more samples of the prediction block, wherein the horizontal displacement and the vertical displacement are used to determine the one or more refinement offsets for modifying the one or more samples of the prediction block according to the gradient - based prediction refinement; Round the horizontal displacement and the vertical displacement for determining the one or more refinement offsets to the same precision level for different inter - frame prediction modes; Determine a first gradient based on a first sample set of the one or more samples of the prediction block; and Determine a second gradient based on a second sample set of the one or more samples of the prediction block, wherein determining the one or more refinement offsets includes: based on the first gradient for the sample at (i,j) among the one or more samples, the rounded horizontal displacement for the sample at (i,j) among the one or more samples, the second gradient for the sample at (i,j) among the one or more samples, and the rounded vertical displacement for the sample at (i,j) among the one or more samples, determine the one or more refinement offsets for modifying the one or more samples of the prediction block according to the gradient - based prediction refinement; According to the gradient - based prediction refinement, modify the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block, wherein modifying the one or more samples includes selecting a shearing process from at least one of the following: Shear the one or more refinement offsets to generate sheared refinement offsets, and modify the one or more samples based on the sheared refinement offsets to generate the modified prediction block, or Modify the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a refinement result, and shear the refinement result to generate the modified prediction block; and Reconstruct the current block based on the modified prediction block.

2. The method according to claim 1, wherein, The first mode among the different inter - frame prediction modes is the Bidirectional Optical Flow (BDOF) mode, and the second mode among the different inter - frame prediction modes is a mode different from the BDOF mode.

3. The method according to claim 1, wherein, The first mode among the different inter - frame prediction modes is the Affine mode, and the second mode among the different inter - frame prediction modes is the Bidirectional Optical Flow (BDOF) mode.

4. The method according to claim 1, further comprising: Determine an inter - frame prediction mode for inter - frame prediction of the current block, wherein determining the horizontal displacement and the vertical displacement includes: determining the horizontal displacement and the vertical displacement based on the determined inter - frame prediction mode.

5. The method according to claim 1, wherein, The precision level is 1 / 64.

6. The method according to claim 1, wherein, The first mode among the different inter-frame prediction modes is an affine mode, and the second mode among the different inter-frame prediction modes is a mode different from the affine mode.

7. The method according to claim 1, wherein, Determining the one or more refinement offsets further includes determining: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient of the sample located at (i,j) in the one or more samples, Δv x (i,j) is the rounded horizontal displacement of the sample located at (i,j) in the one or more samples, g y (i,j) is the second gradient of the sample located at (i,j) in the one or more samples, and Δv y (i,j) is the rounded vertical displacement of the sample located at (i,j) in the one or more samples.

8. The method according to claim 1, wherein, The prediction block is a first prediction block, the current block is a first current block, the horizontal displacement and the vertical displacement are a first horizontal displacement and a first vertical displacement, the one or more refinement offsets are a first one or more refinement offsets, the rounded horizontal displacement and the rounded vertical displacement are a first rounded horizontal displacement and a first rounded vertical displacement, and the modified prediction block is a first modified prediction block. The method further includes: Determining a second prediction block for inter-frame prediction of a second current block; Determining a second horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block; Rounding the second horizontal displacement and the vertical displacement to the same precision level as the precision level to which the first horizontal displacement and the first vertical displacement are rounded, to generate a second rounded horizontal displacement and a second rounded vertical displacement; Determining a second one or more refinement offsets based on the second rounded horizontal displacement and the second rounded vertical displacement; Modifying the one or more samples of the second prediction block based on the determined second one or more refinement offsets to generate a second modified prediction block; and Reconstructing the second current block based on the second modified prediction block.

9. A method for encoding video data, the method comprising: Determining a motion vector for inter-frame prediction of a current block; Based on the determined motion vector, determining a prediction block for inter-frame prediction of the current block; Determining, according to gradient-based prediction refinement, one or more refinement offsets for modifying one or more samples of the prediction block determined based on the determined motion vector, wherein determining the one or more refinement offsets includes: Determining a horizontal displacement and a vertical displacement for the gradient-based prediction refinement of the one or more samples of the prediction block, wherein the horizontal displacement and the vertical displacement are used to determine the one or more refinement offsets for modifying the one or more samples of the prediction block according to the gradient-based prediction refinement; Rounding the horizontal displacement and the vertical displacement for determining the one or more refinement offsets to the same precision level for different inter-frame prediction modes; Determining a first gradient based on a first sample set of the one or more samples of the prediction block; and Determining a second gradient based on a second sample set of the one or more samples of the prediction block Determining the one or more refinement offsets includes: determining, according to gradient-based prediction refinement, one or more refinement offsets for modifying one or more samples of the prediction block based on the first gradient for the sample located at (i,j) in the one or more samples, the rounded horizontal displacement for the sample located at (i,j) in the one or more samples, the second gradient for the sample located at (i,j) in the one or more samples, and the rounded vertical displacement for the sample located at (i,j) in the one or more samples; According to the gradient-based prediction refinement, modifying the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block, wherein modifying the one or more samples includes selecting a clipping process from at least one of the following: Clipping the one or more refinement offsets to generate clipped refinement offsets, and modifying the one or more samples based on the clipped refinement offsets to generate the modified prediction block, or Modifying the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a refinement result, and clipping the refinement result to generate the modified prediction block; Determining a residual value for indicating a difference between the current block and the modified prediction block; and Signaling information for indicating the residual value.

10. The method according to claim 9, wherein, The first mode among the different inter-frame prediction modes is the bidirectional optical flow (BDOF) mode, and the second mode among the different inter-frame prediction modes is a mode different from the BDOF mode.

11. The method according to claim 9, wherein, The first mode among the different inter-frame prediction modes is the affine mode, and the second mode among the different inter-frame prediction modes is the bidirectional optical flow (BDOF) mode.

12. The method according to claim 9, further comprising: Determining an inter-frame prediction mode for inter-frame predicting the current block, wherein determining the horizontal displacement and the vertical displacement includes: determining the horizontal displacement and the vertical displacement based on the determined inter-frame prediction mode.

13. The method according to claim 9, wherein, The accuracy level is 1 / 64.

14. The method according to claim 9, wherein, The first mode among the different inter-frame prediction modes is the affine mode, and the second mode among the different inter-frame prediction modes is a mode different from the affine mode.

15. The method according to claim 9, wherein, Determining the one or more refinement offsets further includes determining: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient of the sample located at (i,j) in the one or more samples, Δv x (i,j) is the rounded horizontal displacement of the sample located at (i,j) in the one or more samples, g y (i,j) is the second gradient of the sample located at (i,j) in the one or more samples, and Δv y (i,j) is the rounded vertical displacement of the sample located at (i,j) in the one or more samples.

16. The method according to claim 9, wherein, The prediction block is a first prediction block, the current block is a first current block, the horizontal displacement and the vertical displacement are a first horizontal displacement and a first vertical displacement, the one or more refinement offsets are a first one or more refinement offsets, the rounded horizontal displacement and the rounded vertical displacement are a first rounded horizontal displacement and a first rounded vertical displacement, the modified prediction block is a first modified prediction block, and the residual value includes a first residual value. The method further includes: Determining a second prediction block for inter-frame predicting a second current block; Determining a second horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block; Round the second horizontal displacement and vertical displacement to the same precision level as the precision level to which the first horizontal displacement and the first vertical displacement are rounded to generate a second rounded horizontal displacement and vertical displacement; Determine a second one or more refinement offsets based on the second rounded horizontal displacement and vertical displacement; Modify the one or more samples of the second prediction block based on the determined second one or more refinement offsets to generate a second modified prediction block; Determine a second residual value for indicating a difference between the second current block and the second modified prediction block; and Signal information for indicating the second residual value.

17. An apparatus for decoding video data, the apparatus comprising: A memory configured to store one or more samples of a prediction block; And A processing circuit configured to: Determine a motion vector for performing inter-frame prediction on a current block; Based on the determined motion vector, determine a prediction block for performing inter-frame prediction on the current block; Determine one or more refinement offsets for modifying the one or more samples of the prediction block according to gradient-based prediction refinement, wherein, to determine the one or more refinement offsets, the processing circuit is configured to: Determine a horizontal displacement and a vertical displacement of the gradient-based prediction refinement for the one or more samples of the prediction block, wherein the horizontal displacement and the vertical displacement are used to determine the one or more refinement offsets for modifying the one or more samples of the prediction block according to the gradient-based prediction refinement; Round the horizontal displacement and the vertical displacement for determining the one or more refinement offsets to the same precision level for different inter-frame prediction modes; Determine a first gradient based on a first sample set of the one or more samples of the prediction block; and Determine a second gradient based on a second sample set of the one or more samples of the prediction block, wherein, to determine the one or more refinement offsets, the processing circuit is configured to: based on the first gradient for the sample located at (i,j) among the one or more samples, the rounded horizontal displacement for the sample located at (i,j) among the one or more samples, the second gradient for the sample located at (i,j) among the one or more samples, and the rounded vertical displacement for the sample located at (i,j) among the one or more samples, determine the one or more refinement offsets for modifying the one or more samples of the prediction block according to gradient-based prediction refinement; Modify the one or more samples of the prediction block based on the determined one or more refinement offsets according to the gradient-based prediction refinement to generate a modified prediction block, wherein, to modify the one or more samples, the processing circuit is configured to select a clipping process from at least one of the following: Clip the one or more refinement offsets to generate clipped refinement offsets and modify the one or more samples based on the clipped refinement offsets to generate a modified prediction block, or Modify one or more samples of the prediction block based on the determined one or more refinement offsets to generate a refinement result, and clip the refinement result to generate the modified prediction block; and Decode the current block based on the modified prediction block.

18. The device according to claim 17, wherein, To decode the current block, the processing circuit is configured to: reconstruct the current block based on the modified prediction block.

19. The device according to claim 17, wherein, To decode the current block, the processing circuit is configured to: Determine a residual value for indicating a difference between the current block and the modified prediction block; and Signal information for indicating the residual value.

20. The device according to claim 17, wherein, The first mode among the different inter-frame prediction modes is the bi-directional optical flow (BDOF) mode, and the second mode among the different inter-frame prediction modes is a mode different from the BDOF mode.

21. The device according to claim 17, wherein, The first mode among the different inter-frame prediction modes is the affine mode, and the second mode among the different inter-frame prediction modes is the bi-directional optical flow (BDOF) mode.

22. The device according to claim 17, wherein, The processing circuit is configured to: Determine an inter-frame prediction mode for performing inter-frame prediction on the current block, wherein, to determine the horizontal displacement and the vertical displacement, the processing circuit is configured to: determine the horizontal displacement and the vertical displacement based on the determined inter-frame prediction mode.

23. The device according to claim 17, wherein, The accuracy level is 1 / 64.

24. The device according to claim 17, wherein, The first mode among the different inter-frame prediction modes is the affine mode, and the second mode among the different inter-frame prediction modes is a mode different from the affine mode.

25. The device according to claim 17, wherein, To determine the one or more refinement offsets, the processing circuit is further configured to determine: g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient of the sample located at (i,j) in the one or more samples, Δv x (i,j) is the rounded horizontal displacement of the sample located at (i,j) in the one or more samples, g y (i,j) is the second gradient of the sample located at (i,j) in the one or more samples, and Δv y (i,j) is the rounded vertical displacement of the sample located at (i,j) in the one or more samples.

26. The device according to claim 17, wherein, The prediction block is the first prediction block, the current block is the first current block, the horizontal displacement and the vertical displacement are the first horizontal displacement and the first vertical displacement, the one or more refinement offsets are the first one or more refinement offsets, the rounded horizontal displacement and the rounded vertical displacement are the first rounded horizontal displacement and the first rounded vertical displacement, and the modified prediction block is the first modified prediction block, and wherein, the processing circuit is configured to: Determine a second prediction block for performing inter-frame prediction on a second current block; Determine a second horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block; Round the second horizontal displacement and the vertical displacement to the same accuracy level to which the first horizontal displacement and the first vertical displacement are rounded, to generate a second rounded horizontal displacement and a vertical displacement; Determine a second one or more refinement offsets based on the second rounded horizontal displacement and the vertical displacement; Modify one or more samples of the second prediction block based on the determined second one or more refinement offsets to generate a second modified prediction block; and Decode the second current block based on the second modified prediction block.

27. The device according to claim 17, further comprising: A display configured to display decoded video data.

28. The device according to claim 17, further comprising: A camera configured to capture the video data to be encoded.

29. The device according to claim 17, wherein, The device includes one or more of a camera, a computer, a wireless communication device, a broadcast receiver device, or a set-top box.

30. A non - transitory computer - readable storage medium storing instructions that, when executed, cause one or more processors to perform the following operations: Determine a motion vector for performing inter - prediction on a current block; Based on the determined motion vector, determine a prediction block for performing inter - prediction on the current block; According to gradient - based prediction refinement, determine one or more refinement offsets for modifying one or more samples of the prediction block, wherein, The instructions that cause the one or more processors to determine the one or more refinement offsets include instructions for causing the one or more processors to perform the following operations: Determine a horizontal displacement and a vertical displacement of the gradient-based prediction refinement for the one or more samples of the prediction block, wherein the horizontal displacement and the vertical displacement are used to determine the one or more refinement offsets for modifying the one or more samples of the prediction block according to the gradient-based prediction refinement; Round the horizontal displacement and the vertical displacement for determining the one or more refinement offsets to the same precision level for different inter-frame prediction modes; Determine a first gradient based on a first set of samples of the one or more samples of the prediction block; and Determine a second gradient based on a second set of samples of the one or more samples of the prediction block, wherein the instructions that cause the one or more processors to determine the one or more refinement offsets include instructions for causing the one or more processors to perform the following operations: determine the one or more refinement offsets for modifying the one or more samples of the prediction block according to the gradient-based prediction refinement based on the first gradient for the sample located at (i, j) among the one or more samples, the rounded horizontal displacement for the sample located at (i, j) among the one or more samples, the second gradient for the sample located at (i, j) among the one or more samples, and the rounded vertical displacement for the sample located at (i, j) among the one or more samples; Modify the one or more samples of the prediction block based on the determined one or more refinement offsets according to the gradient-based prediction refinement to generate a modified prediction block, wherein the instructions that cause the one or more processors to modify the one or more samples include instructions for causing the one or more processors to select a clipping process from at least one of the following: Clip the one or more refinement offsets to generate clipped refinement offsets, and modify the one or more samples based on the clipped refinement offsets to generate a modified prediction block, or Modify the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a refinement result, and clip the refinement result to generate the modified prediction block; and Decode the current block based on the modified prediction block.

31. The non - transitory computer - readable storage medium according to claim 30, wherein, The first mode among the different inter-frame prediction modes is an affine mode, and the second mode among the different inter-frame prediction modes is a bi-directional optical flow (BDOF) mode.

32. An apparatus for decoding video data, the apparatus comprising: A unit for determining a motion vector for inter-frame prediction of a current block; A unit for determining a prediction block for inter-frame prediction of the current block based on the determined motion vector; A unit for determining one or more refinement offsets for modifying one or more samples of the prediction block according to gradient-based prediction refinement, wherein the unit for determining the one or more refinement offsets includes: A unit for determining a horizontal displacement and a vertical displacement of the gradient-based prediction refinement for the one or more samples of the prediction block, wherein the horizontal displacement and the vertical displacement are used to determine one or more refinement offsets for modifying the one or more samples of the prediction block according to the gradient-based prediction refinement; A unit for rounding the horizontal displacement and the vertical displacement for determining the one or more refinement offsets to the same precision level for different inter-frame prediction modes; A unit for determining a first gradient based on a first sample set of the one or more samples of the prediction block; and A unit for determining a second gradient based on a second sample set of the one or more samples of the prediction block, wherein the unit for determining the one or more refinement offsets further includes: a unit for determining one or more refinement offsets for modifying the one or more samples of the prediction block according to the gradient-based prediction refinement based on the first gradient for the sample located at (i, j) in the one or more samples, the rounded horizontal displacement for the sample located at (i, j) in the one or more samples, the second gradient for the sample located at (i, j) in the one or more samples, and the rounded vertical displacement for the sample located at (i, j) in the one or more samples; A unit for modifying the one or more samples of the prediction block based on the determined one or more refinement offsets according to the gradient-based prediction refinement to generate a modified prediction block, wherein the unit for modifying the one or more samples includes a unit for selecting a shearing process from at least one of the following: Shearing the one or more refinement offsets to generate sheared refinement offsets and modifying the one or more samples based on the sheared refinement offsets to generate a modified prediction block, or Modifying the one or more samples of the prediction block based on the determined one or more refinement offsets to generate a refinement result and shearing the refinement result to generate the modified prediction block; and A unit for decoding the current block based on the modified prediction block.

33. The apparatus according to claim 32, wherein, The first mode in the different inter-frame prediction modes is an affine mode, and the second mode in the different inter-frame prediction modes is a bi-directional optical flow (BDOF) mode.